Skip to main navigation Skip to search Skip to main content

Data preprocessing scripts and standardized prompts used in the study: Enabling interoperability across disparate health data sources using Large Language Models

  • Obinwa Nnabuenyi Ozonze (Creator)
  • Uzondu Dike (Creator)
  • Ogechukwu Mercy Okonor (Ravensbourne University) (Contributor)
  • Taiwo Adedeji (Creator)

Dataset

Description

This ZIP file contains: (1) the exact standardized prompt submitted to all three LLMs across all 36 runs (3 LLMs × 4 schemas × 3 repetitions), (2) the complete set of source-to-target mapping pairs used for preprocessing and standardizing CDM element names across PCORnet, OMOP, Sentinel, and i2b2, and (3) the complete mapping results from the three LLMs for all four Common Data Models (PCORnet, OMOP, Sentinel, i2b2) to FHIR across iterations.
Date made available5 Aug 2026
PublisherPublic Library of Science
Date of data productionJan 2025 - Apr 2025

Cite this