Abstract
Background. Health information exchanges (HIEs) improve access to outside records, but transmitted data may still be duplicated, incomplete, or semantically inconsistent. That limits its use in longitudinal review, analytics, quality measurement, and AI-enabled workflows.
Objective. To quantify raw-to-curated retention across clinical categories and illustrate why curation is required before HIE-derived data can perform reliably downstream.
Methods. We conducted a retrospective descriptive analysis of category-level HIE-derived flat data for 8,352 patients in a July 2026 extract. Raw rows were counted before Predoc curation; curated rows were counted after normalization, semantic mapping, deduplication or consolidation, and removal of non-informative records. We report pooled net retention and median patient retention.
Results. Predoc processed 60,654,992 raw rows into 25,437,669 curated rows, or 41.9% net retention across all reported categories. In core clinical categories, 34,746,057 raw rows became 11,780,274 curated rows, or 33.9% net retention. Median patient retention was 74.24% for labs, 66.35% for vital signs, 54.00% for problems, 50.89% for medications, 83.75% for procedures, 83.33% for encounters, and 86.21% for immunizations. Large gaps between net and median retention reflected outlier patients with unusually repetitive records.
Conclusion. HIE data is not inherently ready for downstream use. Curation turns exchanged rows into patient-level facts that can be trended, counted, summarized, and audited.
1. Introduction
Health information exchanges solve an important access problem. They allow provider groups, accountable care organizations, and health systems to retrieve outside clinical history that might otherwise require patient recall, manual outreach, or chart-by-chart reconstruction. But access and performance are different outcomes.
In this paper, performant data means data that can support its intended task without substantial reconstruction. It can be queried, compared over time, counted without double counting, summarized without amplifying repeated evidence, and traced back to its source. Raw HIE-derived data often falls short of that standard. The same event may appear in several documents, equivalent concepts may use different names or units, and some rows contain too little clinical content to support a decision.
Exchange standards make transport possible, but they do not complete this interpretive work. C-CDA packages information into structured clinical documents, while FHIR represents clinical information as discrete resources accessed through APIs. LOINC helps identify what a laboratory test measured. Even so, a receiving system may still need to decide whether two records describe the same event, whether units are equivalent, and whether a row contains a usable fact.[1–3]
Prior research has documented this gap. A SMART C-CDA study found 615 observations of errors and data-expression variation across 91 documents from 21 technologies. Other work has shown that multiple continuity-of-care documents for one patient may contain duplicated or conflicting data, and that laboratory results can remain semantically inconsistent even when LOINC codes are present.[4–6]
This study examines that problem in operational HIE-derived data. It measures how many raw rows remain after curation, compares population-level and patient-level retention, and uses deidentified examples to show three distinct tasks: joining different representations into one trend, consolidating repeated copies of the same event, and preserving records that look similar but represent different points in time.
2. Methods
2.1 Study design and data source
This was a retrospective descriptive analysis of category-level HIE-derived flat data processed through Predoc’s normalization and curation pipeline. The July 2026 extract included 8,352 patients. The analysis did not directly process source C-CDA documents; upstream systems had already extracted the records into category-level rows. The results therefore measure the curation required after upstream extraction, not the completeness or accuracy of C-CDA parsing itself.
Raw rows were counted before Predoc curation. Curated rows were counted after normalization, semantic mapping, deduplication or consolidation, removal of non-informative records, and preservation of source provenance. The analysis included both clinical-fact categories such as labs, problems, medications, procedures, encounters, vital signs, immunizations, allergies, and social histories and context categories such as organizations, practitioners, locations, and document references.
2.2 Measures
Note: A category with 10% retention does not mean that 90% of its clinical information was lost. It means that 90% of the raw rows were not preserved as separate curated rows. This means that 90% of the information presented was not unique or was non-informative for clinical interpretation (e.g., A raw allergy object with no allergen, reaction, status, onset date, code, or note may have been successfully transmitted, but it has not delivered usable clinical information. In contrast, a curated allergy fact should identify what the allergy is, what reaction occurred where available, when it applies, and where the evidence came from.)
2.3 Curation operations
3. Results
Across all reported categories, Predoc processed 60,654,992 raw rows into 25,437,669 curated rows, representing 41.9% net retention. Context categories behaved differently from clinical categories: document references, organizations, and related persons were retained at 100% because they preserve source context rather than represent clinical facts to be consolidated.
Within the nine core clinical categories, 34,746,057 raw rows became 11,780,274 curated rows, or 33.9% net retention. Median patient retention was higher in most categories, indicating that a smaller number of patients with unusually repetitive records accounted for a disproportionate share of the raw volume.
Immunizations showed the clearest divergence: 23.84% net retention versus 86.21% median patient retention. For most patients, much of the immunization content remained. A smaller number of patients had the same vaccine, date, and dose carried across many visits and continuity-of-care documents, producing substantial pooled duplication.
Labs showed 54.87% net retention and 74.24% median patient retention. In one outlier record, approximately 372,000 raw lab rows were reduced to approximately 3,000 curated rows. The curated record still contained a large laboratory history, but its scale was clinically plausible rather than dominated by repeated source representations.
These outliers are operationally important, not merely statistical noise. Patients with the largest records are often medically complex, have many encounters, and are among those for whom a coherent longitudinal record matters most. Net retention captures the total processing burden they create; median patient retention shows what curation looks like for the typical patient.
4. Clinical examples: three forms of curation
4.1 Different representations, one longitudinal trend
A complete blood count with differential may record one red blood cell result as “3.9 MIL/uL” and a later result as “4.26 M/uL.” These are not duplicates: both values should remain in the record. The curation task is to recognize that MIL/uL and M/uL express the same unit family, map both observations to the same RBC concept, preserve their dates and values, and make them eligible for one longitudinal trend.
LOINC 789-8 identifies automated erythrocyte count in blood, and its example units include 10*6/uL. A standard code helps establish what was measured, but trendability still depends on compatible units, values, timing, and source context.[2,3,6] Without that normalization, a clinician may infer the relationship while an analytics or AI application treats the observations as unrelated.
4.2 Many rows, one result
A different laboratory example shows the opposite failure mode. One red blood cell distribution width result appeared 32 times across 30 outpatient progress notes and two episode-summary documents with 15 distinct titles. The clinical content was identical in every row: LOINC 788-0, result 14.5, the same laboratory timestamp, and the same CBC with differential report.
The rows differed in observation identifiers, diagnostic-report identifiers, interpreter references, document titles, rendered narratives, and unit completeness. Thirty omitted both unit fields; two supplied “%.” Those differences are useful provenance, but they do not create 32 clinical observations. The curated representation is one result with links back to the supporting source documents.
4.3 Similar records require selective—not blanket—consolidation
Problems and procedures require more than exact matching. In one deidentified record, Stage 3B chronic kidney disease appeared four times across an ambulatory summary and encounter-summary problem lists. The combined SNOMED and ICD-10 coding, display name, date, and active status were consistent; only the encounter and source links changed. These rows described one active condition and could be consolidated without losing evidence.
A left-foot incision, drainage, and debridement procedure also appeared four times across hospital progress notes and an aggregated continuity-of-care document. The narrative format and source links varied, but procedure code 28005, the January 15, 2026 date, and the providers were consistent. Treating the rows as four independent procedures would create a false event count.
Hyperlipidemia demonstrated why similar-looking rows should not always be merged. Six records shared a display name, SNOMED code, and active status, but some were dated 2021 and others 2025. Those dates may represent distinct assertions in the patient’s longitudinal history. The correct curation decision is therefore selective: consolidate repeated evidence of the same event while preserving clinically meaningful differences in time and context.
5. Discussion
5.1 Retention is a usability metric, not a completeness metric
Retention answers how much transformation the incoming feed required before it could support downstream use. It does not answer how complete the HIE was, nor does it imply that every removed row was a lost clinical fact. A low retention rate may reflect substantial duplication or non-informative content; a high rate may indicate mostly unique data or a category intentionally preserved for provenance.
Net and median patient retention are complementary. Net retention quantifies population-scale burden: storage, processing, duplicate volume, and engineering work. Median patient retention describes the typical patient and is less affected by a few enormous records. Reporting both avoids two errors: presenting outlier-heavy pooled data as typical, or ignoring the operational burden created by the very patients with the most complex histories.
5.2 Why uncurated HIE data can hinder AI applications
For the typical patient, median lab retention was 74.24%. Stated relative to the curated output, the raw feed contained about 35 additional lab rows for every 100 rows ultimately retained. Before an AI application can summarize the chart, identify a trend, or reason over the patient timeline, it must resolve that extra volume: determine which rows are duplicates, which units are equivalent, which observations belong to the same concept, which dates anchor the event, and which sources should be retained as evidence.
That preprocessing burden can increase embedding and inference volume, consume context that could hold more diverse clinical evidence, and return several near-identical records during retrieval. Repetition can also distort apparent evidence strength: a diagnosis documented in 20 source documents is still one condition, not 20 independent findings. Semantic fragmentation creates the opposite problem, as the RBC example shows: all values may be present, yet the application can still miss the trend.
The study did not directly measure model accuracy, latency, or token cost. It does, however, quantify the data volume and patient-level variation that downstream AI systems would otherwise have to reconcile. Curation moves that work into a governed data layer, where the logic and provenance can be inspected before an AI application performs its intended task.
5.3 Implications for provider groups and health systems
The same issues affect human and analytical workflows. Care managers must mentally reconcile repeated problems and procedures. Quality teams must avoid counting duplicate evidence or treating empty rows as proof that a clinical category was addressed. Population-health programs need longitudinal concepts rather than copied-forward problem lists. Data-engineering teams must otherwise rebuild mapping, deduplication, validation, and provenance logic downstream.
The findings suggest that HIE data should be evaluated on more than coverage and transport. A performant data product must also show whether observations can be trended, events can be counted without inflation, source evidence can be audited, and high-volume patient records can be processed without overwhelming the receiving workflow. Predoc’s curation layer is designed to provide that transformation after exchange.
6. Limitations
This analysis reflects one specific sample from a July 2026 extract from Predoc’s HIE-derived flat-data pipeline. Retention may vary by source, specialty, patient mix, extraction method, and upstream vendor behavior. The study did not directly compare source C-CDA documents with curated output or compare multiple HIE vendors.
Retention is not a standalone quality measure. The current analysis does not report duplicate precision, false-merge rates, false-drop rates, or manual validation of a representative sample. It also does not separately quantify the share of records improved through semantic normalization, such as unit-of-measure harmonization, versus deduplication or removal.
The study measures the usability of records received; it does not assess HIE coverage or claim that every possible patient record was available. Finally, the operational implications for clinical review, analytics, and AI are reasoned from the observed data burden and examples. Downstream accuracy, time, cost, and outcomes were not directly measured and should be evaluated in future work.
7. Conclusion
HIE data is not inherently performant. Across 8,352 patients, 60.7 million raw rows became 25.4 million curated rows, while the typical patient retained substantial clinical content in high-value categories. The difference reflects a data layer doing necessary work: normalizing equivalent representations, consolidating repeated events, removing rows without usable clinical meaning, and preserving the source trail.
Transmission makes clinical information available. Curation makes it ready to trend, count, summarize, audit, and use. For provider groups, health systems, and AI applications, that distinction determines whether an HIE feed is merely accessible or actually usable.
References
1. Office of the National Coordinator for Health Information Technology. Consolidated Clinical Document Architecture (C-CDA) Testing & More. Accessed August 4, 2026. Source
2. HL7 International. US Core Laboratory Result Observation Profile. US Core Implementation Guide, version 9.0.0. 2026. Source
3. Regenstrief Institute. LOINC 789-8: Erythrocytes [#/volume] in Blood by Automated count; and LOINC 788-0: Erythrocyte [DistWidth] in Blood by Automated count. LOINC version 2.82. Accessed August 4, 2026. Source
4. D’Amore JD, Mandel JC, Kreda DA, et al. Are Meaningful Use Stage 2 certified EHRs ready for interoperability? Findings from the SMART C-CDA Collaborative. Journal of the American Medical Informatics Association. 2014;21(6):1060–1068. Source
5. Hosseini M, Meade J, Schnitzius J, Dixon BE. Consolidating CCDs from multiple data sources: a modular approach. Journal of the American Medical Informatics Association. 2016;23(2):317–323. Source
6. Lin MC, Vreeman DJ, Huff SM. Investigating the semantic interoperability of laboratory data exchanged using LOINC codes in three large institutions. AMIA Annual Symposium Proceedings. 2011;2011:805–814. Source

.jpg)




