Decomposing growth in a national HL7 CDA clinical document repository
Talvik, H.-A.; Laur, S.; Vilo, J.; Reisberg, S.
Show abstract
Longitudinal evaluations of national electronic health record repositories often track document counts alone, obscuring changes in content size, structure and standards implementation. We decomposed growth in the Estonian Health Information System across document counts, per-document size, section-level structure and version uptake in a 10% random population sample of 4.97 million HL7 Clinical Document Architecture Release 2 documents from 147,819 patients, spanning 2012--2019 and four prespecified document types. Growth patterns differed by document type. Inpatient summaries increased 48.5% in total content volume despite a 2.4% decline in document counts. Section presence and within-section content were highly skewed; 44.6% of 892 data locations carried one fixed value. Code-system diversity increased from 45 to 79, and version uptake took years: inpatient summaries reached 80% organisational uptake after a median 44 months (95% CI 11--78). This decomposition can guide extraction pipelines, secondary use and standards governance in CDA- and FHIR-based repositories.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Increasing Trust in Real-World Evidence Through Evaluation of Observational Data Quality 93%
- Real World Performance of the 21st Century Cures Act Population Level Application Programming Interface 93%
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 92%
Similar papers in this journal
- Is the quality of hospital EHR data sufficient to evidence its ICHOM outcomes performance in heart failure? A pilot evaluation 92%
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 92%
- FHIR-DHP: A Standardized Clinical Data Harmonisation Pipeline for scalable AI application deployment 92%
Similar papers in this journal
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 94%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 94%
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 92%
Similar papers in this journal
Similar papers in this journal
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 94%
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 92%
- Medication information extraction using local large language models 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.