Registry Forge: an open-source end-to-end pipeline for patient-directed SMART on FHIR registries
Boyce, D.; Premasiri, A.; Sullivan, S.; Levine, B.; Vieira, F. G.
Show abstract
Objectives: Patient-directed SMART on FHIR lets registries acquire longitudinal electronic health record data, but the payload requires substantial engineering before use. We present Registry Forge, an open-source pipeline that converts it into research-ready outputs. Materials and Methods: Registry Forge decodes and parses mixed C-CDA, HTML, RTF, PDF, and FHIR inputs, joins records to a canonical patient identifier, and emits a browser-viewable dashboard, an OMOP CDM v5.4 data set, GA4GH Phenopackets v2, a code inventory, and regex extractions of disease-specific narrative content. Results: Applied to the ALS Research Collaborative Study (94 participants, 56 US health systems), it processed 22,686 source files and 1,791 FHIR Bundles (109,599 resources); only 15.0% of files were full C-CDA. Discussion: This pipeline generalizes to any registry acquiring data through patient-directed SMART on FHIR. Conclusion: Registry Forge closes the acquisition-to-analysis gap with no server infrastructure and is openly available.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Increasing Trust in Real-World Evidence Through Evaluation of Observational Data Quality 94%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 93%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
Similar papers in this journal
- MMFP-Tableau: Enabling Precision Mitochondrial Medicine through Integration, Visualization, and Analytics of Clinical and Research Health System Electronic Data 94%
- Development of a COVID-19 Application Ontology for the ACT Network 93%
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 93%
Similar papers in this journal
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 94%
- Genomic Considerations for FHIR; eMERGE Implementation Lessons 92%
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 92%
Similar papers in this journal
Similar papers in this journal
- FHIR-DHP: A Standardized Clinical Data Harmonisation Pipeline for scalable AI application deployment 92%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 90%
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.