Analysis Of Errors In Issuer-Side Qhp Apis And Aca Marketplace Machine-Readable Files (Mrfs) In The United States
Nikolayev, D.
Show abstract
The Centers for Medicare and Medicaid Services (CMS) requires Qualified Health Plan (QHP) issuers on the federal health insurance marketplace to publish machine-readable JSON files describing plans, provider networks, and drug formularies, following the QHP Provider & Formulary API specification (index.json, plans.json, providers.json, and drugs.json). These data are intended to support consumer-facing tools that help people compare coverage options. In parallel, CMSs Center for Consumer Information & Insurance Oversight (CCIIO) publishes Health Insurance Exchange Public Use Files (Exchange PUFs) for plan years 2014-2026, including a Machine-readable URL PUF (MR-PUF). The MR-PUF provides issuer-level records with state, a five-digit HIOS issuer ID, the issuers machine-readable index URL for provider/formulary/plan JSON, and a technical point-of-contact email, and is available for plan years 2016-2026. For plan year 2026, CMS distributes the MR-PUF as machine-readable-url-puf.zip under the 2026 Exchange PUFs. In this study, we analyze 59,899 import errors recorded by the HealthPorta Healthcare Data Dashboard across 735 issuer import logs, covering QHP provider and formulary APIs as of December 2025. Our analysis focuses on plan years 2025 and 2026, which account for the overwhelming majority of observed errors, but we also document that some issuers still publish data for older plan years (2018-2024) inside the same JSON files. Among the 735 issuers, 339 exhibit at least one error. We classify error messages into a small set of recurring categories: missing recommended cost_sharing fields (48.0% of all errors), missing required formulary fields (26.5%), and issuer identifier mismatches relative to the CMS machine-readable URL index (24.8%). A further 0.6% of errors reflect invalid JSON responses, and small residual fractions arise from network/TLS failures, HTTP errors, schema anomalies (e.g., overly long plan IDs), and drug-related field omissions. JSON-related problems include non-UTF-8 bytes, HTML or other non-JSON payloads returned at QHP API endpoints, truncated or incomplete responses, and lexical parse errors. By aggregating errors by source (plans.json, providers.json, drugs.json, and index files), plan year (with special attention to 2025-2026), issuer, state, and endpoint URL, we show that a narrow set of issuer-side problems accounts for the vast majority of observed failures. We then analyze, using the QHP GitHub specification, the CMS Exchange PUF documentation and the Healthcare MRF API import model, why the missing fields we observe (formulary, cost_sharing, IDs, tiers, and policy flags) are structurally necessary and how their absence cascades into broken joins and misleading analytics when QHP data are combined with machine-readable file (MRF) rate data. We conclude with concrete recommendations for issuer and vendor pipelines.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Understanding Data Differences across the ENACT Federated Research Network 94%
- MMFP-Tableau: Enabling Precision Mitochondrial Medicine through Integration, Visualization, and Analytics of Clinical and Research Health System Electronic Data 93%
- Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale 92%
Similar papers in this journal
- From months to minutes: creating Hyperion, a novel data management system expediting data insights for oncology research and patient care 93%
- A data management system for precision medicine 93%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 91%
Similar papers in this journal
- FHIR-DHP: A Standardized Clinical Data Harmonisation Pipeline for scalable AI application deployment 91%
- Can we trust the prediction model? Demonstrating the importance of external validation by investigating the COVID-19 Vulnerability (C-19) Index across an international network of observational healthcare datasets 91%
- Retrospective development and evaluation of prognostic models for exacerbation event prediction in patients with Chronic Obstructive Pulmonary Disease using data self-reported to a digital health application 90%
Similar papers in this journal
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 94%
- Using computable knowledge mined from the literature to elucidate confounders for EHR-based pharmacovigilance 92%
- Mining for Equitable Health: Assessing the Impact of Missing Data in Electronic Health Records 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.