Multi-Model Clinical Validation of an AI-Powered Biomarker Analysis Framework: A Cross-Vendor Benchmark on 4,018 NHANES Patients
Shibakov, D.
Show abstract
BackgroundLarge language models (LLMs) show promise for clinical decision support, yet most validation studies evaluate single models, leaving questions about generalizability and vendor dependence unanswered. We assessed whether a standardized biomarker analysis framework maintains clinical-grade accuracy across multiple LLMs from independent providers. MethodsWe developed a structured prompt-based framework for detecting eight clinical patterns (insulin resistance, diabetes, cardiovascular disease risk, chronic kidney disease risk, systemic inflammation, nutrient deficiency, liver risk, and anemia) from laboratory biomarkers. We evaluated five LLMs from four providers--Grok-3 (xAI), GPT-4o and GPT-4o-mini (OpenAI), Claude Haiku 4.5 (Anthropic), and Gemini 2.0 Flash (Google)--using identical system prompts and inputs on 4,018 adults from the CDC NHANES 2017-2018. Ground truth was established using published clinical criteria (ADA, AHA, KDIGO, WHO). Performance was measured by F1 score with 95% confidence intervals, sensitivity, specificity, and positive predictive value. ResultsAll five models achieved clinical-grade performance (F1 > 0.86) on eight evaluable patterns. Mean F1 scores ranged from 0.865 (95% CI: 0.799-0.931) for GPT-4o-mini to 0.963 (95% CI: 0.930-0.996) for Grok-3. Flagship models significantly outperformed economy-tier models (mean F1: 0.940 vs 0.881; paired t-test p=0.004). Grok-3 achieved near-perfect scores on liver risk (F1=1.000), anemia (0.999), and nutrient deficiency (0.997). Cardiovascular disease risk was the most challenging pattern (F1 range: 0.853-0.885). JSON parse rates exceeded 99.9% for all models. Total benchmark cost was approximately $59 USD. ConclusionsA standardized prompt-based framework achieves clinical-grade accuracy across five LLMs from four independent providers, demonstrating model-agnostic generalizability. These findings support the feasibility of vendor-independent clinical AI systems that can leverage multiple models without requiring framework revalidation.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 94%
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 93%
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 93%
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 95%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 94%
- Causal modeling of chronic kidney disease in a participatory framework for informing the inclusion of social drivers in health algorithms 93%
Similar papers in this journal
- Modeling physician variability to prioritize relevant medical record information 93%
- A deep learning model for clinical outcome prediction using longitudinal inpatient electronic health records 92%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 92%
Similar papers in this journal
Similar papers in this journal
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 91%
- Development and Validation of a Machine Learning Model Integrated with the Clinical Workflow for Inpatient Discharge Date Prediction 91%
- Using Machine Learning to Predict Mortality for COVID-19 Patients on Day Zero in the ICU 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.