Same Inputs, Different EDSS: Measuring Specification Drift in Clinical Scoring Pipelines
Hwang, S.; Mowery, D. L.; Thomas, S.; Williams, H.; Bar-Or, A.; Sharma, V.; Buijs, F.; Perrone, C.
Show abstract
Clinical informatics pipelines increasingly compute validated clinical endpoints from upstream NLP outputs. Even when the endpoint is defined by an established rubric, translating that rubric across representations - natural language instructions, program logic, and reference implementations - can introduce specification drift, where ostensibly equivalent calculators yield meaningfully different scores. We study this phenomenon for the Expanded Disability Status Scale (EDSS), a standard measure of disability in multiple sclerosis. Holding constant a shared set of functional system (FS) subscores extracted by a large language model (LLM), we compare EDSS values computed across three representations of the same scoring rubric: prompt-executed natural language, LLM-generated code, and a canonical reference implementation. We characterize disagreement structure, distributional shifts, and clinically salient boundary flips, and we propose an audit workflow that treats endpoint computation as a first-class verification target in clinical NLP systems.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 94%
- Natural language inference for clinical registry curation 93%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 95%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 93%
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 94%
- MMFP-Tableau: Enabling Precision Mitochondrial Medicine through Integration, Visualization, and Analytics of Clinical and Research Health System Electronic Data 93%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.