A Supervised Text-Embedded Transformer Matching Model to Detect Fall Injuries in Medicare Data
Kane, M.; Greene, E. J.; Esserman, D.; Latham, N. K.; Min, L. C.; Ganz, D. A.
Show abstract
Objective: To develop and validate a supervised text-embedded transformer matching model to identify fall injuries in Medicare data, and evaluate the model's performance -- alongside a validated rule-based algorithm-- against "ground truth" from an external reference standard (self-reported fall injuries leading to medical attention). Materials and Methods: Text embeddings of ICD-10-CM and CPT codes in Medicare claims/encounters from participants in the Strategies to Reduce Injuries and Develop Confidence in Elders (STRIDE) trial served as model inputs. Trained on annotated claims/encounters occurring within +/- one month of self-reported fall injuries leading to medical attention, the transformer model generated a continuous 0-1 probability that each claim/encounter was for a fall injury. The model was then applied to all claims/encounters in STRIDE and compared alongside the rule-based algorithm to the external reference standard. Results: The model achieved an area under the curve (AUC) of > 0.96 against annotated claims/encounters in 9 out of 10 holdout folds and 0.85 in the remaining fold. In the full STRIDE dataset, the model achieved a peak AUC of 0.86 (95% CI, 0.84-0.87) against the external reference standard, with results comparable to the rule-based algorithm. Discussion: Relative to rule-based approaches, which typically generate binary outcomes, the continuous event probability generated by the transformer model could support clinical endpoint adjudication, with high-probability predictions treated as events, moderate-probability predictions being adjudicated, and low-probability predictions treated as non-events. Conclusion: A text-embedded transformer model identified fall injuries with comparable accuracy to a rule-based algorithm, demonstrating "proof of concept" for use in endpoint adjudication.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Preventable deaths involving falls in England and Wales, 2013-2022: a systematic case series of coroners’ reports 91%
- Evaluation of Real-World Mobility Recovery after Hip Fracture using Digital Mobility Outcomes 91%
- Home-based Extended Rehabilitation for Older People with Frailty (HERO): a Randomised Controlled Trial 90%
Similar papers in this journal
- Linking Patient Records at Scale with a Hybrid Approach Combining Contrastive Learning and Deterministic Rules 90%
- A Novel Method for Handling Pre-Existing Conditions in Prediction Models for Covid-19 Death 88%
- A Systematic Review of the Application of Computational Grounded Theory Method in Healthcare Research 88%
Similar papers in this journal
- Applying time series analyses on continuous accelerometry data: a clinical example in older adults with and without cognitive impairment 93%
- A machine learning approach to identifying important features for achieving step thresholds in individuals with chronic stroke 93%
- A Low-Cost Markerless motion capture system to automate Functional Gait Assessment: Feasibility Study 92%
Similar papers in this journal
- Optimizing Temporal Windows for Wearable-Augmented Post-Discharge Risk Prediction: A Methods Study 94%
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 93%
- Temporally-Informed Random Forests for Suicide Risk Prediction 92%
Similar papers in this journal
- Towards Clinical Prediction with Transparency: An Explainable AI Approach to Survival Modelling in Residential Aged Care 94%
- A standardized analytics pipeline for reliable and rapid development and validation of prediction models using observational health data 91%
- SPELL-LLMs: A Scalable and Privacy-Compliant NLP Pipeline Using Locally Hosted Large Language Models for Clinical Information Extraction 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.