Detecting Self-Repairs from Spontaneous Speech with Prompt Ablation Across LLMs and Fine-Tuned Encoder
Wu, R.; Pugh, S.; OCOnnor, K. B.; Xie, K.; O'Brien, K.; Johnson, K.
Show abstract
Self-repairs, in-utterance revisions in which a speaker abandons and reformulates their speech, are a promising interpretable marker for speech-based cognitive screening. Detecting them automatically is difficult because a self-repair is defined by its relationship to surrounding speech rather than by fixed lexical cues. On the DementiaBank ADReSS corpus, we compared the capability of generative LLMs under a five-condition prompt ablation against a fine-tuned DistilBERT token classifier at detecting self-repairs. GPT-5 performed best (test F1 = 0.73) and was largely insensitive to prompt design, whereas the LLaMA (open-weight alternative) was both weaker and far more prompt-sensitive (test F1 = 0.47). DistilBERT, nearly 100 times smaller, matched the open-weight LLM at a fraction of the computational cost. These results suggest that a locally deployable encoder, given sufficient in-domain annotation, is a more plausible route to clinical self-repair detection than scaling model size or prompt complexity.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enhancing Early Detection of Cognitive Decline in the Elderly through Ensemble of NLP Techniques: A Comparative Study Utilizing Large Language Models in Clinical Notes 92%
- Consistent Performance of GPT-4o in Rare Disease Diagnosis Across Nine Languages and 4967 Cases 90%
- irAE-GPT: Leveraging large language models to identify immune-related adverse events in electronic health records and clinical trial datasets 89%
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 92%
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 91%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 91%
Similar papers in this journal
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 91%
- Lesion site and therapy time predict responses to a therapy for anomia after stroke: a prognostic model development study 91%
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 90%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.