Language Model Applications for Early Diagnosis of Childhood Epilepsy
Loyens, J.; Slinger, T.; Doornebal, N.; Braun, K.; Otte, W. M.; van Diessen, E.
Show abstract
ObjectiveAccurate and timely epilepsy diagnosis is crucial to reduce delayed or unnecessary treatment. While language serves as an indispensable source of information for diagnosing epilepsy, its computational analysis remains relatively unexplored. This study assessed - and compared - the diagnostic value of different language model applications in extracting information and identifying overlooked language patterns from first-visit documentation to improve the early diagnosis of childhood epilepsy. MethodsWe analyzed 1,561 patient letters from two independent first seizure clinics. The dataset was divided into training and test sets to evaluate performance and generalizability. We employed two approaches: an established Naive Bayes model as a natural language processing technique, and a sentence-embedding model based on the Bidirectional Encoder Representations from Transformers (BERT)-architecture. Both models analyzed anamnesis data only. Within the training sets we identified predictive features, consisting of keywords indicative of epilepsy or no epilepsy. Model outputs were compared to the clinicians final diagnosis (gold standard) after follow-up. We computed accuracy, sensitivity, and specificity for both models. ResultsThe Naive Bayes model achieved an accuracy of 0.73 (95% CI: 0.68-0.78), with a sensitivity of 0.79 (95% CI: 0.74-0.85) and a specificity of 0.62 (95% CI: 0.52-0.72). The sentence-embedding model demonstrated comparable performance with an accuracy of 0.74 (95% CI: 0.68-0.79), sensitivity of 0.74 (95% CI: 0.68-0.80), and specificity of 0.73 (95% CI: 0.61-0.84). ConclusionBoth models demonstrated relatively good performance in diagnosing childhood epilepsy solely based on first-visit patient anamnesis text. Notably, the more advanced sentence-embedding model showed no significant improvement over the computationally simpler Naive Bayes model. This suggests that modeling of anamnesis data does depend on word order for this particular classification task. Further refinement and exploration of language models and computational linguistic approaches are necessary to enhance diagnostic accuracy in clinical practice.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Real-world smartphone data can trace the behavioural impact of epilepsy: A Case study 92%
- The effectiveness of Cenobamate in patients previously treated with Vagus Nerve Stimulation for drug resistant epilepsy 91%
- Thalamohippocampal atrophy in focal epilepsy of unknown cause at the time of diagnosis 90%
Similar papers in this journal
- Evaluating the generalisability of region-naïve machine learning algorithms for the identification of epilepsy in low-resource settings 95%
- Natural language processing to evaluate texting conversations between patients and healthcare providers during COVID-19 Home-Based Care in Rwanda at scale 91%
- Inferring Gender from First Names: Comparing the Accuracy of Genderize, Gender API, and the gender R Package on Authors of Diverse Nationality 90%
Similar papers in this journal
Similar papers in this journal
- An international study to investigate and optimise the safety of discontinuing valproate in young men and women with epilepsy: protocol 93%
- The use of carbogen for interruption of febrile seizures - the randomized controlled CARDIF trial 91%
- The impact of paediatric epilepsy and co-occurring neurodevelopmental disorders on functional brain networks in wake and sleep 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.