Automated Seizure Classification Using Multimodal Large Language Models
Zhang, L.; Jiang, R.; Monsoor, T.; Pasqua, J. N.; McCrimmon, C. M.; Sinha, P.; Sharma, K.; Alzuabi, M.; Morales, V.; Miranda, H. M.; Manjeshwar, C.; Roychowdhury, V.; Mazumder, R.
Show abstract
ObjectiveAccurately distinguishing between epileptic seizures (ES) and nonepileptic seizures (NES) is a significant clinical challenge that typically requires resource-intensive inpatient video-EEG monitoring. Here, we developed a novel Multimodal Large Language Models (MLLMs)-based method for automated extraction of semiological features from videos of seizure events, and subsequently, classified the events as ES or NES. Methods90 videos of ES and NES events from 29 patients were obtained from an epilepsy monitoring unit at a large academic hospital. Events were labeled as ES or NES based on expert evaluation of video-EEG recordings and simultaneously annotated with 24 clinically relevant semiological features. We implemented a MLLMs framework that integrates open-source vision-language models (VLMs) and audio-language models (ALMs) to analyze the videos and associated audio tracks and automatically extract these 24 features. The performance of the MLLMs-based feature extraction was evaluated against expert annotations. These features were subsequently used to train several classifiers including K-Nearest Neighbors (KNN), XGBoost, and Deep Factorization Machine, to differentiate ES from NES. Model performance was evaluated using leave-one-patient-out (LOPO) cross-validation. ResultsUsing KNN, expert-annotated semiological features achieved precision 0.97, recall 0.97, F1-score 0.97, and AUC 0.99, establishing an upper bound on ES/NES classification performance. The MLLMs pipeline achieved an overall mean recall of 0.71, mean accuracy of 0.58, and a mean F1-score of 0.51 for semiological feature extraction compared to expert annotations. The best performing KNN model (k=7) using MLLMs-extracted features achieved a precision of 0.77, recall of 0.76, F1-score of 0.76, and AUC of 0.76 in classifying ES versus NES; correctly identifying 68 out of 90 events. ConclusionWe demonstrate the feasibility of using MLLMs to automatically extract clinically relevant semiological features from seizure videos and classify ES versus NES. MLLMs-based feature extraction and classification offer a promising clinically interpretable approach to aid diagnosis of epilepsy using videos.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- NLP-based tools for localization of the Epileptogenic Zone in patients with drug-resistant focal epilepsy 96%
- Dynamic network properties of the interictal brain determine whether seizures appear focal or generalised 93%
- Event Driven Neural Network on a Mixed Signal Neuromorphic Processor for EEG Based Epileptic Seizure Detection 93%
Similar papers in this journal
Similar papers in this journal
- Multiscale predictive modeling robustly improves the accuracy of pseudo-prospective seizure forecasting in drug-resistant epilepsy 95%
- Characterizing physiological high-frequency oscillations using deep learning 95%
- Interpretable EEG Biomarkers for Neurological Disease Models in Mice Using Bag-of-Waves Classifiers 94%
Similar papers in this journal
- Longitudinally Tracking Personal Physiomes for Precision Management of Childhood Epilepsy 94%
- Evaluating the generalisability of region-naïve machine learning algorithms for the identification of epilepsy in low-resource settings 93%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.