Back

MLSPred-Bench: ML-Ready Benchmark LeveragingSeizure Detection EEG data for Predictive Models

Mohammad, U.; Saeed, F.

2024-07-19 bioinformatics
10.1101/2024.07.17.604006 bioRxiv
Show abstract

Predicting epileptic seizures is a significantly challenging task as compared to detection. While electroen-cephalography (EEG) data annotated for detection is available from multiple repositories, they cannot readily be used for predictive modeling. In this paper, we designed and developed a strategy that can be used for converting any EEG big data annotated for detection into ML-ready data suitable for prediction. The generalizability of our strategy is demonstrated by executing it on Temple University Seizure (TUSZ) corpus which is annotated for seizure detection. This execution results in 12 ML-ready datasets, collectively called MLSPred-Bench benchmark, which constitutes data for training, validating and testing seizure prediction models. Our strategy uses different variations of seizure prediction horizon (SPH) and the seizure occurrence period (SOP) to make more than 150GB of ML-ready data. To illustrate that the generated data can be used for predictive modeling, we executed an ML model on all the benchmarks which resulted in variable performances when compared with the original model and its performance. We expect that our strategy can be used as a general method to transform seizure detection EEG big data into ML-ready datasets useful for seizure prediction. Our code and related materials will be made available at https://github.com/pcdslab/MLSPred-Bench.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.