Predicting Decision-Making Time for Diagnosis over NGS Cycles: An Interpretable Machine Learning Approach
Khodabakhsh, A.; Loka, T. P.; Boutin, S.; Nurjadi, D.; Renard, B. Y.
Show abstract
MotivationGenome sequencing processes are commonly followed by computational analysis in medical diagnosis. The analyses are generally performed once the sequencing process has finished. However, in time-critical applications, it is crucial to start diagnosis once sufficient evidence has been accumulated. This research aims to define a proof-of-principle for predicting earlier time for decision-making using a machine learning approach. The method is evaluated on Illumina sequencing cycles for pathogen diagnosis. ResultsWe utilized a Long-Short Term Memory (LSTM) approach to make predictions for the early decision-making time in time-critical clinical applications. We modeled the (meta-)information obtained from NGS intermediate cycles to investigate whether there are any changes to expect in the remaining sequencing cycles. We tested our model on different patient datasets, resulting in high accuracy of over 98%, indicating the model is independent of a dataset. Furthermore, we can save several hours of turnaround time by using the early prediction results. We used the SHapley Additive exPlanations (SHAP) framework for the interpretation and assessment of the LSTM classifier. AvailabilityThe source code is available at https://gitlab.com/dacs-hpi/ngs-biclass. ContactBernhard.Renard@hpi.de
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- CRISPR-DIPOFF: An Interpretable Deep LearningApproach for CRISPR Cas-9 Off-Target Prediction 96%
- A Computational Toolset for Rapid Identification of SARS-CoV-2, other Viruses, and Microorganisms from Sequencing Data 95%
- Feature selection with vector-symbolic architectures: a case study on microbial profiles of shotgun metagenomic samples of colorectal cancer 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.