A systematic comparison of computational methods for expression forecasting
Kernfeld, E. M.; Yang, Y.; Weinstock, J. S.; Battle, A.; Cahan, P. M.
Show abstract
Expression forecasting methods use machine learning models to predict how a cell will alter its transcriptome upon perturbation. Such methods are enticing because they promise to answer pressing questions in fields ranging from developmental genetics to cell fate engineering and because they are a fast, cheap, and accessible complement to the corresponding experiments. However, the absolute and relative accuracy of these methods is poorly characterized, limiting their informed use, their improvement, and the interpretation of their predictions. To address these issues, we created a benchmarking platform that combines a panel of 11 large-scale perturbation datasets with an expression forecasting software engine that encompasses or interfaces to a wide variety of methods. We used our platform to systematically assess methods, parameters, and sources of auxiliary data, finding that performance strongly depends on the choice of metric, and especially for simple metrics like mean squared error, it is uncommon for expression forecasting methods to out-perform simple baselines. Our platform will serve as a resource to improve methods and to identify contexts in which expression forecasting can succeed.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- High performance single-cell gene regulatory network inference at scale: The Inferelator 3.0 97%
- Non-negative Independent Factor Analysis disentangles discrete and continuous sources of variation in scRNA-seq data 96%
- Airpart: Interpretable statistical models for analyzing allelic imbalance in single-cell datasets 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Robust differential expression testing for single-cell CRISPR screens at low multiplicity of infection 95%
- Visualizing scRNA-Seq Data at Population Scale with GloScope 95%
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.