Benchmarking gene embeddings from sequence, expression, network, and text models for functional prediction tasks
Zhong, J.; Li, L.; Dannenfelser, R.; Yao, V.
Show abstract
Accurate, data-driven representations of genes are critical for interpreting high-throughput biological data, yet no consensus exists on the most effective embedding strategy for common functional prediction tasks. Here, we present a systematic comparison of 38 gene embedding methods derived from amino acid sequences, gene expression profiles, protein-protein interaction networks, and biomedical literature. We benchmark each approach across three classes of tasks: predicting individual gene attributes, characterizing paired gene interactions, and assessing gene set relationships while trying to control for data leakage. Overall, we find that literature-based embeddings deliver superior performance across prediction tasks, sequence-based models excel in genetic interaction predictions, and expression-derived representations are well-suited for disease-related associations. Interestingly, network embeddings achieve similar performance to literature-based embeddings on most tasks despite using significantly smaller training sets. The type of training data has a greater influence on performance than the specific embedding construction method, with embedding dimensionality having only minimal impact. Our benchmarks clarify the strengths and limitations of current gene embeddings, providing practical guidance for selecting representations for downstream biological applications.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Integrating temporal single-cell gene expression modalities for trajectory inference and disease prediction 94%
- Knowledge-primed neural networks enable biologically interpretable deep learning on single-cell sequencing data 94%
- Cross-species imputation and comparison of single-cell transcriptomic profiles 94%
Similar papers in this journal
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 97%
- An Analysis of Protein Language Model Embeddings for Fold Prediction 95%
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.