PLM-OMG: Protein Language Model-Based Ortholog Detection for Cross-Species Cell Type Mapping
Chau, T. N.; Li, S.
Show abstract
Understanding conserved and divergent cell types across plant species is essential for ad- vancing comparative genomics and improving crop traits. Accurate and scalable ortholog detection is central to this goal, particularly in cross-species single-cell analysis. However, conventional methods are time-consuming and perform poorly with distantly related species, limiting their effectiveness. To address these limitations, we introduce PLM-OMG, a protein language model-based framework for orthogroup classification and cross-species cell type mapping. We benchmark five deep learning models including ESM2, ProGen2, ProteinBERT, ProtGPT2, and LSTM, using a curated 15-species dataset and large-scale monocot and dicot datasets from PLAZA. Transformer-based models, particularly ProtGPT2 and ESM2, achieve superior accuracy and generalization across evolutionary distances. Our results show that PLM-OMG enables scalable and reusable orthogroup detection without recomputing existing groups, significantly reducing computational overhead and highlighting its potential to transform cross-species transcriptomic analysis in plant genomics.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 95%
- ICoN: Integration using Co-attention across Biological Networks 93%
- KSMoFinder - Knowledge graph embedding of proteins and motifs for predicting kinases of human phosphosites 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.