Back

Contrastive modelling of transcription and transcript abundance in legumes using PlanTT

Raymond, N.; Zhang, X.; Sheikh, J.; Daveouis, F.; Verma, R.; Cram, D.; Song, H.; Cao, Y.; Kirzinger, M.; Akaniru, D.; Ubbens, J.; Konkin, D.

2025-11-16 bioinformatics
10.1101/2025.11.15.685414 bioRxiv
Show abstract

Predicting the impacts of sequence variation on gene expression remains a challenging task. Further, in plants, we have a limited understanding of the relative contributions of different gene expression regulatory mechanisms. To address these limitations we generated a comparative multiomic dataset comprising matched 3-RNA-seq and PRO-seq data from matched tissues of reference genotypes of four legumes of the invert repeat lacking clade (Pisum sativum, Vicia faba, Lathyrus sativa and Medicago truncatula). Focused on the challenging task of predicting expression differences between ortholog pairs from unseen orthogroups, we used this dataset and a novel prediction framework to build contrastive models that predict quantitative differences (effect size differences) in transcription and transcript abundance.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.