Back

PPARgene 2.0: leveraging large language models and multi-omics data for enhanced identification and prediction of PPAR target genes

Qin, J.; Xie, W.; Li, H.; Li, H.; Zeng, Y.; Luo, Y.; Li, Y.; Sha, T.; Wang, N.; Li, Y.; Fang, L.

2025-12-04 bioinformatics
10.64898/2025.12.01.691485 bioRxiv
Show abstract

Peroxisome proliferator-activated receptors (PPARs) are ligand-activated transcription factors of the nuclear receptor superfamily. Upon ligand binding, PPARs activate target gene transcription and regulate a variety of important physiological processes such as lipid metabolism, inflammation, wound healing and immune responses. PPARgene is a database that integrates literature-curated and computationally predicted PPAR target genes. It provides gene-level annotations including tissue specificity, species, and supporting PubMed IDs. Computational predictions are generated using a machine learning method that combines PPRE motif analysis with microarray expression data. Here, we introduce PPARgene 2.0, a 10-year update to the original PPARgene database. This update adds 35 newly reported PPAR target genes, 20 PPAR{beta}/{delta} target genes, and 72 PPAR{gamma} target genes, bringing the total number of curated target genes in the database to 337. To retrieve newly reported PPAR target genes from the literature, we used two language models to screen publications after 2016. Candidate papers were then manually reviewed, and verified target genes were added to the updated database. This update also improved the predictive method by expanding the volume of high-throughput gene expression data and incorporating PPAR-related ChIP-seq datasets alongside in silico PPRE analysis. Fivefold cross-validation demonstrated that the new predictive method outperforms the original one. The updated prediction tool is available as part of the PPARgene 2.0 platform. The database is openly accessible at https://www.ppargene.org.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.