PPARgene 2.0: leveraging large language models and multi-omics data for enhanced identification and prediction of PPAR target genes
Qin, J.; Xie, W.; Li, H.; Li, H.; Zeng, Y.; Luo, Y.; Li, Y.; Sha, T.; Wang, N.; Li, Y.; Fang, L.
Show abstract
Peroxisome proliferator-activated receptors (PPARs) are ligand-activated transcription factors of the nuclear receptor superfamily. Upon ligand binding, PPARs activate target gene transcription and regulate a variety of important physiological processes such as lipid metabolism, inflammation, wound healing and immune responses. PPARgene is a database that integrates literature-curated and computationally predicted PPAR target genes. It provides gene-level annotations including tissue specificity, species, and supporting PubMed IDs. Computational predictions are generated using a machine learning method that combines PPRE motif analysis with microarray expression data. Here, we introduce PPARgene 2.0, a 10-year update to the original PPARgene database. This update adds 35 newly reported PPAR target genes, 20 PPAR{beta}/{delta} target genes, and 72 PPAR{gamma} target genes, bringing the total number of curated target genes in the database to 337. To retrieve newly reported PPAR target genes from the literature, we used two language models to screen publications after 2016. Candidate papers were then manually reviewed, and verified target genes were added to the updated database. This update also improved the predictive method by expanding the volume of high-throughput gene expression data and incorporating PPAR-related ChIP-seq datasets alongside in silico PPRE analysis. Fivefold cross-validation demonstrated that the new predictive method outperforms the original one. The updated prediction tool is available as part of the PPARgene 2.0 platform. The database is openly accessible at https://www.ppargene.org.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ChemPert: mapping between chemical perturbation and transcriptional response for non-cancer cells 93%
- PerturbAtlas: A Comprehensive Atlas of Public Genetic Perturbation Bulk RNA-seq Datasets 92%
- Expanding the coverage of regulons from high-confidence prior knowledge for accurate estimation of transcription factor activities 92%
Similar papers in this journal
- Genomic background sequences systematically outperform synthetic ones in de novo motif discovery for ChIP-seq data 92%
- Covering all your bases: incorporating intron signal from RNA-seq data 91%
- Differential Expression Enrichment Tool (DEET): an interactive atlas of human differential gene expression 91%
Similar papers in this journal
Similar papers in this journal
- AnnoMiner: a new web-tool to integrate epigenetics, transcription factor occupancy, and transcriptomics data to predict transcriptional regulators 93%
- A systematic study of HIF1A cofactors in hypoxic cancer cells 92%
- Network and pathway expansion of genetic disease associations identifies successful drug targets 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.