Back

A novel machine learning-based algorithm for eQTL identification reveals complex pleiotropic effects in the MHC region

Li, R. Y.; Su, C.; Qin, Z. S.

2025-05-08 bioinformatics
10.1101/2025.05.06.652558 bioRxiv
Show abstract

Expression quantitative trait loci (eQTLs) are regulatory variants that affect the expression level of their target genes and have significant impact on disease biology. However, eQTL mapping has been done mostly in one tissue at a time, despite the known prevalence of correlations among tissues. Multivariate analyses incorporating multiple phenotypes are available, but they emphasize linear combinations of phenotypes. We present MTClass, a machine learning framework that attempts to classify an individuals genotype based on a vector of multi-phenotype expression levels of a given gene. We conduct simulation studies and multiple case studies using real and imputed data, and we demonstrate that MTClass detects more functionally relevant variants and genes compared to existing single-tissue approaches as well as multi-phenotype association tests. Our results suggest that the importance of expression regulation at the MHC region may have been underestimated, and they provide fresh biological insights into genetic variants that have pleiotropic effects, influencing gene expression in a complex manner. Key pointsO_LIMTClass is a machine learning-based approach that classifies genotypes based on multi-phenotype expression data, providing a novel method for identifying eQTLs. C_LIO_LIMTClass outperforms traditional linear methods like MultiPhen and MANOVA in detecting eQTLs with greater functional impact and in capturing complex genotype-phenotype relationships. C_LIO_LIMTClass identified immune-related variants in the HLA region, suggesting that existing approaches may have underestimated the complexity of these variants effects across tissues. C_LIO_LIMTClass is more flexible and reliable than linear multivariate methods, handling multicollinearity, zero-expressed features, and various input values with greater ease. C_LI

Published in Briefings in Bioinformatics (predicted rank #9) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.