Back

StarPhase: Comprehensive Phase-Aware Pharmacogenomic Diplotyper for Long-Read Sequencing Data

Holt, J. M.; Harting, J.; Chen, X.; Baker, D.; Saunders, C. T.; Kronenberg, Z.; Gonzaludo, N.; Yoo, B.; Hudjashov, G.; Joeloo, M.; Lawlor, J. M. J.; Lim, W. K.; Estonian Biobank Research Team, ; Jamuar, S. S.; Cooper, G. M.; Milani, L.; Pastinen, T.; Eberle, M. A.

2024-12-11 bioinformatics
10.1101/2024.12.10.627527 bioRxiv
Show abstract

Pharmacogenomics is central to precision medicine, informing medication safety and efficacy. Phar-macogenomic diplotyping of complex genes requires full-length DNA sequences and detection of structural rearrangements. We introduce StarPhase, a tool that leverages PacBio HiFi sequence data to diplotype 21 CPIC Level A pharmacogenes and provides detailed haplotypes and supporting visualizations for HLA-A, HLA-B, and CYP2D6. StarPhase diplotypes have high concordance with benchmarks where 99.5% are either exact matches or minor discrepancies. Manual inspection of the 0.5% mismatches indicates they were correctly called by StarPhase. With StarPhase, we update or correct 26.2% of GeT-RM pharmacogenomic diplotypes. Population distributions from StarPhase mostly reflect those of the All of Us cohort, while also highlighting gaps in existing pharmacogenomic databases that long-read sequencing can fill. With a single HiFi whole genome sequencing assay, StarPhase enables robust PGx diplotyping even as additional pharmacogenes and haplotypes are discovered.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.