Back

Pangenomes Aid Accurate Detection Of Large Insertion And Deletions From Gene Panel Data: The Case Of Cardiomyopathies

Mazzarotto, F.; Kalay, O.; Arslan, E.; Cinquina, V.; Turgut, D.; Buchan, R.; Allouba, M.; Bertini, V.; Halawa, S.; Theotokis, P.; Budak, G.; Girolami, F.; Peldova, P.; Bonaventura, J.; Aguib, Y.; Colombi, M.; Olivotto, I.; Gennarelli, M.; Macek, M.; Pelo, E.; Ritelli, M.; Yacoub, M.; Barton, P.; Tetikol, S.; Walsh, R.; Ware, J.; Jain, A.

2024-12-02 health informatics
10.1101/2024.11.27.24318059 medRxiv
Show abstract

Gene panels represent a widely used strategy for genetic testing in a vast range of Mendelian disorders. While this approach aids reliable bioinformatic detection of short coding variants, it fails to detect most larger variants. Recent studies have recommended the adoption of pangenomes to augment detection of large variants from targeted sequencing, potentially providing diagnostic laboratories with the possibility to streamline diagnostic work-ups and reduce costs. Here, we analyze a large-scale cohort comprising 1,952 cardiomyopathy cases and 1,805 technically matched controls and show that a pangenome-based workflow, GRAF, conjugates higher precision and recall (F1 score 0.86) compared with conventional orthogonal methods (F1 0-0.57) in detecting potentially pathogenic [≥]20bp variants from short-read panel data. Our results indicate that pangenome-based workflows aid precise and cost-effective detection of large variants from targeted sequencing data in the clinical context. This will be particularly relevant for conditions in which these variants explain a high proportion of the disease burden.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.