Harnessing the 100,000 Genomes Project whole genome sequencing data - an unbiased systematic tool to filter by biologically validated regions of functionality
Xiao, S.; Kai, Z.; Brown, D.; Genomics England Research Consortium, ; Shovlin, C. L.
Show abstract
Whole genome sequencing (WGS) is championed by the UK National Health Service (NHS) to identify genetic variants that cause particular diseases. The full potential of WGS has yet to be realised as early data analytic steps prioritise protein-coding genes, and effectively ignore the less well annotated non-coding genome which is rich in transcribed and critical regulatory regions. To address, we developed a filter, which we call GROFFFY, and validated in WGS data from hereditary haemorrhagic telangiectasia patients within the 100,000 Genomes Project. Before filter application, the mean number of DNA variants compared to human reference sequence GRCh38 was 4,867,167 (range 4,786,039-5,070,340), and one-third lay within intergenic areas. GROFFFY removed a mean of 2,812,015 variants per DNA. In combination with allele frequency and other filters, GROFFFY enabled a 99.56% reduction in variant number. The proportion of intergenic variants was maintained, and no pathogenic variants in disease genes were lost. We conclude that the filter applied to NHS diagnostic samples in the 100,000 Genomes pipeline offers an efficient method to prioritise intergenic, intronic and coding gDNA variants. Reducing the overwhelming number of variants while retaining functional genome variation of importance to patients, enhances the near-term value of WGS in clinical diagnostics.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Matching whole genomes to rare genetic disorders: Identification of potential causative variants using phenotype-weighted knowledge in the CAGI SickKids5 clinical genomes challenge 93%
- Phasing of de novo mutations using a scaled-up multiple amplicon long-read sequencing approach 93%
- GeneBreaker: Variant simulation to improve the diagnosis of Mendelian rare genetic diseases 93%
Similar papers in this journal
- Towards robust clinical genome interpretation: developing a consistent terminology to characterize disease-gene relationships - allelic requirement, inheritance modes and disease mechanisms 94%
- A gene pathogenicity tool 'GenePy' identifies missed biallelic diagnoses in the 100,000 Genomes Project 94%
- Assessment of the variant prioritisation strategy for genomic newborn screening in the Generation Study 93%
Similar papers in this journal
- GA4GH Phenopacket-Driven Characterization of Genotype-Phenotype Correlations in Mendelian Disorders 93%
- Identification of actionable genetic variants in 4,198 Scottish volunteers from the Viking Genes research cohort and implementation of return of results 93%
- Validating data from Multiplex Assays of Variant Effect (MAVEs): A CanVIG-UK National Survey of NHS Clinical Scientists 92%
Similar papers in this journal
Similar papers in this journal
- Recommendations for clinical interpretation of variants found in non-coding regions of the genome 92%
- Genome-Wide Sequencing as a First-Tier Screening Test for Short Tandem Repeat Expansions 92%
- Evaluating Genome Sequencing Strategies: Trio, Singleton, and Standard Testing in Rare Disease Diagnosis 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.