The unreachable genomic profiling of complex diseases: genotype missingness matters
Abad-Grau, M. M.
Show abstract
The problem of building genome-wide predictors of individual risk to complex diseases seems to be more challenging than it was thought when the first human genome was sequenced on 2003. We have build different enhanced genetic risk predictors from genome-wide data and different complex diseases, making use of haplotypes accurately ascertained from family trios. We confirmed the widely known inability to accurately predict individual risk to complex diseases returned by the state-of-the-art genome-wide predictors. This result is mainly due to the small effect that most genetic variants have in a disease. We also found out that rates of missing genotypes were usually too high for these small-effect variants, as we could force missing imputation in such a tricky way that we would build highly accurate predictors, either by using our own design or the state-of-the-art genetic predictors. We observed that unknown genotypes were not missing at random but related to disease affectation, with more missing genotypes in affected than in non-affected individuals. We were not able to find a way to accurately reduce missing rates to correctly improve accuracy, but we identified a common pattern of missing data in multiple sclerosis, asthma and autism that makes us think that other complex diseases could behave the same way. Because (1) missing rates are high enough to completely change risk prediction due to the small-effect of most of the variants, and (2) there are more missing genotypes in affected than in non affected individuals, we conclude that perhaps the widely known defeat in genomic profiling for complex diseases may be solved by looking closer to the way current genotyping technologies handle genetic variants that may be rare in reference panels but have some effect in a given complex disease.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Towards development of a statistical framework to evaluate myotonic dystrophy type 1 mRNA biomarkers in the context of a clinical trial 96%
- Assessing the performance of genome-wide association studies for predicting disease risk 96%
- HCLC-FC: a novel statistical method for phenome-wide association studies 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Finding disease modules for cancer and COVID-19 in gene co-expression networks with the Core&Peel method 94%
- The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients 94%
- Biobank-scale methods and projections for sparse polygenic prediction from machine learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.