Genome Medicine
○ Springer Science and Business Media LLC
All preprints, ranked by how well they match Genome Medicine's content profile, based on 183 papers previously published here. The average preprint has a 0.18% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Hillary, R. F.; McCartney, D. L.; Bernabeu, E.; Gadd, D. A.; Cheng, Y.; Chybowska, A. D.; Smith, H. M.; Murphy, L.; Wrobel, N.; Campbell, A.; Walker, R. M.; Hayward, C.; Evans, K. L.; McIntosh, A. M.; Marioni, R. E.
Show abstract
BackgroundBlood DNA methylation can inform us about the biological mechanisms that underlie common disease states. Previous epigenome-wide analyses of common diseases often focus solely on the prevalence or incidence of individual conditions and rely on small sample sizes, which may limit power to discover disease-associated loci. ResultsWe conduct blood-based epigenome-wide association studies on the prevalence of 14 common disease states in Generation Scotland (nindividuals[≤]18,413, nCpGs=752,722). We also utilise health record linkage to perform epigenome-wide analyses on the incidence of 19 disease states. We present a structured literature review on existing epigenome-wide analyses for all 19 disease states to assess the degree of replication within the existing literature and the novelty of the present findings. We identify 69 associations between CpGs and the prevalence of four disease states at baseline, of which 58 are novel. We also uncover 64 CpGs that associate with the incidence of two disease states (COPD and type 2 diabetes), of which 56 are novel. These associations were independent from common lifestyle risk factors. We highlight poor replication across the existing literature. Here, replication was defined by the reporting of at least one common gene in >2 studies examining the same disease state. Existing blood-based epigenome-wide analyses showed evidence of replication for only 4/19 disease states (with up-to-15% of unique genes replicated for lung cancer). ConclusionsOur summary data and structured review of the literature provide an important platform to guide future studies that examine the role of blood DNA methylation in complex disease states.
Dann, E.; Teeple, E.; Elmentaite, R.; Meyer, K. B.; Gaglia, G.; Nestle, F.; Savova, V.; de Rinaldis, E.; Teichmann, S.
Show abstract
Whilst the use of single-cell RNA sequencing (scRNA-seq) to understand target biology is well established, its predictive role in increasing the clinical success of therapeutic targets remains underexplored. Inspired by previous work on an association between genetic evidence and clinical success, we used retrospective analysis of known drug target genes to identify potential predictors of target clinical success from scRNA-seq data. We investigated whether successful drug targets are associated with cell type specific expression in a disease-relevant tissue (cell type specificity), and with cell type specific over-expression in disease patients compared to healthy controls (disease cell specificity). Analysing scRNA-seq data across diseases and tissues, we found that both cell type and disease cell specificity are features enriched in targets entering clinical development, and that cell type specificity in the disease-relevant tissue is robustly predictive of target progression from Phase I to II. While scRNA-seq analysis identifies a larger and complementary target space to that of direct genetic evidence, its association with specificity and drug approval appears less clear. We discuss how further expansion and harmonization of single-cell datasets, more sophisticated integration of this data in target discovery, and improved methods for tracking clinical trial outcomes could enhance our ability to leverage scRNA-seq insights in drug development in future.
Jin, H.; Andreopoulos, M.; Viswanadham, V. V.; Park, P. J.
Show abstract
Illumina short-read sequencing underpins clinical cancer genomics, with targeted panels widely used to detect actionable variants. Newer Illumina platforms employ a two-color chemistry to accelerate sequencing, but its impact on variant identification has not been systematically evaluated. Here we show that two-color platforms generate recurrent T>G artifacts in targeted panels at low variant allele fractions, predominantly occurring in specific trinucleotide contexts. These artifacts can produce spurious pathogenic variants in key cancer genes such as TP53 and KIT, and inflate tumor mutational burden, a metric considered when assessing patient eligibility for immunotherapy. Accounting for such artifacts is therefore essential for accurate interpretation of clinical panel data.
Ahmed, Z.; Govindareddy, P.; DeGroat, W.; Narayanan, R.; Peker, E.; Zeeshan, S.
Show abstract
Precision medicine aims to advance our ability from a "one-size-fits-all" approach to personalized and predictive healthcare across diverse populations. It promotes integration of multi-omics and phenotypic data to understand disease mechanisms and discover novel biomarkers and risk factors, which could be used to predict and prevent critical diseases in individual patients across diverse populations. The potential implications of precision medicine approach can accelerate our ability to classify patients at higher risk of developing critical diseases, improve diagnostic capabilities, develop deeper understanding of individual risk, investigate racial differences and demographic characteristics, and find relationships between genetic variants, expressions, and diseases. This study focuses on implementing an innovative and data driven framework of translational bioinformatics and Machine Learning (ML) techniques to analyze multi-omics, including RNA-seq and Whole-Genome Sequencing (WGS) data, generated using blood samples of randomly consented patients. First, we utilized bioinformatics pipelines to identify differentially expressed genes and their pathogenic and likely pathogenic variants for the downstream data analysis, annotation, and visualization. Then, applied a nexus of ML models for multi-omics biomarker discovery, disease prediction, density-based clustering, single-patient profiling, and pathogenicity classification. WGS data analysis supported the exploration of genetic variation and diversity among patients to identify known and novel biomarkers, whereas RNA-seq data analysis improved our understanding of functional and biological pathways that underlying disease states. We classified and clustered pathogenic variants and expressions across various genes and discovered numerous diseases leading risk factors. Our results include gene-disease associations and captured common pathways across the broader population, demonstrating a level of sensitivity and accuracy that has broad clinical implications. We validated our results through clinical records, and state of the science literature. This study delves into the strengths of multi-omics data integration and capabilities of ML application in genetically diverse and complex patient cohorts. Our approach has the potential to elucidate complex gene-disease interactions for genetically diverse populations, which can support earlier diagnoses for patients in many disease realms.
Shovlin, C. L.; Vizcaychipi, M. P.
Show abstract
BackgroundMortality remains very high and unpredictable in CoViD-19, with intense public protection strategies tailored to preceived risk. Males are at greater risk of severe CoViD-19 complications. Genomic studies are in process to identify differences in host susceptibility to SARS-CoV-2 infection. MethodsGenomic structures were examined for the ACE2 gene that encodes angiotensin-converting enzyme 2, the obligate receptor for SARS-CoV-2. Variants in 213,158 exomes/genomes were integrated with ACE2 protein functional domains, and pathogenicity criteria from the American Society of Human Genetics and Genomics/Association for Molecular Pathology. Results483 variants were identified in the 19 exons of ACE2 on the X chromosome. All variants were rare, including nine loss-of-function (potentially SARS-CoV-2 protective) alleles present only in female heterozygotes. Unopposed variant alleles were more common in males (262/3596 [7.3%] nucleotides) than females (9/3596 [0.25%] nucleotides, p<0.0001). 37 missense variants substituted amino acids in SARS-CoV-2 interacting regions or critical domains for transmembrane ACE2 expression. Four upstream open reading frames with 31 associated variants were identified. Excepting loss-of-function alleles, variants would not meet minimum criteria for classification as Likely Pathogenic/beneficial if differential frequencies emerged in patients with CoViD-19. ConclusionsMales are more exposed to consequences from a single variant ACE2 allele. Common risk/beneficial alleles are unlikely in regions subject to evolutionary constraint. ACE2 upstream open reading frames may have implications for aminoglycoside use in SARS-CoV-2-infected patients. For this SARS-CoV-2-interacting protein with pre-identified functional domains, pre-emptive functional and computational studies are encouraged to accelerate interpretations of genomic variation for personalised and public health use.
Ganapathy, K. R.; Broly, M.; Silverstein, S.; Mendoza, M.; Song, E.; Kotis, B.; Hoffman, P.; Pediatric Congenital Genomic Consortium, ; Torkamani, A.; Adams, D. R.; Bonnemann, C.; Lappalainen, T.; Mohammadi, P.
Show abstract
Allele-specific expression (ASE) outlier detection is a powerful tool for identifying genes affected by large effect rare genetic regulatory variants but suffers from data sparsity and noisy signal in low-count genes. Genome phasing can be utilized to aggregate ASE signal along haplotypes to alleviate both sparsity and noise. Yet statistical tools for utilizing haplotype-level ASE data for rare variant interpretation are lacking. Here, we present ANEVA-h, to quantify the amount of genetic variation in gene expression from haplotype-level ASE data in a population, enabling more accurate and comprehensive detection of regulatory effects. We apply ANEVA-h to GTEx project data, along with a compatible dosage outlier test, to show an over 2-fold increase in the number of testable genes, reduction of spurious outlier calls, and improved enrichment for rare high-impact variants. In clinical cohorts of neuromuscular and congenital heart disease, it enhances gene prioritization and identifies candidate diagnoses missed by DROP-MAE and ANEVA. Finally, we analyze globally diverse populations to characterize the impact of ancestry background in reference and the test population. We provide tools and data necessary to facilitate integration of haplotype level ASE outlier testing in rare variant interpretation pipelines
Cheng, Y.; Gieger, C.; Campbell, A.; McIntosh, A. M.; Waldenberger, M.; McCartney, D. L.; Marioni, R. E.; Vallejos, C. A.
Show abstract
Over the last decade, a plethora of blood-based DNA methylation biomarkers have been developed to track differences in ageing, lifestyle, health, and biological outcomes. Typically, penalised regression models are used to generate these predictors, with hundreds or thousands of CpGs included as potential features. However, in such ultra high-dimensional settings, the effectiveness of these methods may be reduced. Here, we introduce Related Trait-based Feature Screening (RTFS), a method for performing CpG pre-selection for incident disease prediction models by utilising associations between CpGs and health-related continuous traits. In a comparison with commonly used CpG pre-selection methods, we evaluate resulting downstream Cox proportional-hazards prediction models for 10-year type 2 diabetes (T2D) onset risk in Generation Scotland (n=18,414). The top performing models utilised incident T2D EWAS (AUC=0.881, PRAUC=0.279) and RTFS (AUC=0.877, PRAUC=0.277). The resulting models also improve prediction over a model using standard risk factors only (AUC=0.841, PRAUC=0.194) and replication was observed in the German-based KORA study (n=4,261) RTFS is a flexible and generalisable framework that can help to refine biomarker development for incident disease outcomes.
Margolin, G.; Bang-Christensen, S.; Grimm, S. A.; Bennett, B. D.; Kilfeather, P.; Fedkenheuer, K.; Pathan, S.; Whitlock, N. C.; Pinto, P. A.; Klimczak, L. J.; Farney, S. K.; Jameel, N.; Chen, Y.-C.; Petrykowska, H. M.; Sowalsky, A. G.; Li, J.-L.; Elnitski, L. L.
Show abstract
Examining DNA in a liquid biopsy for non-invasive cancer detection relies on identifying dilute signal in a high background. This study aims to identify DNA methylation biomarkers for multi-cancer detection. Utilizing large tissue datasets, we apply novel search algorithms to discover confined biomarker panels capable of distinguishing tumor from normal and determining the tissue of origin. We explore the applicability to blood-based testing using targeted methylation sequencing followed by machine learning classification. We present an 8-marker panel, which successfully predicts tumors across 14 types with a 91% average sensitivity, maintaining a low false positive rate (< 0.04%). Additionally, a panel of 39 CpG sites exhibits accuracies ranging from 69% to 98% for identifying tissue of origin. When tested on 114 patient plasma samples (colon, liver, pancreatic, prostate, and stomach cancer), the 8-marker panel obtains an AUC of 0.78 with a 78% sensitivity among 32 early-stage patients (stage I-II), and 60% overall. Using the 39-marker panel in a multi-class classification model selecting only the best match, 54% of tumor samples were on average correctly assigned to the tissue of origin, and up to 80% when allowing more inclusive criteria. Using a limited set of biomarkers, our work contributes to advancing non-invasive cancer diagnostics.
Uria-Regojo, G.; Fernandez-Caballero, L.; Lopez-Alcojor, A.; Lopez-Lopez, L.; Benitez, Y.; Rodilla, C.; Avila Fernandez, A.; Trujillo-Tiebas, M. J.; Osorio, A.; Corton, M.; Almoguera, B.; Ayuso, C.; Minguez, P.
Show abstract
Rare diseases (RDs) remain a major diagnostic challenge. Genetic and phenotypic heterogeneity, incomplete knowledge of disease mechanisms, and limitations in variant clinical interpretation leave many patients without a molecular diagnosis. Meanwhile, the growing volume of genomic data generated in clinical practice offers an opportunity to develop data-driven methodologies for exploring disease mechanisms and improving the reanalysis of unsolved cases. We aggregated real-world genomic data from 11,084 unrelated patients with suspected RD. Patients were clinically classified into 122 diseases. We built a multi-disease genomic variant frequency database (FJD-DB), which enabled the development of variant and gene-disease association scores by means of case-control subcohort comparisons across 32 disease groups. Functional enrichment analyses were then used to highlight disease-associated protein domains, pathways, biological processes, and phenotypes. Finally, the resulting knowledge was integrated into a data-driven framework for the guided reanalysis of unsolved RD patients applied to Inherited Retinal Dystrophies (IRD) patients as first use case. FJD-DB contained more than 45 million unique variants, including ~185,000 potentially pathogenic variants. Disease-specific analyses identified disease-associated pathogenic variants and highlighted both established and candidate disease genes. We detected 179 significantly enriched protein domains across 23 diseases, 124 Human Phenotype Ontology terms across 13 diseases, 79 Reactome pathways across 10 diseases, and 72 Gene Ontology biological processes across 8 diseases, revealing highly disease-specific functional signatures. Integration of disease-specific variant, gene, and functional association signals enabled the development of a data-driven framework for guided reanalysis of unsolved RD cases. Applied to more than 1,100 unsolved IRD cases, the framework generated clinically relevant findings in 26 patients, including four molecular diagnoses, seven candidate diagnoses, and 15 cases upgraded from non-informative findings to variants of uncertain significance. Aggregated real-world genomic data can be leveraged to identify disease-associated molecular signals generating novel biological hypotheses. A unified analytical framework provides a scalable strategy for knowledge discovery and guided reanalysis, facilitating the identification of overlooked and potentially novel genetic causes of RDs.
Ou, Y.; Zhao, S.; Weng, J.; Long, Y.; Li, H.; Hou, Y.; Xiong, Q.; Liu, S.; Wang, Z.; Xu, Y.; Pan, H.; Zhang, H.; Sun, L.
Show abstract
Gut microbial structural variation captures strain-level genomic diversity beyond species abundance, yet its contribution to rheumatoid arthritis (RA) remains unclear. Integrating gut metagenomic datasets from four independent cohorts (n = 491), we identified reproducible RA-associated structural variants (SVs), most of which occurred in species without differential abundance. Incorporating SVs into machine learning models consistently improved disease discrimination across independent validation cohorts. Functional analyses identified a core deletion SV in Agathobacter rectalis that removes an XRE-family transcriptional regulator. Motif discovery and 3D structural modeling demonstrated sequence-specific binding of the regulator to the promoter of a short-chain fatty acid biosynthetic gene, supporting a strain-level regulatory mechanism. Together, these findings establish microbial structural variation as a complementary functional layer beyond taxonomy for RA discrimination and mechanistic interpretation.
Buonaiuto, S.; Marsico, F.; Mohammed, A.; Chinthala, L. K.; Amos-Abanyie, E. K.; Genetics Center, R.; Prins, P.; Mozhui, K.; Rooney, R. J.; Williams, R. W.; Davis, R. L.; Finkel, T. H.; Brown, C. W.; Colonna, V.
Show abstract
The Biorepository and Integrative Genomics (BIG) Initiative in Tennessee has developed a pioneering resource to address gaps in genomic research by linking genomic, phenotypic, and environmental data from a diverse Mid-South population, including underrepresented groups. We analyzed 13,152 exomes from BIG and found significant genetic diversity, with 50% of participants inferred to have non-European or several types of admixed ancestry. Ancestry within the BIG cohort is stratified, with distinct geographic and demographic patterns, as African ancestry is more common in urban areas, while European ancestry is more common in suburban regions. We observe ancestry-specific rates of novel genetic variants, which are enriched for functional or clinical relevance. Disease prevalence analysis linked ancestry and environmental factors, showing higher odds ratios for asthma and obesity in minority groups, particularly in the urban area. Finally, we observe discrepancies between self-reported race and genetic ancestry, with related individuals self-identifying in differing racial categories. These findings underscore the limitations of race as a biomedical variable. BIG has proven to be an effective model for community-centered precision medicine. We integrated genomics education, and fostered great trust among the contributing communities. Future goals include cohort expansion, and enhanced genomic analysis, to ensure equitable healthcare outcomes.
Lennon, N. J.; Kottyan, L. C.; Kachulis, C.; Abul-Husn, N.; Arias, J.; Belbin, G.; Below, J. E.; Berndt, S.; Chung, W.; Cimino, J. J.; Clayton, E. W.; Connolly, J. J.; Crosslin, D.; Dikilitas, O.; Edwards, D. R. V.; Feng, Q.; Fisher, M.; Freimuth, R.; Ge, T.; Glessner, J. T.; Gordon, A.; Guiducci, C.; Hakonarson, H.; Harden, M.; Harr, M.; Hirschhorn, J.; Hoggart, C.; Hsu, L.; Irvin, R.; Jarvik, G. P.; Karlson, E. W.; Karlson, E. W.; Khan, A.; Khera, A.; Kiryluk, K.; Kullo, I.; Larkin, K.; Limdi, N.; Linder, J. E.; Loos, R.; Luo, Y.; Malolepsza, E.; Manolio, T.; Martin, L. J.; McCarthy, L.; Mei
Show abstract
Polygenic risk scores (PRS) have improved in predictive performance supporting their use in clinical practice. Reduced predictive performance of PRS in diverse populations can exacerbate existing health disparities. The NHGRI-funded eMERGE Network is returning a PRS-based genome-informed risk assessment to 25,000 diverse adults and children. We assessed PRS performance, medical actionability, and potential clinical utility for 23 conditions. Standardized metrics were considered in the selection process with additional consideration given to strength of evidence in African and Hispanic populations. Ten conditions were selected with a range of high-risk thresholds: atrial fibrillation, breast cancer, chronic kidney disease, coronary heart disease, hypercholesterolemia, prostate cancer, asthma, type 1 diabetes, obesity, and type 2 diabetes. We developed a pipeline for clinical PRS implementation, used genetic ancestry to calibrate PRS mean and variance, created a framework for regulatory compliance, and developed a PRS clinical report. eMERGEs experience informs the infrastructure needed to implement PRS-based implementation in diverse clinical settings.
Nnam, C. F.; Salas, L.; Mboya, E. A.; Li, Y.; Zhang, M.; Kolling, F.; Perrard, L.; Palys, T. J.; Pflugradt, E.; Pioli, P. A.; Ernstoff, M. S.; Seigne, J. D.; Pettus, J. R.; Ren, B.; Song, L.; Christensen, B. C.
Show abstract
BackgroundRetrotransposable elements (RE) comprise approximately 45% of the human genome and are typically repressed by DNA methylation to preserve genomic integrity. In cancer, global DNA hypomethylation can lead to RE derepression, resulting in genomic instability and activation of innate immune pathways through viral mimicry. While individual RE classes have been examined in clear cell renal cell carcinoma (ccRCC), the integrated epigenetic landscape of multiple RE families and their clinical relevance remain incompletely characterized. MethodsWe performed a genome-wide prediction of DNA methylation across three major RE classes (Alu, LINE-1, and LTR elements) using a validated computational framework applied to Illumina methylation array data from two independent ccRCC tumor cohorts. Integrated unsupervised clustering of RE methylation profiles was used to define the epigenetic subtypes. Associations with clinicopathologic variables, tumor immune microenvironment composition (DNA Methylation-derived), hypoxia signaling, innate immune activation, and overall survival were evaluated. Prognostic relevance was assessed using multivariable Cox regression models adjusting for age, sex, AJCC stage or AUA risk group, and immune and angiogenic tumor microenvironment features. Key findings were then externally validated in CPTAC-ccRCC and independently replicated in an institutional Dartmouth Cancer Center (DCC) cohort with matched methylation and RNA-sequencing data. ResultsIntegrated clustering identified three reproducible RE methylation subtypes, Repressed, Transient, and Active. In the discovery cohort, the Active subtype showed significantly worse overall survival than the Repressed subtype, with a graded survival pattern across RE methylation states that persisted after multivariable adjustment. RE hypomethylation was associated with reduced EPAS1 (HIF2A) expression, increased immune infiltration, elevated PD-1 expression, and heightened cGAS-STING and interferon signaling, consistent with an immune-inflamed yet immunosuppressed tumor state. In the external CPTAC validation cohort, RE methylation subtypes recapitulated key molecular features and showed supportive survival trends. In the independent DCC replication cohort, an Active RE state was again associated with poorer survival, lower EPAS1 expression, increased PD-1 expression, greater CD8 T-cell and Treg infiltration, and elevated T-cell exhaustion signatures, supporting the reproducibility of the prognostic and immune-exhausted phenotype across cohorts. ConclusionsWe identified RE methylation subtypes with distinct molecular, immunologic, and prognostic features in ccRCC. External validation in CPTAC and independent replication in DCC support the robustness of this RE methylation framework across large-scale and institutional cohorts. These findings highlight the prognostic potential of RE methylation profiles and support their integration into molecular classification strategies to improve risk stratification in ccRCC.
Salvucci, M.; Poynter, L.; Minerzami, R.; Carberry, S.; O'Byrne, R.; Alexander, J.; Lambrechts, D.; Veselkov, K.; Kinross, J.; Prehn, J. H. M.
Show abstract
ObjectiveHere, we systematically investigated alterations in the bacteriome and mycobiome of CRC patients in tumours and matched adjacent mucosa resulting in the identification of microbiome-based subtypes associated with host clinico-pathological and molecular characteristics. DesignDiversity and composition of bacteriome and mycobiome of tumour and adjacent mucosa, resulting subtypes were computationally deconvoluted from RNA sequencing, using >10000 samples from in-house and publically available patient cohorts. ResultsThe bacteriome of tumours had higher dominance and lower -diversity compared to matched adjacent local and distant mucosa. Tumours were enriched with Proteobacteria (Gammaproteobacteria class), Fusobacteria (including Fusobacterium Nucleatum species) and Basidiomycota fungi (Malasseziaceae family). Tumours were depleted of Bacteroidetes (Bacteroidia class), Firmicutes (Clostridia class) and Ascomycota (Sordariomycetes and Saccharomycotina). Tumours and adjacent mucosa samples were classified into 4 microbial subtypes, termed C1 to C4, based on the bacteriome and mycobiome composition. The bacterial Propionibacteriaceae, Enterobacteriaceae, Fusobacteriaceae, Bacteroidaceae and Ruminococcaceae and the fungal Malasseziaceae, Saccharomycetaceae and Aspergillaceae were among the key families driving the microbial subtyping. Microbial subtypes were associated with distinct tumour histology and patient phenotypes and served as an independent prognostic marker for disease-free survival. Key associations between microbial subtypes and alterations in host immune response and signalling pathways were validated in the TCGA pan-cancer cohort. The microbial subtyping demonstrated stratification value in the pan-cancer settings beyond merely representing differences in survival by cancer type. ConclusionsThis study demonstrates the translational potential of microbial subtyping in CRC patient stratification, and provides avenues to design tailored microbiota modulation therapy to further precision oncology. Statement of significanceO_ST_ABSWhat is already known on this subject?C_ST_ABSO_LIThe microbiome has been implicated in the pathogenesis, progression and therapeutic response in patients diagnosed with CRC and other cancers. C_LIO_LIThe vast majority of studies to date has focussed on investigating the bacteriome while the critical role played by the mycobiome in shaping cancer has begun to be explored more recently. C_LIO_LIThe bacterial and fungal composition and diversity in on-tumour tissue compared to matched local and distal off-tumour mucosa is largely unexplored. C_LIO_LITumorigenesis may be promoted via alterations of the microbiome ecosystem that may be better recapitulated by a multi-kingdom microbial signature rather than by the abundance of individual bacterial or fungal microorganisms. C_LI What are the new findings?O_LIOn-tumour tissue was enriched with Fusobacteria, Proteobacteria, Basidyomicota and depleted of Bacteroidetes, Firmicutes, Ascomycota compared with adjacent off-tumour mucosa. C_LIO_LIWe stratified >600 CRC patients into four distinct microbial-based subtypes (C1-C4) according to their bacteriome and mycobiome composition. The microbial subtypes were associated with distinct prognosis, clinical phenotypes such as staging, tumour location, history, lymphovascular invasion, TP53 status and clinical outcome. C_LIO_LIFurthermore, the majority of matched adjacent mucosa samples were classified as C1 and paired tumour-matched normal samples demonstrated a robust shift towards the C1 subtype in off-tumour tissue, suggesting that the C1 subtype may recapitulate a healthier-like microbiome. This hypothesis was supported by the microbial subtyping of colon samples from healthy subjects that were categorised as C1 almost exclusively. C_LIO_LIThe identified microbial subtyping demonstrated stratification value in the pan-cancer settings (n=28 additional solid cancer indications spanning >9000 samples) beyond merely representing differences in survival by cancer type, providing the strongest stratification in liver cancer. C_LI How might it impact on clinical practise in the foreseeable future?O_LIThis study laids the foundation to develop a microbial signature as biomarker to clinically manage CRC and other solid cancers and potentially lead to microbiome-based companion diagnostics to design microbiota modulation therapy tailored to specific patients subgroups with distinct bacteriomes and mycobiomes. C_LI
Dong, X.; Wu, B.; Wang, H.; Yang, L.; Chen, X.; Ni, Q.; Wang, Y.; Liu, B.; Lu, Y.; Zhou, W.
Show abstract
BackgroundQuantitatively describe the phenotype spectrum of pediatric disorders has remarkable power to assist genetic diagnosis. Here, we developed a matrix which provide this quantitative description of genomic-phenotypic association and constructed an automatic system to assist the diagnose of pediatric genetic disorders. Results20,580 patients with genetic diagnostic conclusions from the Childrens Hospital of Fudan University during 2015 to 2019 were reviewed. Based on that, a phenotype spectrum matrix -- cGPS (clinical Genes Preferential Synopsis) -- was designed by Naive Bayes model to quantitatively describe genes contribution to clinical phenotype categories. Further, for patients who have both genomic and phenotype data, we designed a ConsistencyScore based on cGPS. ConsistencyScore aimed to figure out genes that were more likely to be the genetic causal of the patients phenotype and to prioritize the causal gene among all candidates. When using the ConsistencyScore in each sample to predict the causal gene for patients, the AUC could reach 0.975 for ROC (95% CI 0.972-0.976 and 0.575 for precision-recall curve (95% CI 0.541-0.604). Further, the performance of ConsistencyScore was evaluated on another cohort with 2,323 patients, which could rank the causal gene of the patient as the first for 75.00% (95% CI 70.95%-79.07%) of the 296 positively genetic diagnosed patients. The causal gene of 97.64% (95% CI 95.95%-99.32%) patients could be ranked within top 10 by ConsistencyScore, which is much higher than existing algorithms (p <0.001). ConclusionscGPS and ConsistencyScore offer useful tools to prioritize disease-causing genes for pediatric disorders and show great potential in clinical applications.
Ferraro, F.; Drost, M.; van der Linde, H.; Bardina, L.; Smits, D.; de Graaf, B. M.; Schot, R.; van Bever, Y.; Brooks, A.; Donker Kaat, L.; Brosens, E.; Deng, R.; Barakat, T. S.; Verhoeven, V. J. M.; van Ham, T. J.; Kleefstra, T.; Rots, D.
Show abstract
Multiple developmental and congenital disorders due to genetic variants or environmental exposures are associated with unique genome-wide alterations in DNA methylation (DNAm). Consequently, these patterns referred to as DNAm signatures, can be leveraged for diagnostic purposes by developing artificial intelligence (AI) models that enable molecular subclassification of individuals. Notably, DNAm signature application has been particularly successful for diagnosing individuals affected by developmental disorders and congenital anomalies, especially those with defects in genes encoding the Mendelian epigenetic machinery. So far, over 100 DNAm distinctive signatures have been reported for these disorders. However, the translation into diagnostic practice remains challenging not only because of the scarcity of samples from these (ultra)rare disorders needed to train DNAm AI models, but also due to the privacy regulations that restrict the sharing of affected individuals data, lack of methods for standardization, limited replication across different centers, and the emergence of commercial entities with competing interests. In this study, we show that synthetic cases, meaning in silico cases generated from publicly available DNAm data from unaffected individuals, and summarized data derived from anonymized study cohorts of affected individuals with certain disorders, can be used to train DNAm classifiers. We demonstrate that these DNAm classifiers trained on a large cohort of synthetic cases have an improved performance compared to previously published classifiers trained on cohorts of affected individuals only, which typically are limited in size due to the rarity of these conditions. Furthermore, they improve the classification of variants with intermediate effect and mosaic cases and do not require any private affected individual data for training. Finally, to facilitate dissemination of these models, we release 169 synthetic cases-trained DNAm classifiers for 89 disorders with MethaDory, an open-access tool for simultaneous testing of these DNAm signatures.
Bassiouni, R.; Jin, Y.; Gibbs, L. D.; Qian, J.; Rotimi, S. O.; Miller, H.; Webb, M. G.; Rajpara, S.; Arias-Stella, J.; Craig, D. W.; Roman, L.; Carpten, J. D.
Show abstract
The mortality rate of ovarian cancer remains disproportionately high compared to its incidence. This is partly due to a high level of intra-tumoral heterogeneity that promotes disease recurrence and treatment failure. In this study, we describe degrees of heterogeneity revealed by single-cell whole genome sequencing and spatial transcriptomics of five epithelial ovarian carcinomas. At the cellular level, we describe pseudo-diploid cells that match the malignant cell population in both somatic variant and copy number patterns. At the clonal and subclonal levels, we describe diversification associated with copy number gains and whole genome doubling. In multi-clonal samples, we infer evolutionary relationships from single cell copy number, loss of heterozygosity analysis, and somatic variant detection, and correlate these with tissue histology and gene expression programs. In one sample, we identify functionally consequential copy number alterations that contribute to molecular diversity, cell proliferation, and inflammation in a minor clone that persisted without major expansion alongside a more complex major clone. In another, we describe a complex evolutionary history including a spontaneous reversion of a driver mutation in a secondary clone, which correlated with a switch in oncogenic expression programs.
Schobers, G.; Derks, R.; Ouden, A. D.; Swinkels, H.; Reeuwijk, J. V.; Bosgoed, E.; Lugtenberg, D.; Sun, S. M.; Corominas Galbani, J.; Weiss, J.; Blok, R.; Olde Keizer, R.; Hofste, T.; Hellebrekers, D.; Leeuw, N. D.; Stegmann, A.; Kamsteeg, E.-J.; Paulussen, A.; Ligtenberg, M.; Bradley, X. Z.; Peden, J.; Gutierrez, A.; Pullen, A.; Payne, T.; Gilissen, C.; Wijngaard, A. V.; Brunner, H.; Nelen, M.; Yntema, H.; Vissers, L.
Show abstract
BackgroundTo diagnose the full spectrum of hereditary and congenital diseases, genetic laboratories use many different workflows, ranging from karyotyping to exome sequencing. A single generic high-throughput workflow would greatly increase efficiency. We assessed whether genome sequencing (GS) can replace these existing workflows aimed at germline genetic diagnosis for rare disease. MethodsWe performed GS (NovaSeq6000; 37x mean coverage) on 1,000 cases with 1,271 known clinically relevant variants, identified across different workflows, representative of our tertiary diagnostic centers. Variants were categorized into small variants (single nucleotide variants and indels <50 bp), large variants (copy number variants and short tandem repeats) and other variants (structural variants and aneuploidies). Variant calling format files were queried per variant, from which workflow-specific true positive rates (TPRs) for detection were determined. A TPR of [≥]98% was considered the lower threshold for transition to GS. A GS-first scenario was generated for our laboratory, using diagnostic efficacy and predicted false negative as primary outcome measures. As input, we modeled the diagnostic path for all 24,570 individuals referred in 2022, combining the clinical referral, the transition of the underlying workflow(s) to GS, and the variant type(s) to be detected. ResultsOverall, 95% (1,206/1,271) of variants were detected. Detection rates differed per variant category: small variants in 96% (826/860), large variants in 93% (341/366), and other variants in 87% (39/45). TPRs varied between workflows (79-100%), with 7/10 being replaceable by GS. Models for our laboratory indicate that a GS-first strategy would be feasible for 84.9% of clinical referrals (750/883), translating to 71% of all individuals (17,444/24,570) receiving GS as their primary test. An estimated false negative rate of 0.3% could be expected. ConclusionGS can capture clinically relevant germline variants in a GS-first strategy for the majority of clinical indications in a genetics diagnostic lab.
Kenkre, R.; Larson, J. D.; Chapman, O. S.; Luebeck, J.; Lo, Y. Y.; Paul, M.; Zhang, W.; Bafna, V.; Wechsler-Reya, R.; Chavez, L.
Show abstract
Extrachromosomal DNA (ecDNA) is a powerful oncogenic driver linked to poor prognosis in pediatric cancers. Whole-genome sequencing of 338 patient-derived xenograft (PDX) samples and 127 matched primary tumors across multiple childhood cancer types was used to compare ecDNA prevalence, sequence conservation, and clonal dynamics. ecDNA in PDX models frequently mirrored oncogene amplifications observed in patient tumors (e.g., MYCN, MYC, MDM2) and showed high sequence conservation. Medulloblastoma and neuroblastoma PDXs exhibited significantly higher ecDNA prevalence, consistent with strong selection or de novo formation during tumor propagation. Although ecDNA copy numbers were generally preserved, some neuroblastoma PDXs displayed marked MYCN copy gains. Single-cell multiome profiling revealed that ecDNA-positive clones either persisted or expanded dramatically in PDXs, in one case growing from a minor subpopulation to nearly all tumor cells. These findings establish PDX models as valuable systems for ecDNA research and underscore the selective growth advantage conferred by ecDNA during tumor evolution.
Rota Negroni, M.; Billato, I.; Romualdi, C.
Show abstract
Copy number alterations (CNAs) are major contributors to genomic instability in cancer, and copy number signatures (CNS) provide a compact representation of the processes shaping CNA landscapes. However, the relationships among existing CNS frameworks and their predictability from molecular data other than whole-genome sequencing remain unclear. Here, we compare three major CNS compendia across more than 5,800 TCGA cancer samples, evaluating their overlap, complementarity, biological relevance, and prognostic associations. Individual signatures showed limited cross-study concordance, whereas signature-derived clusters identified biologically distinct patient groups, including favorable-outcome clusters observed across all frameworks. Using gene expression, DNA methylation, somatic mutation features, age, and tumor purity, XGBoost models predicted cluster membership with framework-dependent performance, achieving high F1-scores for the Drews and Steele compendia but limited performance for Tao. Feature importance analysis highlighted expression-driven predictors and pathways linked to genomic instability. These findings show that current CNS frameworks capture complementary rather than interchangeable dimensions of tumor genome instability and suggest that multi-omic profiles can extend signature-based stratification to cohorts without whole-genome sequencing.