aiAtlas: High Fidelity Cell Simulations of Genetic Perturbations in Rare Diseases and Cancers
Danter, W. R.
Show abstract
BackgroundLinking genetic perturbations to cellular phenotypes remains a central challenge in translational biology. Experimental iPSC and organoid models are powerful but constrained by scalability, variability, and difficulty modeling rare or polygenic states. MethodsWe developed aiAtlas v1.2, a high-fidelity simulation platform that integrates Large Concept Model (LCM) logic with aiPSC-derived modeling. We evaluated 136 virtual cell lines spanning wild-type, single-mutation, multiple-mutation, human tumor-derived, and gene-fusion cohorts. Twenty-five features covering DNA damage/repair, replication stress, epigenetic remodeling, pluripotency, and stress responses were quantified. Statistical analysis used the Mann-Whitney U test with Bonferroni correction, Hodges-Lehmann estimators (HLE) for median differences, and Cliffs delta effect sizes with bootstrap 95% confidence intervals. Robustness measures included early stopping, bagging, and 5-fold cross-validation. ResultsaiAtlas v1.2 reliably separated wild-type and mutant cohorts, revealing consistent disruptions in DNA damage accumulation, replication stress, epigenetic dysfunction, and loss of pluripotency, while identifying stable features (e.g., core nucleotide-excision repair processes and selected apoptosis measures). Subgroup analyses showed shared systemic effects and context-specific vulnerabilities: single mutations frequently produced measurable divergence; multiple mutations amplified instability; tumor-derived and gene-fusion lines yielded distinct but partially overlapping phenotypes. Large effect sizes (Cliffs {delta}) with narrow bootstrap CIs supported reproducibility across cohorts. Conclusions/ImpactaiAtlas v1.2 provides a robust virtual subject framework that uses aiCRISPR-Like (aiCRISPRL) virtual gene editing system that complements wet-lab CRISPR models by scaling to diverse genomic contexts and highlighting both disruption and stability. The platform can guide therapeutic prioritization, gene-editing strategy design, and regulatory innovation consistent with the FDA Modernization Act 2.0, accelerating therapy development in rare diseases and cancer. Significance StatementaiAtlas introduces a scalable, reliable simulation framework that integrates advanced large concept model (LCM) logic with iPSC-derived cellular modeling. aiAtlas overcomes major limitations of experimental systems by capturing both broad and subgroup-specific phenotypic divergence across single mutations, multiple mutations, tumor-derived cell lines, and gene fusions. This reliable platform establishes a new opportunity for rare diseases and cancer modeling, offering reproducible insights that can accelerate discovery and translational applications where traditional wet-lab approaches are impractical. Furthermore, in situations where the target mutational profile has been defined but no cellular models yet exist, aiAtlas can quickly generate custom virtual cell lines that accurately reproduce the corresponding genomic and phenotypic features.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CNpare: matching DNA copy number profiles 92%
- Single Nucleotide Polymorphism (SNP) and Antibody-based Cell Sorting (SNACS): A tool for demultiplexing single-cell DNA sequencing data 92%
- Using Cancer Profiles to Identify Synthetic Lethal Therapeutic Targets and Predictive Biomarkers in Cancer Gene Dependency Data 92%
Similar papers in this journal
- Improved Quality Metrics for Association and Reproducibility in Chromatin Accessibility Data Using Mutual Information 93%
- scMuffin: an R package for disentangling solid tumor heterogeneity from single-cell expression data 92%
- Drug mechanism enrichment analysis improves prioritization of therapeutics for repurposing 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.