Back

AI-driven framework modeling perturbation in brain organoids reveals candidate genes for autism

Koh, I. G.; Chang, E.; Choi, Y. S.; Kim, S.-W.; Kim, Y.; Lee, H.; Byeon, G.; Ryu, Y.; Kim, S.; Lee, J.; Park, H.; Sim, H.; Ryu, Y.; Shim, W.; Lee, J.; Salazar, N. B.; de Aquino, M. M.; Engchuan, W.; Zhou, X.; Son, J. H.; Lee, J.; Bong, G.; Kim, I. B.; Han, J. H.; Werling, D. M.; Kim, S. H.; Oh, M.; Kim, M.-S.; Lee, D.; Kim, J.; Lee, Y.-S.; Sun, W.; Kim, E.; Scherer, S. W.; Jeon, M.; Yoo, H. J.; An, J.-Y.

2026-08-26 bioinformatics
10.64898/2026.08.22.746387 bioRxiv
Show abstract

Autism gene discovery is constrained by the rarity and heterogeneity of damaging variants, requiring large cohorts to identify susceptibility genes. Neural organoids and single-cell foundation models enable perturbation modeling in neurodevelopmental contexts. Here, we show that perturbation-informed foundation modeling of neural organoids can provide functional context for prioritizing candidate genes with genomic and clinical support. We constructed a 3.6-million-cell organoid atlas and trained models to predict genome-wide perturbation responses. Benchmarking 17 models identified a telencephalic neuron-specific model best preserving autism-relevant perturbation structure. Genome-wide profiling revealed two clusters associated with mid-fetal synaptic neuronal processes and early radial glia ubiquitin signaling. These clusters were supported by damaging-variant enrichment and clinical phenotypes across 89,916 family-based samples. Logistic-regression prioritization identified 343 candidates, including 167 in the key clusters, with convergence across TADA signals and recurrent evidence for NBEA and KLHDC10. This framework integrates predicted perturbation effects with genomic evidence to support autism candidate prioritization.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.