dnaSORA - A Unified Diffusion Transformer for DNA point clouds
Njie, e. G.; Koreniuk, O.
Show abstract
The relatively obscure Hawaiian experiment collapses diverse phenotypes, including nearly all human genetic diseases to a singular Gaussian-like point cloud feature, structuring unstructured information. The uniformity of the feature space provides a straightforward way for AI models to learn all three billion tokens for reading the human genome as a first language. We propose a diffusion transformer, dnaSORA, for learning these features. dnaSORA has generative capacity similar to Stable Diffusion but for DNA point clouds. The models architecture is novel because it is unified; thus, it also functions as a discriminator that uses a frozen latent representation for classification. dnaSORA transfer learns from synthetic data emulating real genome point clouds to classify misrepresented tokens in C. elegans Hawaiian data at state-of-the-art 0.3 Mb resolution. Pre-training large genome models typically requires expensive and difficult-to-obtain genomes. However, our solution provides nearly unlimited synthetic training data at negligible compute costs. Inference for new token assignments (e.g., new diseases) requires genomes from several dozen rather than thousands of individuals. These efficiencies, combined with state-of-the-art resolution, provide a pathway for rapid, massive scaling of token annotation of the entire human genome at orders of magnitude below expected costs.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Private information leakage from functional genomics data: Quantification with calibration experiments and reduction via data sanitization protocols 94%
- Evolving super stimuli for real neurons using deep generative networks 93%
- SPLASH: a statistical, reference-free genomic algorithm unifies biological discovery 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.