Back

Nature

Springer Science and Business Media LLC

Preprints posted in the last 30 days, ranked by how well they match Nature's content profile, based on 645 papers previously published here. The average preprint has a 0.63% match score for this journal, so anything above that is already an above-average fit.

1
Ancient DNA reveals matrilineal organisation and recurrent unions between dominant matrilines in Iron Age Britain

Olalde, I.; Armit, I.; Büster, L.; Lillie, M.; Urkixo F. de Zuazo, E.; Ringbauer, H.; Akbari, A.; Castells Navarro, L.; Esteve-Gomez, D.; Goodchild, H.; Hamilton, D.; Puig i Riera, N.; Sanchez-Sanz, A.; Buckberry, J.; Budd, C.; Caffell, A.; Halkon, P.; Holst, M.; Jerand, P.; Panagiotakopulu, E.; Stephens, M.; Stubbings, M.; Ware, P.; Bleasdale, M.; Booth, T.; Callan, K.; Caughran, E.; Fischer, C.-E.; Frost, T.; Iliev, L.; Kearns, A.; Legge, M.; Mah, M.; Masters, M. K.; Manjila, N.; Nawaz, M.; Oppenheimer, J.; Ponce, P.; Primeau, C.; Silva, M.; Skoglund, P.; Swali, P.; Qiu, L.; Gregory, S.; Wo

2026-08-06 genetics 10.64898/2026.08.03.742615 medRxiv
Top 0.1%
56.7%
Show abstract

Kinship practices underpin all traditional societies, forming the basis for socially sanctioned reproductive unions, residence patterns and the inheritance of rights and property1,2. Although the relationship between biological relatedness and kinship is not always straightforward, ancient DNA studies are increasingly used to examine the extent to which biological relatedness underpinned social constructs of kinship in prehistoric societies3-6. Here, we report the analysis of genome-wide data for 534 individuals from the Arras Culture of Middle Iron Age northeast England (including 390 from Wetwang Slack, 100 from Pocklington, and 29 from Melton), finding evidence for communities with kinship systems structured along matrilineal lines. At Wetwang Slack, we reconstruct a 13-generation pedigree comprising 195 individuals structured around female-line connections: matrilineal transmissions greatly outnumbered patrilineal ones and male reproductive partners were largely absent from the cemetery, plausibly because they were buried in their own natal communities. Furthermore, the three main sites with robust sample sizes were characterised by non-overlapping dominant mitochondrial haplogroups, implying a maternal clan-based structure. Reproductive unions at Wetwang Slack suggest a recurrent alliance between two dominant maternal descent groups, with members of each group never reproducing with members of their own maternal lineage. Meanwhile, three individuals from lavishly furnished chariot burials at Wetwang Slack were close maternal relatives belonging to a lineage with consecutive generations of close kin unions, a pattern largely absent among other individuals at the site. These results indicate highly distinctive social practices among an elite group embedded in the wider kinship network of the Arras community.

2
Time-resolved lineage recording reveals a pre-existing, heritable cell state underlying metastatic potential

Park, J.; Chang, Y.; Schiffman, J. S.; Koyyalagunta, D.; Somayaji, H.; McQuillen, C. N.; Chan, J.; Morris, Q.; Landau, D.; Kim, H. H.; Choi, J.

2026-08-11 cancer biology 10.64898/2026.08.10.744013 medRxiv
Top 0.1%
51.5%
Show abstract

Metastasis causes most cancer deaths1,2, yet no recurrent mutation specifically drives it3,4, raising the possibility that metastatic potential is a non-genetic yet heritable cell state. Classic experiments established that metastatically predisposed subclones pre-exist within a tumor and that these predispositions are inherited over many cell divisions5, but what molecular states or factors underlie this predisposition remain unknown. While previous lineage recording studies6,7 mapped how tumors disseminate, their recording sites saturate too quickly to resolve when a lineage branched, or to attribute a state to its founder. Here we show, using a DNA Typewriter lineage recorder8 with nearly 1,000 recording sites in lung cancer cells, that metastatic potential is already present before dissemination, with colonization predicted by a pre-existing glycolytic state and further spread by expression of ENO1, a glycolytic enzyme that also moonlights as a cell-surface plasminogen receptor9. Profiling the pre-transplant cells and the post-transplantation tumors for both their transcriptomes and their lineage recordings, we reconstructed time-resolved lineage trees across three orthotopically transplanted mice. These trees trace each liver metastasis to a single founder of known pre-transplant state, dating each dissemination event from the primary lung. When every clone was scored before transplant against 349 genes recurrently heritable in vitro, both that set and the glycolytic state independently shifted a clones odds of colonizing the lung. At the gene level, sixteen genes were both heritable and predictive of colonization, and ENO1 alone also predicted which established clones spread further. Hypoxia, the program most strongly associated with phylogenetic fitness within the metastases, did not predict colonization when scored before transplant, separating niche-selected traits from the inherited cell state. Metastatic potential in this system is therefore transmitted along the lineage rather than acquired after seeding. Looking forward, we anticipate that time-resolved lineage recorders will enable the separation of the heritable and acquired components of the cellular heterogeneity seen in single-cell studies of tumor progression and drug tolerance.

3
Covalent bond formation caught in a LOV photoreceptor

Olasz, B.; Gotthard, G.; Nag, P.; Gonzalez-Viegas, M.; Mous, S.; Koczurowska, A.; Johnson, P. J. M.; Sato, T.; Shankar, M. K.; Wranik, M.; Solecka, A.; Melo, D.; Ke, R.; James, D.; Nass, K.; Langner, P.; Valerio, J.; Furrer, A.; Gashi, D.; de Wijn, R.; Popelar, T.; Pachota, M.; Wisniewska, M.; Dietze, T.; E, J.; Asghar, A.; Zabelskii, D.; Trost, F.; Koua, F. H. M.; Smyth, P.; Sobolev, E.; Kim, C.; Ozerov, D.; Letrun, R.; Turkot, O.; Doerner, K.; Bielecki, J.; Han, H.; Dworkowski, F.; Cirelli, C.; Schertler, G. F. X.; Schulz, J.; Bacellar, C.; Standfuss, J.; Bean, R.; Milne, C.; Heberle, J.; Sc

2026-08-06 biophysics 10.64898/2026.08.02.742302 medRxiv
Top 0.1%
42.5%
Show abstract

Light-oxygen-voltage (LOV) domains are blue-light photoreceptors of plants, algae and fungi, and among the most widely used tools in optogenetics. They switch on by forming a covalent thioether bond between a conserved cysteine and their flavin chromophore, in a reaction that needs a proton to cross from the cysteine to the flavin through a pocket containing essentially no water. Its mechanism has been debated for two decades1, and because the chemistry is over within a microsecond its elementary steps have stayed hidden. Here we combine 10 time-resolved serial femtosecond crystallography snapshots and infrared spectroscopy with QM/MM calculations to resolve the entire sequence of events at 1.4 [A] resolution: from excited-state distortion of the flavin ring (10-100 ps), through hydration of a surface channel (10 ns) and a single ordered water reaching the active site as the reactive cysteine shifts between its conformations (100-500 ns), to the thioether bond itself, caught half-formed at 1 {micro}s (half the molecules reacted, half still poised) and complete at 10-100 {micro}s. That water bridges the cysteine and the flavin and shuttles the proton, lowering the barrier from [~]35 to [~]15 kcal/mol and accelerating the reaction by roughly fourteen orders of magnitude (without it, the half-life would be [~]237,000 years), then departs before the bond forms. Proteins can therefore hydrate a dehydrated active site transiently and on demand to overcome otherwise prohibitive reaction barriers, a catalytic strategy that reaches well beyond photoreceptors.

4
The Pan-European Impact of the Balkan Hunter-Gatherers

Tabin, D. R.; Boric, D.; Sumer, A. P.; Soos, G.; Alhaique, F.; Bonsall, C.; Boroneant, A.; Candilio, F.; Cristiani, E.; Cullen, T.; Fiore, I.; Fotiadou, C. M.; Katsarou, S.; Lepic, A.; Maric, A.; Masciana, A.; Melis, R. T.; Mussi, M.; Papathanasiou, A.; Perles, C.; Price, T. D.; Korzow, K.; Sampson, A.; Soficaru, A.; Tagliacozzo, A.; Uno, K.; Yavuz, O. E.; Mallick, S.; Lazaridis, I.; Patterson, N.; Posth, C.; Pinhasi, R.; Rohland, N.; Hajdinjak, M.; Brielle, E.; Reich, D.

2026-08-09 genetics 10.64898/2026.08.05.743150 medRxiv
Top 0.2%
33.8%
Show abstract

During the Last Glacial Maximum (LGM) 26-19 thousand years ago (kya), Europe had three principal refugia not covered by ice: the Iberian, Italian, and Balkan peninsulas1,2. While the Iberian and Italian refugia have been studied with ancient DNA3-7, the legacy of the Balkan refugium has remained a mystery due to a lack of genetic data 30-12 kya. We present genome-wide data from nine newly reported individuals: three Epipaleolithic from Romania, one Epipaleolithic from Bosnia and Herzegovina, three Mesolithic from Greece, one Epipaleolithic from Southern Italy, and one Mesolithic from Sardinia. We find that the Epipaleolithic individuals from Romania and the Mesolithic individual from mainland Greece were from a previously unsampled population that was the primary source for later European hunter-gatherers. Westward expansion of Balkan Epipaleolithic people into the Italian Peninsula and mixture with a minority contribution from pre-LGM Italians, produced a population that then spread further west to become the primary ancestry of west-ern European Mesolithic people. Northeastward expansion, bypassing Italy, contributed the European ancestry source of Scandinavian and Eastern European hunter-gatherers. Southeastward expansion to the Aegean led to Anatolian-European mixtures before the spread of farming.

5
Emergency-granulopoiesis trajectory, but not the CD4/NK lymphocyte trajectory, is directionally reproduced as a correlate of sepsis mortality: a cross-cohort transcriptomic study with external testing and clinical-analogue triangulation

Su, L.; Zhang, L.; Huang, W.; Gui, C.; Gong, F.

2026-08-06 intensive care and critical care medicine 10.64898/2026.08.04.26359671 medRxiv
Top 0.2%
31.6%
Show abstract

In early sepsis the direction in which blood immune-cell transcriptional programmes move may carry prognostic information beyond a single baseline measurement, but whether such trajectory associations survive independent testing is unknown. We scored five immune modules, frozen before analysis, in three public longitudinal whole-blood microarray sepsis cohorts and fitted a logistic model ladder fixed in advance to the change per 24 hours in the two cohorts with mortality data (82 patients, 24 deaths), pooling by inverse-variance fixed-effect meta-analysis with Benjamini-Hochberg control. No association survived correction for multiple testing. The two leading signals were a rising CD4/NK lymphocyte trajectory associated with lower mortality (pooled odds ratio 0.53, 95% confidence interval 0.31 to 0.90) and a rising emergency-granulopoiesis trajectory associated with higher mortality (1.60, 0.92 to 2.79), both per one standard deviation. We then tested both in an independent transcriptomic cohort with serial sampling (63 patients, 15 deaths), scored by the identical frozen method, and against their cell-count analogues in an intensive-care database of 12,607 adults meeting Sepsis-3 criteria, of whom 744 to 4,206 had the serial measurements each analogue required. Independent testing separated the two signals, in the order opposite to the one discovery had suggested. The emergency-granulopoiesis association was reproduced in direction and effect size without reaching conventional significance on its own (validation odds ratio 1.72, 0.92 to 3.20, p=0.088; pooled 1.65, 1.09 to 2.50), was positive in all nine sensitivity analyses, each fixed before the estimates were examined, and was supported by two of its three analogues, including the neutrophil-to-lymphocyte ratio (1.31, 1.20 to 1.42). The CD4/NK association did not reproduce (1.06, 0.59 to 1.89), was null in the window most favourable to it, and received no support from an analogue well powered to detect the discovery effect. The discovery signal that looked most consistent failed independent testing.

6
Sequence architecture shapes human allele frequencies through ectopic gene conversion

Steyaert, W.

2026-08-28 genetics 10.64898/2026.08.26.747289 medRxiv
Top 0.4%
26.3%
Show abstract

Population-genetic interpretations of allele frequency commonly begin from a single-origin assumption: that a variant arose once and was subsequently shaped by selection, drift and demographic history. Here we show that for a substantial fraction of human variation this assumption fails, and fails predictably: ectopic gene conversion reintroduces the same nucleotide change generation after generation, at rates strongly structured by two fixed properties of genome architecture, homologous template length and donor-acceptor distance. Across 534 million gnomAD v4.1 variants, this recurrent input accounts for an estimated 4% of the rarest variants and more than 15% of common ones. De novo mutations in 11,963 trios show the process directly: the mutation rate is elevated up to 28-fold precisely where the same variant is already common, with the newly arising allele matching the paralogous donor in 94% of such cases at long templates. The effect is strongest within segmental duplications but extends along every chromosome. Affected positions show reduced linkage disequilibrium and, at common frequencies, are depleted by 30-40% among reported GWAS associations. For this fraction of human variation, allele frequency is encoded in genome sequence architecture rather than set by population-genetic processes alone.

7
Selection and surveillance of 5S ribosomal RNA genes in human populations

Sengl, L.; Bagaric, I.; Conil, C.; Seeleuthner, Y.; Mueller, M.; Klughammer, J.; Mages, S.; Cobat, A.; Bohlen, J.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.27.26361558 medRxiv
Top 0.4%
23.0%
Show abstract

The 5S ribosomal RNA gene is present in the human genome not once but in ~80 copies, arranged head to tail in a single array of ribosomal DNA on chromosome 1 -one of the most repetitive and least explored regions of the genome. Its product is one of the four RNAs in every ribosome and, when ribosome assembly fails, it activates the tumour suppressor p53. Whether these copies vary in sequence between people, and whether such variation has physiological or pathological consequences, is unknown. Using telomere-to-telomere genome assemblies, whole-genome sequences from ~490 000 UK Biobank participants, and ~940 GTEx transcriptomes, we find that every person carries copies bearing substitutions or indels, and that ~10% of people express such variant 5S rRNA. Mutating every position of the gene in vitro, we find that variants blocking incorporation into the ribosome map to the uL5/uL18 interface and activate p53. Remarkably, these same variants are depleted from human populations: selection has acted on the step that p53 monitors. Ribosomal DNA is thus a functional source of human genetic variation, long invisible to genome-wide analysis and shaped by the p53 pathway it controls.

8
Canonically minimal RNA-guided insertion sequences expand into large elements that disseminate antimicrobial resistance

Hu, K.; Xie, B.; Yang, H.; Rubin, B. E.

2026-08-21 microbiology 10.64898/2026.08.14.744561 medRxiv
Top 0.5%
22.0%
Show abstract

IS110 has emerged as a powerful genome-editing tool because it is the smallest RNA-guided system capable of diverse programmable insertions. Naturally existing elements are conventionally modeled as compact[~] 1.5-kb systems comprising a single transposase and a bridge RNA (bRNA). Using high-throughput junction mapping together with large-scale comparative genomics, we redefined the in vivo structural boundaries, growth, and mobilization of IS110 elements. We uncovered a previously unrecognized size continuum extending to[~] 100 kb, driven by progressive local expansion, with expanded loci being widespread across bacterial genomes. Experiments confirmed the activity of natural IS110s both well below and above the size range of previously characterized elements. These large systems preferentially accumulate adaptive cargo, including antimicrobial resistance determinants and heavy-metal detoxification systems, and are strongly enriched for plasmid-derived DNA. Boundary configurations at expanded loci and the range of partial excision intermediates they produce both indicate flexible sequence recognition by IS110, most commonly through half-matches between the bRNA and complementary DNA sequence. This sequence tolerance allows loci to expand with diverse cargo. Together, these findings redefine IS110 from a compact insertion sequence into a dynamic platform that disseminates adaptive cargo.

9
Contrasting selective pressures shape human pepsinogen A gene copy-number variation across Eurasia

Chen, Q.; Wang, F.; Liu, A.; Zhu, Z.; Wang, H.; Li, X.; Wu, D.; Zhang, G.

2026-08-21 genomics 10.64898/2026.08.17.745382 medRxiv
Top 0.6%
19.2%
Show abstract

Diet has repeatedly shaped human genomes, yet adaptation of protein digestion remains poorly understood. Using 1,348 haplotype-resolved assemblies, we reconstructed the structural evolution of the human pepsinogen A (PGA) locus and found a west-to-east increase in copy-number across Eurasia that tracked regional reliance on plant-derived protein. Modern and ancient genomes further revealed a recent selective sweep in the East but signatures of balancing selection in the West. We traced this divergence primarily to expansion of PGA34A, the most proteolytically active paralog in vitro, and showed that recurrent nonallelic homologous recombination continually generated structural diversity in this locus. Independent PGA expansions were also enriched in plant-dominant mammals. These findings link paralog-specific dosage variation in protein digestion to dietary adaptation across human populations and diverse mammalian lineages.

10
Spatial multi-omics and single-cell transcriptomics uncover senescence-associated cellular programs during colon adenoma to cancer progression

Yu, M.; Xie, Y.; Allen, L.; Carter, K.; Donato, E.; Cheng, L.; Kendziorski, C.; Matkowskyj, K. A.; De Marzo, A. M.; Hicks, J.; Reddi, D.; Newell, E. W.; Sun, W.; Grady, W. M.

2026-08-07 cancer biology 10.64898/2026.08.06.743310 medRxiv
Top 0.6%
18.8%
Show abstract

Colorectal cancer develops through a normal-adenoma-carcinoma sequence, yet only 5-10% of adenomas progress to malignancy, and the cellular programs governing that sequence remain poorly defined. Here we generate a spatial multi-omics atlas of human colon adenomas, combining Visium CytAssist and protein co-detection across 24 nonadvanced and advanced tubular adenomas with single-cell resolution Xenium Prime 5K profiling of 101 patient-matched normal, adenoma, and carcinoma cores from 16 patients. Integrating whole-transcriptome and 31-plex protein data identifies nine spatial clusters and two dysplastic epithelial populations that co-express stemness, proliferation, and senescence programs. These programs occupy a shared, spatially confined epithelial niche that expands from adenoma to carcinoma. Spatial analysis revealed GDF15, a senescence-associated secretory factor, mediated the coupling between senescence and stemness in advanced adenomas, and that GDF15-high epithelium locally excludes CD8+ T cells in adenoma and, more broadly, in carcinoma. These findings position senescence as a spatially instructive rather than merely tumor-suppressive program during colorectal carcinogenesis and suggest GDF15 might be a potential candidate target for cancer prevention and interception in the colon.

11
Corpusome, a cross-body-site human microbiome corpus for representation learning

Xuan, H.; Huang, Y.; Bian, J.

2026-08-29 microbiology 10.64898/2026.08.28.747922 medRxiv
Top 0.7%
18.6%
Show abstract

Machine-learning models of the human microbiome are trained mostly on stool samples from single cohorts, limiting cross-body-site representation and cross-study generalization. Progress is constrained less by algorithms than by the absence of a harmonized multi-body-site corpus carrying the technical metadata needed to model, rather than ignore, batch structure. Here we release Corpusome, a harmonized two-tier cross-body-site human microbiome corpus for representation learning: a harmonized corpus of 187,546 human microbiome samples integrating standardized profiles from curatedMetagenomicData, the American Gut Project, and the EBI MGnify platform. Corpusome follows a two-tier design preserving both functional depth and cross-body-site breadth: a shotgun tier (22,588 samples, 93 studies) with species- and pathway-level profiles, and a 16S tier (164,958 samples, from a full pull of 708 MGnify studies) with genus-level profiles extending coverage to oral, skin, respiratory, and urogenital sites. It spans six body sites and two modalities, with harmonized metadata for batch-aware modelling. Body-site signal exceeds technical/source variance in the 16S tier by approximately 2.4-fold.

12
FBXW11 Activity Regulates Radial Glial Expansion in Human Cerebral Organoids

Moreno, C. L.; King, H.; Trabish, S.; Thijs, M.; Beck, D.; Bergamasco, M.; D'Araujo, T. Y.; Mosca, T. J.; Weatheritt, R.; Neely, G. G.

2026-08-21 neuroscience 10.64898/2026.08.20.746070 medRxiv
Top 0.7%
18.4%
Show abstract

Human brain development depends on tightly coordinated gene-regulatory programs and the emergence of complex tissue architecture, making large scale functional interrogation difficult using conventional screen models. To overcome this challenge, we used a pooled CRISPR screening approach. Guided by neuro-specific whole-genome screens in Drosophila, we tested 129 poorly characterised human orthologs and found 8 that modify cerebral organoid development. Candidates were validated using individual CRISPR knockouts and mosaic competition assays. Among these candidates we describe FBXW11, a substrate-recognition component of the SCF E3 ubiquitin ligase complex, as a potent negative regulator of cerebral organoid expansion. FBXW11 loss increases radial glial abundance, expands ventricular-like domains, and impairs neuronal maturation. Mechanistically, FBXW11 associates with {beta}-catenin and alters WNT signalling. FBXW11 mutations cause the autosomal-dominant Mendelian syndrome Neurodevelopmental, Jaw, Eye and Digital syndrome (NEDJED), and we found that disease-associated variants mapped preferentially to WD40 substrate-binding repeats and {beta}-catenin contact regions, linking impaired substrate recognition to neurodevelopmental disease. Together, these findings identify FBXW11 as a conserved negative regulator of {beta}-catenin-dependent radial glial expansion and neuronal maturation during human cerebral brain development.

13
All-by-All Cytokine Receptor Pairing Network Unlocks Coding of Non-Natural T Cell States

Tao, P.; Tsui, K. C. Y.; Zhao, Y.; Jiang, H.; Good, Z.; Garcia, K. C.

2026-08-24 immunology 10.64898/2026.08.19.745857 medRxiv
Top 0.7%
18.2%
Show abstract

Cytokine receptor pairing rules, set by evolution, confine JAK-STAT signaling to a narrow region of a far larger combinatorial space. Of more than 1,200 pairings theoretically possible among the [~]36 JAK-associated human cytokine receptors, only [~]30-40 exist in nature. Using a double-orthogonal platform, we enforced pairings across the full all-by-all receptor matrix and resolved a fine-grained STAT atlas richer than the natural repertoire. Selected non-natural pairings generated emergent T cell states unpredictable from either parental receptor, with pairing orientation encoding signaling specificity. A synthetic IL-21R>IL-2R{beta} pairing, but not its reciprocal, drove a cytotoxic Tc17-like state, whereas natural IL-9R/{gamma}c drove a Tc1 fate despite similar STAT activation, showing that rebalancing quantitative STAT combinatorics can modulate T cell fate. Recombining IL-31R, not expressed in T cells, with STAT-biased receptor partners generated diverse states, several with superior antitumor efficacy. These findings define a non-natural pairing code for engineering synthetic T cell fates.

14
Mutagenic bypass of uracil-derived abasic sites underlies the SBS17 mutational signature

Di Ciccio, S.; van Strien, J.; Jiang, Y.; Abbasi, A.; Bakker, L.; Weertman, N.; Leiteris, L.; Willems, J.; Alexandrov, L. B.; Garaycoechea, J. I.

2026-08-21 molecular biology 10.64898/2026.08.21.746158 medRxiv
Top 0.8%
18.1%
Show abstract

Understanding how mutations arise is central to predicting and preventing cancer, yet most of the mutational signatures found in cancer genomes still have no known mechanistic cause. Among these, single-base-substitution signature 17 (SBS17) is the dominant mutational process in gastric and oesophageal adenocarcinoma. SBS17 has been linked to oxidative damage and to the chemotherapeutic agent 5-fluorouracil (5-FU), but the molecular events that generate it are unknown. Here we demonstrate that SBS17 causes driver mutations in gastrointestinal cancer. We then combine genetics in cell lines and organoids with whole-genome sequencing to define the mechanistic aetiology of SBS17 mutagenesis. We disprove the prevailing oxidative damage hypothesis and instead show that SBS17 arises through the misincorporation of dUTP during DNA replication, followed by uracil excision by the glycosylase UNG and mutagenic bypass of the resulting abasic (AP) sites by the translesion synthesis machinery. Importantly, we reveal that the same mechanism underlies both spontaneous and chemotherapy-induced mutations. Together, our results resolve the origin of gastric mutagenesis, opening therapeutic avenues for curbing ongoing mutagenesis, tumour evolution and drug resistance in gastrointestinal cancers.

15
Probabilistic mouse-human brain correspondence by multimodal optimal transport

Koch, S. P.; Perales, M.; Sil, T.; Hiew, S.; Hartig, J.; Weigl, B.; Muthuraman, M.; Lange, F.; Li, N.; Kuhn, A.; Breininger, K.; Ip, C. W.; Volkmann, J.; Reich, M.; Boehm-Sturm, P.; Peach, R. L.

2026-08-27 neuroscience 10.64898/2026.08.24.746652 medRxiv
Top 0.8%
17.9%
Show abstract

The mouse is the principal model for brain mechanism and disease, allowing experiments that cannot be performed in humans. However, findings often translate poorly because homologous regions differ in relative size and some human territories have no clear mouse counterpart. Here we present OTTER, which learns mouse-human brain correspondence as a probabilistic, parcel-resolution coupling using multimodal fused Gromov-Wasserstein optimal transport, integrating functional and structural connectivity with spatial position and curated homologies. OTTER recovers established homologues on a transcriptomic benchmark and preserves broad cross-species organisation along the cortical areal hierarchy. Applying the coupling to human functional connectivity reveals a graded decline in mouse-based reconstruction across evolutionarily expanded association cortex, with the lowest values in lateral prefrontal territory. Finally, the bidirectional map generates testable human predictions from mouse experiments and ranks mouse circuits corresponding to human clinical targets.

16
Pangenome discovery and characterization of human protein-coding duplicated genes

Ren, L.; Yoo, D.; Vlajic, K.; Dishuck, P. C.; Guitart, X.; Kwon, Y.; Lin, J.; Munson, K. M.; Hoekzema, K.; Stergachis, A.; Vollger, M. R.; Schweppe, D. K.; Eichler, E. E.

2026-08-06 genomics 10.64898/2026.08.05.743125 medRxiv
Top 0.8%
16.6%
Show abstract

Protein-coding genes mapping to high-identity segmental duplications (SDs) have been difficult to annotate and characterize and are the source of most previously unknown protein-coding genes being discovered as part of the human pangenome. Here, we combine long-read assembled human genomes (298) and long-read transcriptome data (5.6 billion full-length cDNA from 83 tissues) to phylogenetically interrogate 493 gene families discovering 2713 potentially copy number polymorphic genes not present in the human reference genome. For reference SD gene families where paralog specificity can be assigned, we find that 60.0% are expressed and maintain open reading frames, with 45.7% showing high expression in brain, embryo, or testis. We revise 386 gene models, including 150 that absent or different from current T2T-CHM13 gene annotation and 236 (35.1%) pseudogenes as protein-coding where we find evidence of transcription, an open reading frame, and chromatin-accessible promoters. We find that 24.2% of SD genes show evidence of constraint for both copy number and amino acid mutation. The majority of these constraint genes are ancestral, whereas only 16.2% of derived duplicated genes that emerged recently in the human lineage show evidence of constraint. The pangenome provides unparalleled specificity to understand genetic variation in SD genes allowing us to distinguish functional genes from pseudogenes and highlighting potential gene innovations that arose most recently in human evolution.

17
Functional Pairing of TCR-peptide-MHC Interactomes by Single-cell Clonal Expansion

Liu, L.; Shin, S. W.; Joslin, K.; Wang, C.; Pai, J.; Ma, R.; Xiang, X.; Clark, I. C.; Garcia, K. C.

2026-08-23 immunology 10.64898/2026.08.18.745550 medRxiv
Top 0.8%
16.3%
Show abstract

T cell antigen-specific immunity depends on pairwise interactions between T cell receptors and peptide-MHC, yet isolating the TCR-pMHC pairs that drive productive engagement remains a major obstacle for antigen-specific therapeutics and for decoding TCR specificity. We overcome this by co-encoding TCR and pMHC in a single founder cell, then clonally expanding it inside a semi-permeable capsule so that genetically identical daughter cells engage in trans. T cell activation, rather than binding affinity, is used to sort cells with functional pairs, and a single PCR on the clone's linked genomic library captures both partners. This platform, LINC-seq, recovered known cognate pairs from pooled libraries at up to 95% accuracy and performed simultaneous, library-on-library deep mutational scanning of both partners. Wild-type clonotypes ranked among the top-enriched sequences in complex mixtures, and the screens resolved co-evolutionary epistasis and cross-reactivity rules inaccessible to one-sided mutagenesis. The approach generalizes to any receptor-ligand pair whose trans-engagement drives a reporter.

18
A comprehensive atlas of somatic mutation rates and mutational signatures in normal human cells

Pham, M. H.; Harvey, L. M. R.; Oliver, T. R. W.; Dunstone, E.; Lawson, A. R. J.; Nicola, P. A.; Sanghvi, R.; Hooks, Y.; Mitchell, E.; Jarman, G. L.; Wang, Y.; Abascal, F.; Jung, H.; Neville, M. D. C.; Ishida, Y.; Fowler, J. C.; Le, A. P.; Moody, S.; Marshall, H.; Brzozowska, N.; Ding, C.; Pac, C. A.; Machado, H. E.; O'Neill, L.; Latimer, C.; Humphreys, L.; Saeb-Parsy, K.; Mahbubani, K. T. A.; Baxter, J.; Rassl, D. M.; Vicario, R.; Geissmann, F.; Kabashima, K.; Bleys, R. L. A. W.; Moore, L.; Heer, R.; Coorens, T. H. H.; Behjati, S.; Hoare, M.; Campbell, P. J.; Jones, P. H.; Martincorena, I.; Ra

2026-08-29 genomics 10.64898/2026.08.28.747772 medRxiv
Top 0.9%
15.3%
Show abstract

Over the course of a lifetime, somatic mutations accrue in normal human cells, causing variation in cell phenotype and engendering somatic evolution with outcomes ranging from the adaptive immune system to cancer. To inform understanding of somatic evolution in the human body we report the mutation rates and mutational signatures of 53 normal cell types. Most show evidence of linear mutation accumulation over time with single base substitution mutation rates ranging from ~3.5/year/diploid genome in spermatogonia and sperm, to ~20/year in postmitotic neurons, ~50/year in mitotically active colorectal epithelial cells, ~60/year in kidney proximal tubule cells and hepatocytes, 100s/year in sun-exposed skin epidermal cells and 10-50/year in the remainder. Certain cell types, including skin epidermis, cardiac myocytes, bladder urothelium, kidney proximal tubule cells, and hepatocytes, show substantial variability in mutation burdens around the linear age trend, indicating the influence of additional factors which differ between individuals and modulate mutation accumulation, including exogenous mutagen exposures. At least 18 single-base substitution and nine small insertion and deletion mutational signatures are present, some in all cell types, some in a subset and others in a single cell type. Known exogenous mutagen exposures and endogenous mutational processes account for some mutational signatures, but the origins and mechanisms underlying many are uncertain. This comprehensive survey of mutagenesis provides a foundation for understanding somatic evolution of human cell populations in health and disease.

19
A Cytokine Receptor Signaling Atlas Reveals How STAT Mosaics Fine-Tune T Cell Function

Tao, P.; Rastogi, R.; Jiang, H.; Zhao, Y.; Su, L. L.; Jude, K.; Kundaje, A.; Garcia, K. C.

2026-08-24 immunology 10.64898/2026.08.19.745852 medRxiv
Top 0.9%
15.3%
Show abstract

The extent to which JAK/STAT cytokine signaling is functionally redundant or selective remains debated. Here we engineered a double orthogonal IL-2/IL-2R{beta}/{gamma}c ternary system enabling programmable, interference-free activation of each of the 36 mammalian cytokine receptors, and their downstream six STATs, in T cells. At the membrane-proximal level, comprehensive phospho-signaling profiling revealed that while each receptor activates a dominant STAT, unique STAT activation fingerprints derived from combinatorial biases fine-tune nuanced T cell fates. At the membrane-distal level, single-cell transcriptomic atlas of all cytokine receptors confirmed that these STAT mosaics sensitively specify non-redundant transcriptional programs. STAT5-dominant receptors drove proliferative expansion at the expense of stemness; STAT3-driven programs instructed a continuum from stem cell memory to terminal effector states with preserved cytotoxic capacity and mediated superior curative antitumor responses; while other STATs specified highly restricted phenotypes. These findings decode a STAT signaling vocabulary that defines the intrinsic functional bandwidth of natural cytokines.

20
G2T: Tissue Reconstruction from Gene Expression via Embedding-Distance Flow Matching

Birk, S.; Theis, F. J.; Lotfollahi, M.

2026-08-28 genomics 10.64898/2026.08.25.746917 medRxiv
Top 0.9%
15.1%
Show abstract

Single-cell RNA sequencing (scRNA-seq) profiles transcriptomes at high resolution but discards the spatial context of cells within a tissue -- information that is essential for studying intercellular mechanisms and tissue architecture. Spatial transcriptomics (ST) retains coordinates but, depending on the assay, trades this off against gene-panel breadth, spatial resolution, or cost. We present G2T (Gene-to-Tissue), a generative deep learning model that reassembles a tissue from gene expression -- its only observed input -- by predicting the matrix of pairwise distances between cells in a learned embedding space. G2T uses an attention-based Transformer with an Euclidean-Distance-Matrix (EDM) output head and is trained with conditional flow matching: the network learns to denoise corrupted cell positions, conditioned on the slice's gene expression, by predicting per-cell embeddings whose pairwise squared distances match the ground-truth distance matrix. At inference, a fast locally-optimal-block (LOBPCG) multidimensional scaling step turns the predicted distance matrix into 2-D coordinates. On a published MERFISH mouse primary motor cortex benchmark, G2T improves over the previous state-of-the-art method, LUNA, across all three standard metrics -- Spearman correlation of pairwise-distance ranks, Contact F1, and per-cell-class Sum RSSD -- and even larger relative gains on the mouse central-nervous-system scRNA-seq atlas, evaluated against an imputed spatial reference (STARmap PLUS-integrated locations, not measured coordinates). By predicting this geometry in a higher-dimensional embedding space rather than regressing 2-D coordinates, G2T relaxes the 2-D output parameterisation of prior diffusion-based methods and yields a compact, scalable building block for reconstructing tissue from dissociated cells, enabling downstream spatial niche and cell-cell communication analysis.