Back

Scalable search of massively pooled nucleic acid samples enabled by a molecular database query language

Berleant, J. D.; Banal, J. L.; Rao, D. K.; Bathe, M.

2024-04-15 health informatics
10.1101/2024.04.12.24305660 medRxiv
Show abstract

Conventional collection, preservation, and retrieval of nucleic acid specimens, particularly unstable RNA, require costly cold-chain infrastructure and rely on inefficient robotic sample handling, hindering downstream analyses. These generate critical bottlenecks for global pathogen surveillance and genomic biobanking efforts, prohibiting large-scale nucleic acid sample collection and analyses that are needed to empower pathogen tracing, as well as rare disease diagnostics1. Here, we introduce a scalable nucleic acid storage system that enables rapid and precise retrieval on pooled nucleic acid samples--stored at room-temperature with minimal physical footprint2,3--using versatile database-like queries on barcoded, encapsulated samples. Queries can incorporate numerical ranges, categorical filters, and combinations thereof, which is a significant advancement beyond previous demonstrations limited to single-sample retrieval or Boolean classifiers. We apply our system to a pool of ninety-six mock SARS-CoV-2 genomic samples identified with theoretical patient data including patient age, geographic location, and diagnostic state, allowing rapid, multiplexed nucleic acid sample retrieval in a scalable manner to empower genomic analyses. By avoiding expensive and cumbersome freezer storage and retrieval systems, our approach in principle scales to millions of samples without loss of fidelity or throughput, thereby supporting the development of large-scale pathogen and genomic repositories in under-resourced or isolated regions of the US and worldwide.

Published in Nature Communications (predicted rank #6) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.