Systematic tissue annotations of -omics samples by modeling unstructured metadata
Hawkins, N. T.; Maldaver, M.; Yannakopoulos, A.; Guare, L. A.; Krishnan, A.
Show abstract
There are currently >1.3 million human -omics samples that are publicly available. This valuable resource remains acutely underused because discovering particular samples from this ever-growing data collection remains a significant challenge. The major impediment is that sample attributes are routinely described using varied terminologies written in unstructured natural language. We propose a natural-language-processing-based machine learning approach (NLP-ML) to infer tissue and cell-type annotations for -omics samples based only on their free-text metadata. NLP-ML works by creating numerical representations of sample descriptions and using these representations as features in a supervised learning classifier that predicts tissue/cell-type terms. Our approach significantly outperforms an advanced graph-based reasoning annotation method (MetaSRA) and a baseline exact string matching method (TAGGER). Model similarities between related tissues demonstrate that NLP-ML models capture biologically-meaningful signals in text. Additionally, these models correctly classify tissue-associated biological processes and diseases based on their text descriptions alone. NLP-ML models are nearly as accurate as models based on gene-expression profiles in predicting sample tissue annotations but have the distinct capability to classify samples irrespective of the -omics experiment type based on their text metadata. Python NLP-ML prediction code and trained tissue models are available at https://github.com/krishnanlab/txt2onto.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Neighborhood nonnegative matrix factorization identifies patterns and spatially-variable genes in large-scale spatial transcriptomics data 95%
- Sfaira accelerates data and model reuse in single cell genomics 95%
- N-of-one differential gene expression without control samples using a deep generative model 95%
Similar papers in this journal
- Automated assignment of cell identity from single-cell multiplexed imaging and proteomic data 93%
- An efficient not-only-linear correlation coefficient based on machine learning 93%
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 93%
Similar papers in this journal
Similar papers in this journal
- mcRigor: a statistical method to enhance the rigor of metacell partitioning in single-cell data analysis 94%
- ROSIE: AI generation of multiplex immunofluorescence staining from histopathology images 94%
- FastCCC: A permutation-free framework for scalable, robust, and reference-based cell-cell communication analysis in single cell transcriptomics studies 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.