Cell phenotypes in the biomedical literature: a systematic analysis and text mining corpus
Rotenberg, N. H.; Leaman, R.; Islamaj, R.; Kuivaniemi, H.; Tromp, G.; Fluharty, B.; Richardson, S.; Eastwood, C.; Diller, M.; Xu, B.; Pankajam, A. V.; Osumi-Sutherland, D.; Lu, Z.; Scheuermann, R. H.
Show abstract
The variety of cell phenotypes identified by single-cell technologies is rapidly expanding, yet this knowledge is dispersed across the scientific literature and incompletely represented in structured resources. We present the CellLink corpus, a manually annotated collection of over 22,000 mentions of human and mouse cell populations in recent journal articles, distinguishing specific cell phenotypes, heterogeneous cell populations, and vague cell populations, and linking to Cell Ontology (CL) terms as either exact or related matches, covering nearly half of the terms in the current CL. A systematic analysis reveals lineage-specific patterns in how authors utilize anatomical context, molecular signatures, functional roles, developmental stage, and other attributes in cell naming. We show that fine-tuning transformer-based models on CellLink yields strong performance for named entity recognition, while embedding-based approaches support zero-shot entity linking and distinguishing exact from related matches. We further demonstrate the utility of CellLink to expand and refine the chondrocyte branch of CL.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.