Designing Cell Cell Relation Extraction Models: A Systematic Evaluation of Entity Representation and Pre-training Strategies
Yoshikawa, M.; Mizuno, T.; Ohto, Y.; Fujimoto, H.; Kusuhara, H.
Show abstract
Extracting cell-cell relations from biomedical literature is essential for understanding intercellular communication in immunity, inflammation, and tissue biology. However, cell-cell relation extraction has not been established as a standalone biomedical relation extraction task, and no benchmark corpus or systematic evaluation framework currently exists. Fully manual corpus construction is costly and difficult to scale, limiting literature-based analyses of cell-cell communication. Here, we define a sentence-level cell-cell relation extraction task and construct complementary manually annotated corpora under realistic annotation constraints. To enable scalable annotation, rule-based literature mining is used solely as an annotation accelerator to identify candidate sentences, while all relation labels are assigned manually. In addition, an independently annotated PubMed corpus without rule-based filtering is constructed to evaluate robustness on natural sentence distributions. Using these resources, we evaluate representative model configurations involving entity indication strategies, classification architectures, and continued pre-training. Our results show that cell-cell relation extraction remains challenging under realistic conditions. Increasing training data size yields consistent performance gains, and specific combinations of entity-aware architectures and continued pre-training provide modest robustness improvements. Nevertheless, performance on unfiltered PubMed sentences remains in the 70% accuracy range, and error analyses indicate that failures cannot be readily explained by simple surface-level factors. Comparisons with general-purpose large language models further suggest that task complexity, rather than model class, is the primary limiting factor. Together, these findings establish a practical foundation for literature-scale cell-cell relation extraction while clarifying its intrinsic limitations.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- LSD600: the first corpus of biomedical abstracts annotated with lifestyle–disease relations 95%
- A Sequence Labeling Framework for Extracting Drug-Protein Relations from Biomedical Literature 95%
- RegulaTome: a corpus of typed, directed, and signed relations between biomedical entities in the scientific literature 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.