Back

Interaction-finder: automated literature-based discovery of biological entity associations with quote-level provenance

Chapman, T. E.; Lassmann, T.

2026-07-10 bioinformatics
10.64898/2026.07.07.736901 bioRxiv
Show abstract

Identifying interactions between biological entities is a cornerstone of molecular research, but assembling such lists from the literature is slow and tedious. For many research questions, no curated database exists, leaving researchers to survey the relevant literature themselves. We present interaction-finder, a tool that automates this process: given a topic string and user-defined entity types, it discovers relevant literature through O_SCPLOWLLMC_SCPLOW-guided iterative search, extracts candidate associations from full-text articles, and produces a ranked list where every association is backed by quoted passages verified against the source text. A self-contained interactive O_SCPLOWHTMLC_SCPLOW report enables rapid triage of the results. Evaluated across 60 topics in three domains (celltype-cellmarker, disease-gene, and ligand-receptor), interaction-finder recalls 1.2-4.3x as many known associations as single-shot prompting and an off-the-shelf deep-research framework, with all extracted quotes verified against source text. To assess candidates unrecognised from the gold-standard databases, we scored each candidate using an independent O_SCPLOWLLMC_SCPLOW judge blind to the tools reasoning. Across the three domains, unverified candidates score similarly to gold-standard associations. We find the gold-standard associations are enriched at the top of our ranked candidates, with an overall recall@20 of 0.61. Interaction-finder is freely available at https://github.com/tecosaur/interaction_finder under an O_SCPLOWMITC_SCPLOW licence.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.