Curation at Scale with EPITOME: Extraction Pipeline for Immunological Texts and Open-Source Multimodal Enquiry
Richardson, E.; Lischwe, P.; Bennett, J.; Blazeska, N.; Greenbaum, J.; Harlan, B.; Lallmamode, N.; Marrama, D.; Talbott, M.; Vita, R.; Sette, A.; Kuether, K.; Peters, B.
Show abstract
The Immune Epitope Database (IEDB, iedb.org) has manually curated epitope data from over 26,000 publications across two decades. With PubMed adding [~]5,000 articles daily, traditional curation methods face scalability challenges. Given the multimodality of data contained in scientific papers, we have sought to build an open-source vision language model (VLM)-based tool that human curators can use to speed up and automate biological data curation. Here we present a multimodal document ingestion and Question-Answering (QnA) pipeline that ties traditional Optical Character Recognition (OCR) and text matching with Vision-Language Model (VLM) capabilities. The system, which we call EPITOME, implements three-stage processing: regex-based epitope and MHC molecule identification, visual element extraction from PDFs, and contextual indexing that links peptide sequences, MHC molecules, and assays to their locations across text, tables, and figures. This indexing is used to supply context for further VLM QnA. Our preliminary results from EPITOME demonstrate promising zero-shot performance of open-source VLMs that suggest promise for accelerating biocuration through a curator-in-the-loop process, with our evaluation identifying strategic points where curator-in-the-loop intervention can enhance overall system accuracy.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CROssBAR: Comprehensive Resource of Biomedical Relations with Deep Learning Applications and Knowledge Graph Representations 93%
- Language model-based B cell receptor sequence embeddings can effectively encode receptor specificity 92%
- DeepSpaceDB: a spatial transcriptomics atlas for interactive in-depth analysis of tissues and tissue microenvironments 92%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.