Back

Dual-LLM Adversarial Framework for Information Extraction from Research Literature

Li, Z.; Yu, Y.; Gu, W.; Zhu, T.; Song, H.; Guo, W.; Yang, X.; Zhu, Z.

2025-09-16 bioinformatics
10.1101/2025.09.11.675507 bioRxiv
Show abstract

Information Extraction (IE) is a fundamental task in Natural Language Processing (NLP) that aims to automatically identify relevant information from unstructured or semi-structured data. Information extraction from lengthy research literature, particularly in multi-omics studies, faces significant challenges due to their complex narratives and extensive context. To address this, we present a novel dual-LLM adversarial framework in which one large language model (LLM) performs the extraction and another provides iterative feedback to refine the results. This process systematically reduces errors, enhances consistency across heterogeneous data sources, and converges toward more accurate outputs. We evaluated our approach against manual and single-LLM extraction, using LLMs as evaluators. Experimental results show that our adversarial framework outperforms these baselines, highlighting its effectiveness for extracting structured information from lengthy scientific texts.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.