Dual-LLM Adversarial Framework for Information Extraction from Research Literature
Li, Z.; Yu, Y.; Gu, W.; Zhu, T.; Song, H.; Guo, W.; Yang, X.; Zhu, Z.
Show abstract
Information Extraction (IE) is a fundamental task in Natural Language Processing (NLP) that aims to automatically identify relevant information from unstructured or semi-structured data. Information extraction from lengthy research literature, particularly in multi-omics studies, faces significant challenges due to their complex narratives and extensive context. To address this, we present a novel dual-LLM adversarial framework in which one large language model (LLM) performs the extraction and another provides iterative feedback to refine the results. This process systematically reduces errors, enhances consistency across heterogeneous data sources, and converges toward more accurate outputs. We evaluated our approach against manual and single-LLM extraction, using LLMs as evaluators. Experimental results show that our adversarial framework outperforms these baselines, highlighting its effectiveness for extracting structured information from lengthy scientific texts.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Species-Agnostic Transfer Learning for Cross-species Transcriptomics Data Integration without Gene Orthology 93%
- The current landscape and emerging challenges of benchmarking single-cell methods 92%
- Evaluation of out-of-distribution detection methods for data shifts in single-cell transcriptomics 92%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 96%
- Knowledge Graph-based Thought: a knowledge graph enhanced LLMs framework for pan-cancer question answering 94%
- ShinyLearner: A containerized benchmarking tool for machine-learning classification of tabular data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.