Back

ARCADE: Enhancing Automated Document Analysis Through Adversarial Multi-Agent Validation

Kumar, C. J.; Pearson, H. J.; Reddy, C. L.

2025-12-23 health informatics
10.64898/2025.12.21.25342744 medRxiv
Show abstract

We present ARCADE (Adversarial Critique Architecture for Document Evaluation), a multi-agent architecture addressing three limitations of traditional retrieval-augmented generation for automated document analysis: incomplete information extraction, shallow analytical depth, and framework paraphrasing. We compared ARCADE against Single-Pass RAG using 95 policy documents (50 National Cancer Control Plans and 45 Cardiovascular Disease plans) evaluated on 36 metrics across six capabilities: Natural Language Understanding, Logical Reasoning, Fluency, Coherence, Factual Grounding, and Synthesis. Results demonstrate three primary improvements. First, ARCADE achieves 100% extraction success compared to 60% for Single-Pass RAG, eliminating systematic failures that affect nearly half of framework dimensions. Second, analytical depth increases substantially (93-150% longer responses) while readability simultaneously improves by 12 points on the Flesch Reading Ease scale, moving from "very difficult" to "difficult" reading level. Third, multiple independent metrics confirm genuine critical evaluation rather than redundancy: 60% score disagreement between methods, 17-30% factual content overlap, and decreased framework semantic similarity. These gains require minimal trade-off: a 9% reduction in lexical diversity suggests more consistent, if slightly repetitive, terminology. These findings establish adversarial multi-agent validation as an effective paradigm for automated document assessment, with clear applications in policy analysis, regulatory compliance, and clinical protocol evaluation.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.