Back

IAN: An Intelligent System for Omics Data Analysis and Discovery

Nagarajan, V.; Shi, G.; Horai, R.; Yu, C.-R.; Gopalakrishnan, J.; Yadav, M.; Liew, M.; Gentilucci, C.; Caspi, R. R.

2025-03-11 bioinformatics
10.1101/2025.03.06.640921 bioRxiv
Show abstract

IAN is an R package that addresses the challenge of integrating, analyzing and interpreting high-throughput "omics" data, using a multi-agent artificial intelligence (AI) system. IAN leverages popular pathway and regulatory datasets (KEGG, WikiPathways, Reactome, GO, ChEA) and the STRING database for protein-protein interactions to perform standard enrichment analysis. The individual enrichment results are then used to generate insightful summaries, for each of the datasets, using a large language model (LLM) through a multi-agent architecture. These summaries are then contextually integrated and interpreted by the LLM, guided by carefully engineered prompts and grounding instructions, to provide insightful explanations, system overview, key regulators, novel observations etc. We demonstrate IANs potential to facilitate biological discovery from complex omics data, by reanalyzing two already published data and evaluating the results. We also show remarkable performance of IAN, in terms of avoiding hallucination. IAN package, along with installation instructions and example usage, is available on https://github.com/NIH-NEI/IAN.

Published in Cell Reports Methods · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.