Back

Dual Inference Routes for Natural Language Hypothesis Generation in Multimodal Drug Discovery: An Early Experimentation Study

WU, R.; Matsuno, H.; Ge, Q.-W.

2025-09-07 bioinformatics
10.1101/2025.09.02.673688 bioRxiv
Show abstract

This study presents an early experimentation of a modular AI framework for multimodal drug discovery that integrates natural product based therapies with modern pharmaceuticals. The system combines structured biomedical data, knowledge graphs, and large language models (LLMs) to generate explicit natural language hypotheses. The architecture has four phases: data aggregation, hypothesis generation, dynamic simulation, and in silico evaluation, and supports dual inference routes (compound [->] gene [->] disease/symptom and disease/symptom [->] gene [->] compound). As a case study, Phases 1 and 2 were applied to the Kampo formula Shakuyaku-kanzo-to, a typical example of a multicomponent and multi-target natural therapy. The framework originally arose from challenges in conventional filtering, where important but poorly annotated compounds were often overlooked. However, the focus has since shifted beyond filtering, toward uncovering hidden relationships across fragmented biomedical knowledge. This early implementation demonstrates the potential of natural language hypothesis generation to restructure fragmented knowledge into interpretable insights, providing a blueprint toward future multimodal drug discovery.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.