Back

A Large-Scale Concordance Study of Toxicity Findings Across Preclinical Species and Humans for Small Molecules and Biologics in Drug Development

Liu, X.; Fan, F.

2026-01-30 bioinformatics
10.64898/2026.01.29.702667 bioRxiv
Show abstract

Translating preclinical safety findings into reliable insights for human risk assessment remains a fundamental challenge in drug development. Prior preclinical-clinical concordance studies have been constrained by limited drug coverage, reliance on identical-term matching for adverse events (AEs), and insufficient consideration of species, modality, exposure, and biological or mechanistic context. To address these gaps, we assembled a large cross-species concordance dataset, integrating standardized preclinical and clinical safety data for 7,565 marketed and investigational drugs from PharmaPendium and OFF-X. Our framework employs likelihood ratios to reduce prevalence bias and extends concordance assessment beyond identical-term matches to include semantically and mechanistically related AE pairs. Stratified analyses by species, modality, and exposure-matched subsets further refined translational relevance, while integration of on- and off-target annotations supports mechanistic interpretation and potential screening. Using this approach, we identified 850 significant identical-term AEs and 2,833 additional unique endpoints from cross-term associations. To promote reproducibility and transparency in animal research, we provide open access to the analytic code and statistical results via an interactive web application. An accompanying multi-agent AI system (ToxAgents) enables standardized querying and interpretation of concordance results. Together, these resources extend previous foundational efforts and establish a shared, data-driven platform to advance translational safety science, support evidence-based study design aligned with the 3Rs, and ultimately contribute to the development of safer medicines to improve human health.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.