Point-in-time evidence and cross-area clinical precedent anticipate clinical entry across 100 focal areas: retrospective validation of the Intangia triage layer
Elliott, T. O.; Molnar, S.; Peeters, G.; Collart, O.
Show abstract
Early-opportunity teams face a combinatorial problem: once a focal target, mechanism or indication is fixed, the space of plausible partners runs to thousands of candidates per area. Intangia's triage layer ranks that space from point-in-time evidence (how much literature, patent and clinical activity a candidate pairing has accumulated, and whether the partner already has clinical precedent in other contexts) so that review starts where clinical activity is most likely to begin next. This preprint validates that capability retrospectively across 100 focal areas spanning drug targets, mechanisms and disease indications, replaying 24.1 million historically scored combination-years with every area scored by a model trained on the other 99 and never on itself. The headline is operational. At a twenty-partner review shortlist per focal area, the median area's four-year first-alert precision is 0.234, against a matched random-ranker median of 0.008: roughly one in four shortlisted partners subsequently entered the focal clinical context within four years, about 38 times each area's own background rate (95% CI 31 to 45). A panel-level permutation puts the result at p = 0.0005. Discrimination generalises: the full 13-feature specification reaches a median leave-one-focal-out ROC-AUC of 0.922 (95% CI 0.911 to 0.929), with no area below chance and all 100 areas beating their strongest count-based baseline. Shortlisted entrants are anticipated with a median observed lead of two years within the evaluation window, and three years (interquartile range one to five) once the window cap is removed and every realised entrant is counted. The core ranking is carried by two interpretable signal families: cumulative co-occurrence counts and leave-one-area-out clinical precedent. Burst detection serves a complementary role: it supplies the time-stamped, source-specific momentum evidence attached to every recommendation (what is accelerating, and why now) rather than additional ranking power. A conditional view of the same landscape ranks candidates with no cross-area precedent against one another, enriched relative to matched random ranking, supporting a lower-yield emerging-opportunities capability. Two worked examples, PD-1 combination immunotherapy and CTLA-4, are point-in-time historical replays of the same architecture in familiar territory, showing what an alert looked like with the dated evidence behind it. The endpoint throughout is first clinical entry, not clinical success; prospective validation is the next stage.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Cell-to-cell and type-to-type heterogeneity of signaling networks: Insights from the crowd 91%
- A tissue-aware machine learning framework enhances the mechanistic understanding and genetic diagnosis of Mendelian and rare diseases 91%
- Predicting antigen-specificity of single T-cells based on TCR CDR3 regions 90%
Similar papers in this journal
- Bi-level Graph Learning Unveils Prognosis-Relevant Tumor Microenvironment Patterns in Breast Multiplexed Digital Pathology 91%
- Knowledge-guided deep learning models of drug toxicity improve interpretation 91%
- Contrastive Learning Enables Epitope Overlap Predictions for Targeted Antibody Discovery 90%
Similar papers in this journal
- CELLector: Genomics Guided Selection of Cancer in vitro Models 92%
- Markov Field network integration of multi-modal data predicts effects of immune system perturbations on intravenous BCG vaccination in macaques 91%
- Integration of multi-modal measurements identifies critical mechanisms of tuberculosis drug action 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.