Back

Automated Discovery of Therapeutic Biomaterial for Renally Impaired Hyperuricemia Patients by Natural Language Processing and Machine Learning

Zeng, X.; Qiu, J.; Zhao, X.; Liu, K.; Zhao, L.; An, J.; Xiang, L.; Liu, T.; Wang, Z.; Xie, W.; Wang, M.; Luo, J.; Zhang, S.

2025-03-11 biochemistry
10.1101/2025.03.05.641578 bioRxiv
Show abstract

The exponential growth of scientific publications presents opportunities for researchers to identify valuable knowledge, especially in the highly interdisciplinary field --- biomaterials, where exploiting possible connections between unmet clinical needs and materials properties from literatures is crucial. However, with traditional literature reading, it is extremely challenging to marry unmet clinical needs with existing materials reported for different applications or other purposes. Here, to provide a not-renally cleared therapeutics for renally impaired hyperuricemia patients, we designed a multi-tiered framework MatWISE that fuses state-of-the-art natural language processing, semantic relationship mapping, and machine learning to automate the complex process of material discovery from a sea of scientific literatures published until December of 2022, and successfully identified and optimized {delta}-MnO2 into an orally administered, nonabsorbable uric acid (UA) lowering biomaterial. {delta}-MnO2 had superior serum and urine UA-lowering effect in three hyperuricemia mouse models, by comparing with a standard of care drug. {delta}-MnO2 is highly promising to serve as a safe and effective UA-lowering drug for renally impaired hyperuricemia patients. We demonstrated a new research paradigm for biomaterials that combining state-of-the-art machine learning techniques and a handful of experiments to discover a translationally relevant material from the massive existing research, for an unmet clinical need.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.