Constructing cancer-specific patient similarity network with clinical significance
Zhang, R.; Liu, Z.; Zhu, C.; Cai, H.; Yin, K.; Zhong, F.; Liu, L.
Show abstract
Clinical molecular genetic testing and molecular imaging dramatically increase the quantity of clinical data. Combined with the extensive application of electronic health records, medical data ecosystem is forming, which summons big-data-based medicine model. We tried to use big data analytics to search for similar patients in a cancer cohort and to promote personalized patient management. In order to overcome the weaknesses of most data processing algorithms that rely on expert labelling and annotation, we uniformly adopted one-hot encoding for all types of clinical data, calculating Euclidean distance to measure patient similarity, and subgrouping via unsupervised learning model. Overall survival was investigated to assess the clinical validity and clinical relevance of the model. Thereafter, we built a high-dimensional network cPSN (clinical patient similarity network). When performing overall survival analysis, we found Cluster_2 had the longest survival rates while Cluster_5 had the worst prognosis among all subgroups. Because patients in the same subgroup share some clinical characteristics, clinical feature analysis found that Cluster_2 harbored more lower distal GCs than upper proximal GCs, shedding light on the debates. Overall, we constructed a cancer-specific cPSN with excellent interpretability and clinical significance, which would recapitulate patient similarity in the real-world. The constructed cPSN model is scalable, generalizable, and performs well for various data types. The constructed cPSN could be used to accurately "locate" interested patients, classify the patient into a disease subtype, support medical decision making, and predict clinical outcomes.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ViPal: A Framework for Virulence Prediction of Influenza Viruses with Prior Viral Knowledge Using Genomic Sequences 94%
- Discovering Signature Disease Trajectories in Pancreatic Cancer and Soft-tissue Sarcoma from Longitudinal Patient Records 94%
- A Utility-Based Machine Learning-Driven Personalized Lifestyle Recommendation for Cardiovascular Disease Prevention 94%
Similar papers in this journal
- Unsupervised Discovery of Risk Profiles on Negative and Positive COVID-19 Hospitalized Patients 95%
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 95%
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 95%
Similar papers in this journal
- Using Automated-Machine Learning to Predict COVID-19 Patient Survival: Identify Influential Biomarkers 95%
- Clinical Characteristics And Prognostic Factors For ICU Admission Of Patients With COVID-19 Using Machine Learning And Natural Language Processing 92%
- Fear of Infection and Sufficient Vaccine Reservation Information Might Drive Rapid Coronavirus Disease 2019 Vaccination in Japan: Evidence from Twitter Analysis 92%
Similar papers in this journal
- Regional medical inter-institutional cooperation in medical provider network constructed using patient claims data from Japan 97%
- Machine learning based prediction of recurrence after curative resection for rectal cancer 95%
- Predicting Adverse Drug Effects: A Heterogeneous Graph Convolution Network with a Multi-layer Perceptron Approach 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.