Hide and Seek: Privacy-Preserving Artificial Intelligence with a Feasibility Study in Rare Disease Diagnosis
Rajaganapathy, S.; St. Sauver, J.; Pinto e Vairo, F.; Iyer, P. G.; Liu, H.; Fan, J. W.
Show abstract
BackgroundIntegrating advanced artificial intelligence (AI) into clinical decision-support often requires the sharing of sensitive patient data with external services, raising privacy concerns. Homomorphic encryption (HE) allows computing directly on encrypted data, without revealing the underlying patient information. ObjectivesTo develop a large language model (LLM)-assisted diagnosis framework while preserving patient privacy in the clinical text analysis, by leveraging HE and using rare disease (RD) diagnosis as a representative application. To demonstrate HE does not hinder the system performance. Materials and MethodsTexts from patient histories and a RD knowledge base were embedded by LLMs into vectors, then encrypted using HE to obscure private information while retaining the semantic nuances. Diagnostic recommendations were generated by computing and ranking the similarities between the patient history and RD vectors in the encrypted space. The system was evaluated using 50 synthetic case reports (5 RDs, each with 10 reports). ResultsApplying HE did protect private information from reverse-embedding attacks. HE imposed little disruption to the diagnostic accuracy, with normalized discounted cumulative gains (nDCG) of 0.6108 {+/-} 0.3412 (encrypted) versus 0.6083 {+/-} 0.3415 (unencrypted). The accuracy and computational performance were tunable and consistent, as demonstrated across five different LLMs. DiscussionOur privacy-preserving framework opens tremendous opportunities toward hosting and serving powerful AI solutions across institution boundaries, which would remove the need for local deidentification and incentivize users to access secure external decision-support services. ConclusionsIntegrating HE with LLM retrieval can promote the dissemination of nonredundant, high-capacity AI services by preserving both privacy and accuracy.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Building a Best-in-Class De-identification Tool for Electronic Medical Records Through Ensemble Learning 93%
- Privacy-Preserving Federated Neural Network Learning for Disease-Associated Cell Classification 93%
- Structuring clinical text with AI: old vs. new natural language processing techniques evaluated on eight common cardiovascular diseases 92%
Similar papers in this journal
- Information retrieval in an infodemic: the case of COVID-19 publications 92%
- A benchmark of online COVID-19 symptom checkers 92%
- Optimal policy determination in sequential systemic and locoregional therapy of oropharyngeal squamous carcinomas: A patient-physician digital twin dyad with deep Q-learning for treatment selection 91%
Similar papers in this journal
- Privacy Protection of Sexually Transmitted Infections Information from Chinese Electronic Medical Records 92%
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 92%
- Synthetic data for privacy-preserving clinical risk prediction 91%
Similar papers in this journal
- Explainable deep learning for disease activity prediction in chronic inflammatory joint diseases 93%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 92%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.