Back

Hide and Seek: Privacy-Preserving Artificial Intelligence with a Feasibility Study in Rare Disease Diagnosis

Rajaganapathy, S.; St. Sauver, J.; Pinto e Vairo, F.; Iyer, P. G.; Liu, H.; Fan, J. W.

2026-01-17 health informatics
10.64898/2026.01.15.26344228 medRxiv
Show abstract

BackgroundIntegrating advanced artificial intelligence (AI) into clinical decision-support often requires the sharing of sensitive patient data with external services, raising privacy concerns. Homomorphic encryption (HE) allows computing directly on encrypted data, without revealing the underlying patient information. ObjectivesTo develop a large language model (LLM)-assisted diagnosis framework while preserving patient privacy in the clinical text analysis, by leveraging HE and using rare disease (RD) diagnosis as a representative application. To demonstrate HE does not hinder the system performance. Materials and MethodsTexts from patient histories and a RD knowledge base were embedded by LLMs into vectors, then encrypted using HE to obscure private information while retaining the semantic nuances. Diagnostic recommendations were generated by computing and ranking the similarities between the patient history and RD vectors in the encrypted space. The system was evaluated using 50 synthetic case reports (5 RDs, each with 10 reports). ResultsApplying HE did protect private information from reverse-embedding attacks. HE imposed little disruption to the diagnostic accuracy, with normalized discounted cumulative gains (nDCG) of 0.6108 {+/-} 0.3412 (encrypted) versus 0.6083 {+/-} 0.3415 (unencrypted). The accuracy and computational performance were tunable and consistent, as demonstrated across five different LLMs. DiscussionOur privacy-preserving framework opens tremendous opportunities toward hosting and serving powerful AI solutions across institution boundaries, which would remove the need for local deidentification and incentivize users to access secure external decision-support services. ConclusionsIntegrating HE with LLM retrieval can promote the dissemination of nonredundant, high-capacity AI services by preserving both privacy and accuracy.

Published in Journal of Clinical and Translational Science · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.