Quantitative Assessment of the Exposure Risk and Targeted Protection of Sensitive Information related to Mental Health Disorders in China's Electronic Medical Records
Gong, M.; que, m.; ouyang, z.; cai, e.; liu, c.; zhou, x.; liu, q.; lv, m.; zeng, z.; shi, w.; xiao, y.
Show abstract
BackgroundMental health issues affect populations worldwide, with depression, schizophrenia, and dementia being particularly prevalent in China, where the China Mental Health Survey (CMHS) reported a 7.4% lifetime prevalence of mood disorders. Mental disorders have become a leading cause of disability. While the widespread adoption of electronic medical records (EMRs) has significantly improved healthcare efficiency and resource allocation, the sensitivity of medical data poses serious privacy breach risks. Many patients withhold medical information due to data security concerns, increasing the risk of treatment discontinuation. Currently, the lack of unified management policies and technical standards for electronic health records (EHRs) has led to frequent unauthorized access, leaks, and illegal trading of data, exacerbating doctor-patient conflicts and societal stigma against individuals with mental illness. MethodsThis study developed a novel personally identifiable information (PII) desensitization protocol (EPPDI) to mitigate privacy risks through comprehensive database scanning (as opposed to traditional field-specific desensitization). The protocol incorporates the following technological innovations: (1) Expansion of a lexicon of 20 mental health-related keywords using a Word2Vec vector space model; (2) Application of regular expressions to replace sensitive information and its surrounding 10 characters with asterisks (*). ResultsAmong 1,235,651 patients (8,016,263 records), the EPPDI protocol achieved 97.60% precision and 95.40% recall, with a privacy protection efficacy rate of 97.85%. Diagnosis records (31.84%) and medication data (45.41%) were identified as primary leakage sources. Regional disparities were notable, with Beijing showing a PD identification rate of 21.06%, far exceeding Qinghais 1.66%. Regarding data utility preservation, among 2,000 pieces of PDI patient information, 48 (2.4% false positive rate) contained no sensitive content, while 92 (4.6% false negative rate) pieces of non-PDI patient information included sensitive data among 2,000 pieces of non-PDI patient information. ConclusionThe EPPDI protocol addresses challenges such as ambiguity in Chinese terminology and adaptation to unstructured narratives, providing a technical framework for implementing Chinas Personal Information Protection Law in mental health. Future efforts should focus on balancing privacy protection with research needs through dynamic, tiered desensitization approaches.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Deep Learning Approach for Transgender and Gender Diverse Patient Identification in Electronic Health Records 93%
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 93%
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 93%
Similar papers in this journal
- Development and Evaluation of Machine Learning Models for the Detection of Emergency Department Patients with Opioid Misuse from Clinical Notes 93%
- Long COVID symptoms from Reddit: Characterizing post-COVID syndrome from patient reports 92%
- Automatic Gender Detection in Twitter Profiles for Health-related Cohort Studies 92%
Similar papers in this journal
- Simulated Misuse of Large Language Models and Clinical Credit Systems 92%
- Predicting critical state after COVID-19 diagnosis: Model development using a large US electronic health record dataset 92%
- Natural Language Processing Techniques to Detect Delirium in Hospitalized Patients from Clinical Notes: A Systematic Review 91%
Similar papers in this journal
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 93%
- Emergence and Evolution of Big Data Analytics in HIV Research: Bibliometric Analysis of Federally Sponsored Studies 2000-2019 92%
- Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.