Back

Collaborative and privacy-preserving workflows on a clinical data warehouse: an example developing natural language processing pipelines to detect medical conditions

Petit-Jean, T.; Gerardin, C.; Berthelot, E.; Chatellier, G.; Frank, M.; Tannier, X.; Kempf, E.; Bey, R.

2023-09-11 health informatics
10.1101/2023.09.11.23295069 medRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSObjectiveC_ST_ABSTo develop and validate advanced natural language processing pipelines that detect 18 conditions in clinical notes written in French, among which 16 comorbidities of the Charlson index, while exploring a collaborative and privacy-preserving workflow. Materials and methodsThe detection pipelines relied both on rule-based and machine learning algorithms for named entity recognition and entity qualification, respectively. We used a large language model pre-trained on millions of clinical notes along with clinical notes annotated in the context of three cohort studies related to oncology, cardiology and rheumatology, respectively. The overall workflow was conceived to foster collaboration between studies while complying to the privacy constraints of the data warehouse. We estimated the added values of both the advanced technologies and the collaborative setting. ResultsThe 18 pipelines reached macro-averaged F1-score positive predictive value, sensitivity and specificity of 95.7 (95%CI 94.5 - 96.3), 95.4 (95%CI 94.0 - 96.3), 96.0 (95%CI 94.0 - 96.7) and 99.2 (95%CI 99.0 - 99.4), respectively. F1-scores were superior to those observed using either alternative technologies or non-collaborative settings. The models were shared through a secured registry. ConclusionsWe demonstrated that a community of investigators working on a common clinical data warehouse could efficiently and securely collaborate to develop, validate and use sensitive artificial intelligence models. In particular, we provided efficient and robust natural language processing pipelines that detect conditions mentioned in clinical notes.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.