Structured Understanding of Assessment and Plans in Clinical Documentation
Stupp, D.; Barequet, R.; Lee, I.-C.; Oren, E.; Feder, A.; Benjamini, A.; Hassidim, A.; Matias, Y.; Ofek, E.; Rajkomar, A.
Show abstract
Physicians record their detailed thought-processes about diagnoses and treatments as unstructured text in a section of a clinical note called the assessment and plan. This information is more clinically rich than structured billing codes assigned for an encounter but harder to reliably extract given the complexity of clinical language and documentation habits. We describe and release a dataset containing annotations of 579 admission and progress notes from the publicly available and de-identified MIMIC-III ICU dataset with over 30,000 labels identifying active problems, their assessment, and the category of associated action items (e.g. medication, lab test). We also propose deep-learning based models that approach human performance, with a F1 score of 0.88. We found that by employing weak supervision and domain specific data-augmentation, we could improve generalization across departments and reduce the number of human labeled notes without sacrificing performance.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 95%
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 94%
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 94%
Similar papers in this journal
- Deep representation learning for clustering longitudinal survival data from electronic health records 94%
- Segmenting functional tissue units across human organs using community-driven development of generalizable machine learning algorithms 92%
- Deep transfer learning for reducing health care disparities arising from biomedical data inequality 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.