Parsing Clinical Trial Eligibility Criteria for Cohort Query by a Multi-Input Multi-Output Sequence Labeling Model
Tian, S.; Yin, P.; Zhang, H.; Erdengasileng, A.; Bian, J.; He, Z.
Show abstract
To enable electronic screening of eligible patients for clinical trials, free-text clinical trial eligibility criteria should be translated to a computable format. Natural language processing (NLP) techniques have the potential to automate this process. In this study, we explored a supervised multi-input multi-output (MIMO) sequence labelling model to parse eligibility criteria into combinations of fact and condition tuples. Our experiments on a small manually annotated training dataset showed that that the performance of the MIMO framework with a BERT-based encoder using all the input sequences achieved an overall lenient-level AUROC of 0.61. Although the per-formance is suboptimal, representing eligibility criteria into logical and semantically clear tuples can potentially make subsequent translation of these tuples into database queries more reliable.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Novel Question-Answering Framework for Automated Abstract Screening Using Large Language Models 94%
- Active Neural Networks to Detect Mentions of Changes to Medication Treatment in Social Media 94%
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 94%
Similar papers in this journal
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 94%
- Ontology-based expansion of virtual gene panels to improve diagnostic efficiency for rare genetic diseases 93%
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 92%
Similar papers in this journal
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 93%
- The role of natural language processing in cancer care: a systematic scoping review with narrative synthesis 92%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.