Optimized BERT-based NLP outperforms Zero-Shot Methods for Automated Symptom Detection in Clinical Practice
Diaz Ochoa, J. G.; Layer, N.; Mahr, J.; Mustafa, F. E.; Menzel, C. U.; Mueller-Schilling, M.; Schilling, T.; Illerhaus, G.; Knott, M.; Krohn, A.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSBO_SCPLOWACKGROUNDC_SCPLOWC_ST_ABSLarge Language Nodels (LLMs) have raised broad expectations for clinical use, particularly in the processing of complex medical narratives. However, in practice, more targeted Natural Language Processing (NLP) approaches may offer higher precision and feasibility for symptom extraction from real-world clinical texts. NLP provides promising tools for extracting clinical information from unstructured medical narratives. However, few studies have focused on integrating symptom information from free texts in German, particularly for complex patient groups such as emergency department (ED) patients. The ED setting presents specific challenges: high documentation pressure, heterogeneous language styles, and the need for secure, locally deployable models due to strict data protection regulations. Furthermore, German remains a low-resource language in clinical NLP. MO_SCPLOWETHODSC_SCPLOWWe implemented and compared two models for zero-shot learning--GLiNER and Mistral--and a fine-tuned BERT-based SCAI-BIO/BioGottBERT model for named entity recognition (NER) of symptoms, anatomical terms, and negations in German ED anamnesis texts in an on-premises environment in a hospital. Manual annotations of 150 narratives were used for model validation. The postprocessing steps included confidence-based filtering, negation exclusion, symptom standardization, and integration with structured oncology registry data. All computations were performed on local hospital servers in an on-premises implementation to ensure full data protection compliance. RO_SCPLOWESULTSC_SCPLOWThe fine-tuned SCAI-BIO/BioGottBERT model outperformed both zero-shot approaches, achieving an F1 score of 0.84 for symptom extraction and demonstrating superior performance in negation detection. The validated pipeline enabled systematic extraction of affirmed symptoms from ED-free text, transforming them into structured data. This method allows large-scale analysis of symptom profiles across patient populations and serves as a technical foundation for symptom-based clustering and subgroup analysis. CO_SCPLOWONCLUSIONSC_SCPLOWOur study demonstrates that modern NLP methods can reliably extract clinical symptoms from German ED free text, even under strict data protection constraints and with limited training resources. Fine-tuned models offer a precise and practical solution for integrating unstructured narratives into clinical decision-making. This work lays the methodological foundation for a new way of systematically analyzing large patient cohorts on the basis of free-text data. Beyond symptoms, this approach can be extended to extracting diagnoses, procedures, or other clinically relevant entities. Building upon this framework, we apply network-based clustering methods (in a subsequent study) to identify clinically meaningful patient subgroups and explore sex- and age-specific patterns in symptom expression.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comparing neural language models for medical concept representation and patient trajectory prediction 93%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 93%
- The role of natural language processing in cancer care: a systematic scoping review with narrative synthesis 92%
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 94%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 94%
- Ontology-based expansion of virtual gene panels to improve diagnostic efficiency for rare genetic diseases 92%
Similar papers in this journal
Similar papers in this journal
- Improving dictionary-based named entity recognition with deep learning 94%
- Lifestyle factors in the biomedical literature: An ontology and comprehensive resources for named entity recognition 94%
- CoCoScore: Context-aware co-occurrence scoring for text mining applications using distant supervision 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.