systematic evaluation and benchmarking of text summarization methods for biomedical literature: From word-frequency methods to language models
Baumgärtel, F.; Bono, E.; Galou, L.; Keska-Izworska, K.; Walter, S.; Andorfer, P.; Kratochwill, K.; Perco, P.; Ley, M.
Show abstract
The rapid expansion of biomedical literature demands automated summarization tools that can reliably condense research articles into concise, accurate overviews. We benchmarked 62 text summarization methods - ranging from frequency-based and TextRank extractors to modern encoder-decoder models (EDMs) and large language models (LLMs) - on a set of 1,000 biomedical abstracts for which author-generated highlights sections were available as reference summaries. Models were evaluated using a composite suite of metrics covering lexical overlap (ROUGE-1/2/L, BLEU, METEOR), embedding-based semantic similarity (RoBERTa, DeBERTa, all-mpnet-base-v2), and factual consistency (AlignScore). Our results indicate that general-purpose language models (LMs) achieve the highest overall scores across both lexical and semantic metrics, outperforming both reasoning-oriented and domain-specific models. Within the general-purpose group, medium-sized models, typically runnable on a single node, often outperform frontier-scale counterparts, suggesting an optimal balance between model capacity and computational efficiency. Statistical extractive methods lag behind all neural approaches. These findings provide a systematic reference for selecting summarization tools in biomedical research and highlight that broad pretraining remains more effective than narrow domain adaptation for generating high-quality scientific summaries.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Knowledge Graph-based Thought: a knowledge graph enhanced LLMs framework for pan-cancer question answering 95%
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 93%
- ShinyLearner: A containerized benchmarking tool for machine-learning classification of tabular data 92%
Similar papers in this journal
- A Novel Question-Answering Framework for Automated Abstract Screening Using Large Language Models 97%
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 95%
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 95%
Similar papers in this journal
Similar papers in this journal
- Relation extraction between bacteria and biotopes from biomedical texts with attention mechanisms and domain-specific contextual representations 96%
- Optimizing biomedical information retrieval with a keyword frequency-driven Prompt Enhancement Strategy 96%
- SKiM-GPT: Combining Biomedical Literature-Based Discovery with Large Language Model Hypothesis Evaluation 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.