Systematic assessment of text summarization methods for biomedical literature from frequency methods to language models
Systematic assessment of text summarization methods for biomedical literature from frequency methods to language models

Systematic assessment of text summarization methods for biomedical literature from frequency methods to language models

iScience. 2026 Sep 3;29(9):117438. doi: 10.1016/j.isci.2026.117438. eCollection 2026 Sep 18.

ABSTRACT

The rapid expansion of biomedical literature demands automated summarization tools that reliably condense research articles into concise, accurate summaries. We benchmarked 62 summarization methods, ranging from frequency-based and TextRank extractors to encoder-decoder models (EDMs) and large language models (LLMs), on 1,000 biomedical abstracts from 20 journals across ScienceDirect and Cell Press, using author-written highlights as reference summaries. Models were evaluated with a composite suite of lexical, semantic, and factual metrics, including ROUGE, BLEU, METEOR, embedding-based similarity, and factuality scores. General-purpose models (e.g., Mistral, GPT, and Llama) achieved the highest overall performance across lexical and semantic dimensions, outperforming reasoning-oriented (e.g., DeepSeek and Magistral) and domain-specific (e.g., BioGPT and BioMistral) models. Notably, medium-sized models outperformed large-scale models, suggesting an optimal balance between model capacity and efficiency, while classical extractive methods lagged behind neural approaches. These findings provide a systematic reference for selecting biomedical summarization tools and highlight that broad pretraining outperforms narrow domain adaptation.

PMID:42733895 | PMC:PMC13571750 | DOI:10.1016/j.isci.2026.117438