Direkt zum Inhalt

Walter, Nike ; Amanatullah, Derek F. ; Debarbieux, Laurent ; Doub, James B. ; Ferry, Tristan ; Gross, Justus ; Międzybrodzki, Ryszard ; Mirzaei, Mohammadali Khan ; Deng, Li ; Rācenis, Kārlis ; Suh, Gina A. ; Que, Yok-Ai ; Górski, Andrzej ; Rupp, Markus

Performance of large language models as a source of clinical information on bacteriophage therapy

Artikel

Walter, Nike , Amanatullah, Derek F., Debarbieux, Laurent, Doub, James B., Ferry, Tristan, Gross, Justus , Międzybrodzki, Ryszard, Mirzaei, Mohammadali Khan, Deng, Li, Rācenis, Kārlis, Suh, Gina A., Que, Yok-Ai, Górski, Andrzej und Rupp, Markus (2026) Performance of large language models as a source of clinical information on bacteriophage therapy. npj Viruses 4, S. 41.

DOI zum Zitieren dieses Dokuments: 10.5283/epub.80717


Zusammenfassung

Bacteriophage therapy is re-emerging as a potential strategy to address antimicrobial resistance, but standardized patient education materials are limited. Large language models (LLMs) are increasingly used for patient-facing medical information. The quality of LLM-generated responses to 20 patient-relevant questions was evaluated by 12 clinicians and research experts in bacteriophage therapy ...

Bacteriophage therapy is re-emerging as a potential strategy to address antimicrobial resistance, but standardized patient education materials are limited. Large language models (LLMs) are increasingly used for patient-facing medical information. The quality of LLM-generated responses to 20 patient-relevant questions was evaluated by 12 clinicians and research experts in bacteriophage therapy independently rated each response for accuracy, completeness, clarity, and tone/empathy using 5-point Likert scales. Expert suggestions for improvement were recorded. A total of 960 ratings were analyzed. Adjusted mean scores ranged from 3.36 to 3.96 across domains, indicating generally favorable evaluations for all models. Significant differences among LLMs were observed for completeness and tone/empathy (Holm-adjusted p = 0.042 for both), but not for accuracy or clarity. Differences were small in magnitude (Cohen’s d = 0.12–0.29). Claude scored significantly lower than the other models for completeness and tone/empathy, while Perplexity achieved the highest completeness scores. Experts recommended improvements for 34–40% of responses; wrong information was given in 20%. The best responses were revised into an expert-informed patient guide provided as Supplementary Material, presenting a hybrid model in which LLMs generate draft patient information that is subsequently refined by clinical experts, particularly in rapidly evolving therapeutic domains lacking standardized educational resources.



Beteiligte Einrichtungen


Details

DokumentenartArtikel
Titel eines Journals oder einer Zeitschriftnpj Viruses
VerlagSpringer Nature
Open Access ArtDEAL (Springer Gold)
Band4
SeitenbereichS. 41
Datum31 August 2026
Veröffentlichungsdatum14 Sep 2026 05:52
InstitutionenMedizin > Lehrstuhl für Unfallchirurgie
Identifikationsnummer
WertTyp
10.1038/s44298-026-00224-2DOI
Dewey-Dezimal-Klassifikation600 Technik, Medizin, angewandte Wissenschaften > 610 Medizin
StatusVeröffentlicht
BegutachtetJa, diese Version wurde begutachtet
An der Universität Regensburg entstandenZum Teil
URN der UB Regensburgurn:nbn:de:bvb:355-epub-807178
Dokumenten-ID80717

Bibliographische Daten exportieren

Nur für Besitzer und Autoren: Kontrollseite des Eintrags

nach oben