S'abonner

Empowering front-line physicians with AI: Evaluating large language models in everyday ENT care - 19/02/26

Doi : 10.1016/j.ajem.2026.01.029 
Sholem Hack a, , 1 , Habib G. Zalzal b, Rebecca Attal a, Armin Farzad c, Lilia Ann Crew d, Idit Tessler e, f, h, Talya Frankel g, Ben Gvili e, f, h, Shaked Shivatzki e, f, h, Amit Wolfovitz e, f, h, Noa Rozendorn e, f, h
a City St. George's University London School of Medicine, Program Delivered by University of Nicosia at the Chaim Sheba Medical Center, Ramat Gan, Israel 
b Division of Otolaryngology-Head and Neck Surgery Children's National Hospital Washington District of Columbia, USA 
c School of Medicine, Cit St.George's University of London, Crammer Terrace, London SW17 0RE, United Kingdom 
d Touro Collge of Osteopathic Medicine - Montanna, United Sates 
e Department of Otolaryngology, Head and Neck Surgery, Sheba Medical Center, Tel Hashomer, Israel 
f Gray Faculty of Medical & Health Sciences, Tel-Aviv University, Israel 
g Clalit Health Services, Tel aviv, Israel 
h Faculty of Medicine, Tel Aviv University, Tel Aviv, Israel 

Corresponding author at: City St. Georges University London School of Medicine, Program Delivered by University of Nicosia at the Chaim Sheba Medical Center, Derech Sheba 2, Tel Hashomer, 5621, Israel City St. Georges University London School of Medicine Program Delivered by University of Nicosia at the Chaim Sheba Medical Center Derech Sheba 2 Tel Hashomer 5621 Israel

Abstract

Purpose

Artificial intelligence systems known as large language models are being evaluated for clinical decision support, yet their role in emergency and primary care remains limited. Physicians in these settings often encounter ear, nose, and throat conditions where diagnostic uncertainty, unnecessary testing, and inappropriate referrals contribute to patient risk and healthcare inefficiency. This study compared the performance of advanced large language models with physicians in diagnosis, management, and referral across common and high-acuity otolaryngologic scenarios.

Methods

Twelve clinical vignettes representing routine and urgent presentations were developed and validated by otolaryngologists. One hundred practicing physicians in family medicine and emergency medicine, including residents and attending physicians, completed all vignettes by providing a diagnosis, management plan, and referral decision. Four large language models (Gemini-2.0, ChatGPT-4.0, ChatGPT-5, and OpenEvidence) were tested using identical prompts. Model outputs were anonymized, randomized, and rated by a blinded expert panel using the Quality Analysis of Medical Artificial Intelligence tool, which assesses accuracy, clarity, completeness, sourcing, relevance, and usefulness.

Results

Physicians achieved mean diagnostic accuracy of 91.6% and management accuracy of 87.9%. In non-urgent cases, 30.4% of responses represented inappropriate referral. Only half recognized the need for urgent referral in a cerebrospinal fluid leak scenario. Large language models demonstrated comparable diagnostic and management accuracy with higher referral appropriateness.

Conclusions

Large language models showed consistent, guideline-concordant reasoning in simulated emergency and primary-care otolaryngology cases. Their potential lies in supporting, not replacing, clinical judgment through responsible integration and real-world validation.

Le texte complet de cet article est disponible en PDF.

Keywords : Artificial intelligence, Large language models, Otolaryngology, Clinical decision support, Referral patterns, Diagnostic accuracy


Plan


© 2026  Elsevier Inc. Tous droits réservés.
Ajouter à ma bibliothèque Retirer de ma bibliothèque Imprimer
Export

    Export citations

  • Fichier

  • Contenu

Vol 102

P. 90-97 - avril 2026 Retour au numéro
Article précédent Article précédent
  • A CT-based multimodal fusion model for predicting outcomes in blunt chest trauma: A multicenter study
  • Tingting Zhao, Dong Li, Mengshan Wu, Chenyuan Zhang, Xiaoyuan Qu, Xin Tian, Yixi Zhang, Chunlin Song, Xiaoran Wang, Xianghong Meng, Zhi Wang
| Article suivant Article suivant
  • Remote real-time ambulatory ECG monitoring of the longest RR interval and corresponding arrhythmias during altitude ascent
  • Li-Hong Zheng, Yan Wang, Xue-Wen Huang, Miao Li, Hai-Ying Xian, Chun-Xia Guo, Zi-Yang He, Lin Ma

Bienvenue sur EM-consulte, la référence des professionnels de santé.
L’accès au texte intégral de cet article nécessite un abonnement.

Déjà abonné à cette revue ?

Elsevier s'engage à rendre ses eBooks accessibles et à se conformer aux lois applicables. Compte tenu de notre vaste bibliothèque de titres, il existe des cas où rendre un livre électronique entièrement accessible présente des défis uniques et l'inclusion de fonctionnalités complètes pourrait transformer sa nature au point de ne plus servir son objectif principal ou d'entraîner un fardeau disproportionné pour l'éditeur. Par conséquent, l'accessibilité de cet eBook peut être limitée. Voir plus

Mon compte


Plateformes Elsevier Masson

Déclaration CNIL

EM-CONSULTE.COM est déclaré à la CNIL, déclaration n° 1286925.

En application de la loi nº78-17 du 6 janvier 1978 relative à l'informatique, aux fichiers et aux libertés, vous disposez des droits d'opposition (art.26 de la loi), d'accès (art.34 à 38 de la loi), et de rectification (art.36 de la loi) des données vous concernant. Ainsi, vous pouvez exiger que soient rectifiées, complétées, clarifiées, mises à jour ou effacées les informations vous concernant qui sont inexactes, incomplètes, équivoques, périmées ou dont la collecte ou l'utilisation ou la conservation est interdite.
Les informations personnelles concernant les visiteurs de notre site, y compris leur identité, sont confidentielles.
Le responsable du site s'engage sur l'honneur à respecter les conditions légales de confidentialité applicables en France et à ne pas divulguer ces informations à des tiers.


Tout le contenu de ce site: Copyright © 2026 Elsevier, ses concédants de licence et ses contributeurs. Tout les droits sont réservés, y compris ceux relatifs à l'exploration de textes et de données, a la formation en IA et aux technologies similaires. Pour tout contenu en libre accès, les conditions de licence Creative Commons s'appliquent.