S'abonner

Auto Machine Learning for Diabetic Retinopathy Screening: A Head-to-Head Multiplatform Comparison Against Human Graders and IDx-DR - 07/07/26

Doi : 10.1016/j.ajo.2026.04.030 
Tomasz Krzywicki a, f, Ceren Durmaz Engin b, c, Andrzej Grzybowski d, e,
a From the Faculty of Mathematics and Computer Science (T.K.), University of Warmia and Mazury, 10-719 Olsztyn, Poland 
b Department of Ophthalmology (C.D.E.), Democracy University, Izmir, Turkey 
c Izmir Health Technologies Development and Accelerator (BioIzmir) (C.D.E.), Dokuz Eylul University, Izmir, Turkey 
d Department of Ophthalmology (A.G.), University of Warmia and Mazury, 10-719 Olsztyn, Poland 
e Institute for Research in Ophthalmology (A.G.), Foundation for Ophthalmology Development, 61-553 Poznan, Poland 
f Faculty of Electronics, Telecommunications and Informatics, Gdansk University of Technology, Poland 

Inquiries to Andrzej Grzybowski, Department of Ophthalmology, University of Warmia and Mazury, Olsztyn, Poland Department of Ophthalmology University of Warmia and Mazury Olsztyn Poland

Highlights

We performed a head-to-head benchmark of AutoML platforms for DR screening.
Amazon SageMaker Canvas and AutoGluon achieved the highest AUCs for RDR and STDR.
Model performance varied markedly across platforms and probability thresholds.
Several AutoML tools showed moderate agreement with the FDA-approved IDx-DR system.
External validation and threshold calibration are essential for real-world deployment.

Le texte complet de cet article est disponible en PDF.

Résumé

Purpose

To benchmark multiple automated machine learning (AutoML) platforms for diabetic retinopathy (DR) screening from fundus photographs using a unified training and evaluationframework, with human consensus grading and an FDA-approved autonomous system (IDx-DR) as reference standards.

Design

Retrospective, diagnostic performance and benchmarking study.

Methods

Image classifiers were trained on large public datasets labeled according to the International Clinical Diabetic Retinopathy (ICDR) scale (APTOS, n = 5590; DDR, n = 12,524; EyePACS, n = 31,557) after automated image-quality filtering. Performance was evaluated on an independent, institutionally collected patient-level test cohort (n = 726) using the highest DR grade across all images for patient-level classification. The evaluated platforms included Google Vertex AI, Amazon Rekognition, Amazon SageMaker Canvas, AutoGluon, AutoKeras, and Apple CreateML. Models were assessed for 3 screening endpoints—any DR, referable DR (RDR), and sight-threatening DR (STDR)—across probability thresholds of 20%, 50%, and 70%. The primary endpoints were the Area Under the Receiver Operating Characteristic Curve (AUC) and sensitivity for RDR at a 50% decision threshold. Secondary endpoints included specificity, positive predictive value, negative predictive value, accuracy, and F1-score with 95% confidence intervals. Pairwise comparisons were performed using bootstrap testing (n = 1000) for AUC differences and McNemar’s test (at a 50% threshold) for binary outcomes, both subjects to Bonferroni correction ( P < .0033). Grad-CAM was applied to locally deployable convolutional neural network–based models.

Results

Using human consensus grading as the reference standard, Amazon SageMaker Canvas and AutoGluon demonstrated the strongest overall discrimination, achieving AUC values up to0.96 for STDR and 0.93 to 0.94 for RDR. At the 50% decision threshold, Canvas showed the most balanced performance for RDR (sensitivity 88.3%, specificity 85.5%, accuracy 86.0%), whereasAutoGluon favored sensitivity (any DR sensitivity 95.9%) at the expense of specificity. Vertex AI showed consistently weaker and unstable performance (any DR AUC 0.58; RDR AUC 0.38). Relative to IDx-DR, Amazon Rekognition and Canvas showed the highest agreement, particularly for STDR (AUC up to 0.88-0.90; κ up to ∼0.56). Agreement with human graders was generallylow to moderate (κ ≈ 0.3-0.6) and increased at higher probability thresholds.

Conclusions

AutoML platforms can achieve clinically meaningful performance for DR screening. Differences across tools and thresholds reflect their adaptability to diverse clinical settings, underscoring the importance of external validation and threshold calibration.

Le texte complet de cet article est disponible en PDF.

Plan


© 2026  Elsevier Inc. Tous droits réservés.
Ajouter à ma bibliothèque Retirer de ma bibliothèque Imprimer
Export

    Export citations

  • Fichier

  • Contenu

Vol 288

P. 267-283 - août 2026 Retour au numéro
Article précédent Article précédent
  • Genotype–Phenotype Correlations in RPGRIP1-Associated Retinal Dystrophy in a Nationwide Japanese Cohort
  • Kei Mizobuchi, Taiga Inooka, Takuya Aoki, Hazuki Anzai, Kaoruko Torii, Kazuki Hashimoto, Akiko Suga, Ryo Ando, Miki Hiraoka, Taro Kominami, Shigeru Sato, Motokazu Tsujikawa, Kohji Nishida, Yusuke Murakami, Toru Nakazawa, Akiko Maeda, Kazuki Kuniyoshi, Yasuhiro Ikeda, Hiroyuki Kondo, Mineo Kondo, Koji M. Nishiguchi, Akira Murakami, Maki Fukami, Sachiko Nishina, Takeshi Iwata, Hirotomo Saitsu, Kazushige Tsunoda, Shinji Ueno, Yoshihiro Hotta, Tadashi Nakano, Takaaki Hayashi
| Article suivant Article suivant
  • Tubercular Retinitis: Clinical Spectrum and Multimodal Imaging Features of an Insufficiently Characterized Entity
  • Atul Arora, Manu Sharma, Pietro Gentile, Aniruddha Agarwal, Mohit Dogra, Aman Sharma, Vishali Gupta

Bienvenue sur EM-consulte, la référence des professionnels de santé.
L’accès au texte intégral de cet article nécessite un abonnement.

Déjà abonné à cette revue ?

Elsevier s'engage à rendre ses eBooks accessibles et à se conformer aux lois applicables. Compte tenu de notre vaste bibliothèque de titres, il existe des cas où rendre un livre électronique entièrement accessible présente des défis uniques et l'inclusion de fonctionnalités complètes pourrait transformer sa nature au point de ne plus servir son objectif principal ou d'entraîner un fardeau disproportionné pour l'éditeur. Par conséquent, l'accessibilité de cet eBook peut être limitée. Voir plus

Mon compte


Plateformes Elsevier Masson

Déclaration CNIL

EM-CONSULTE.COM est déclaré à la CNIL, déclaration n° 1286925.

En application de la loi nº78-17 du 6 janvier 1978 relative à l'informatique, aux fichiers et aux libertés, vous disposez des droits d'opposition (art.26 de la loi), d'accès (art.34 à 38 de la loi), et de rectification (art.36 de la loi) des données vous concernant. Ainsi, vous pouvez exiger que soient rectifiées, complétées, clarifiées, mises à jour ou effacées les informations vous concernant qui sont inexactes, incomplètes, équivoques, périmées ou dont la collecte ou l'utilisation ou la conservation est interdite.
Les informations personnelles concernant les visiteurs de notre site, y compris leur identité, sont confidentielles.
Le responsable du site s'engage sur l'honneur à respecter les conditions légales de confidentialité applicables en France et à ne pas divulguer ces informations à des tiers.


Tout le contenu de ce site: Copyright © 2026 Elsevier, ses concédants de licence et ses contributeurs. Tout les droits sont réservés, y compris ceux relatifs à l'exploration de textes et de données, a la formation en IA et aux technologies similaires. Pour tout contenu en libre accès, les conditions de licence Creative Commons s'appliquent.