Generative Models for Medical Image Synthesis: A Critical Evaluation of Validation Practices, Clinical Readiness, and Translational Challenges - 02/09/26
Cet article a été publié dans un numéro de la revue, cliquez ici pour y accéder
Abstract |
Objectives |
Medical image synthesis using Generative Adversarial Networks (GANs) has gained increasing attention to address data scarcity, missing modalities, and image quality enhancement in clinical workflows. Despite encouraging technical performance, concerns persist regarding the clinical reliability, generalizability, and safety of GAN-based approaches. This systematic review critically examines experimental and observational studies applying GANs to medical image synthesis, with emphasis on methodological design, evaluation practices, clinical validation, and sources of bias.
Material and Methods |
A PRISMA-guided search of PubMed, IEEE Xplore, Scopus, and Web of Science was conducted to identify studies published between January 2018 and December 2025. Studies were included if they focused on GAN-based medical image synthesis using human subject data, provided quantitative evaluation measures, and reported sufficient methodological detail. Data extraction covered study characteristics, GAN architecture, evaluation metrics, validation strategies, and clinical relevance. Methodological quality and risk of bias were assessed using adapted frameworks.
Results |
Twenty studies met inclusion criteria, spanning multiple imaging modalities (MRI: 40%, CT: 35%, PET: 20%) and anatomical regions. Modality translation (40%) and super-resolution (20%) were the predominant applications, with conditional GAN architectures dominating (75% paired training). While improvements in image similarity metrics such as PSNR and SSIM were frequently reported, external validation (0–5%), multi-center evaluation (0–5%), and clinical reader studies (0–5%) were rare. Common methodological limitations included overreliance on pixel-level metrics, inadequate reporting of failure cases (85%), and dataset biases arising from single-center cohorts (80%).
Conclusion |
Although GAN-based medical image synthesis demonstrates substantial technical promise, current evidence remains insufficient to support routine clinical deployment. Critical barriers include evaluation bias, limited generalizability, hallucination risks, and inadequate safety assessment. Future work should prioritize clinically meaningful endpoints, robust multi-center validation, standardized reporting, and explicit safety protocols. By concentrating on GAN-based methods with demonstrated experimental maturity, this work provides a clinically grounded evidence synthesis intended to inform both researchers and clinicians regarding the responsible use of image synthesis technologies in medical settings.
Le texte complet de cet article est disponible en PDF.Highlights |
• | PRISMA-guided review of 20 GAN-based medical synthesis studies. |
• | Conditional GANs dominate modality translation and super-resolution. |
• | Evaluation bias: over-reliance on pixel-level metrics (PSNR, SSIM). |
• | External validation and clinical reader evaluations remain scarce. |
• | Hallucination risks and dataset bias hinder clinical translation. |
Keywords : Generative adversarial networks, Medical image synthesis, Systematic review, Modality translation, Super-resolution, Clinical applications
Plan
Bienvenue sur EM-consulte, la référence des professionnels de santé.
L’accès au texte intégral de cet article nécessite un abonnement.
Déjà abonné à cette revue ?

