Suscribirse

Are answers obtained from artificial intelligence models for information purposes repeatable? - 04/10/25

Doi : 10.1016/j.ortho.2025.101071 
Yasemin Tunca 1, Volkan Kaplan 2, Murat Tunca 1, ⁎
1 Department of Orthodontics, Faculty of Dentistry, Kutahya Health Sciences University, Kutahya, Turkey 
2 Department of Oral and Maxillofacial Surgery, Faculty of Dentistry, Tekirdag Namık Kemal University, Tekirdag, Turkey 

⁎ Murat Tunca, Department of Orthodontics, Faculty of Dentistry, Kutahya Health Sciences University, Kutahya, Turkey. Department of Orthodontics, Faculty of Dentistry, Kutahya Health Sciences University Kutahya Turkey

Highlights

•
The repeatability of orthodontic responses generated by large language models (LLMs) over time is of significant importance.
•
While ChatGPT-3.5 demonstrated the highest level of consistency, the Gemini models exhibited moderate repeatability.
•
The temporal variability in model performance underscores the need for caution when utilizing AI tools in patient communication.

El texto completo de este artículo está disponible en PDF.

Summary

Introduction

The objective of this study was to assess the repeatability of orthodontic responses generated by multiple large language models across repeated time points.

Methods

This experimental study assessed the answers provided by ChatGPT-3.5, ChatGPT-4.0, Gemini, and Gemini-Advanced to 40 frequently asked orthodontic questions. Each model was prompted with the same questions at three time points (T0: day 0, T1: day 7, and T2: day 14). Two blinded orthodontic experts independently evaluated responses using a 3-point accuracy scale. Cohen's Kappa and ICC were applied to assess inter-rater agreement and repeatability, respectively. In addition, Friedman test with Bonferroni post-hoc analysis and Spearman correlation were used for temporal comparisons.

Results

Cohen's Kappa values between raters ranged from 0.624 to 0.749, indicating substantial inter-rater agreement. ICC values for repeatability ranged from 0.666 (Gemini) to 0.960 (ChatGPT-3.5). Friedman test results revealed significant differences in model accuracy at T0 and T2 ( P < 0.001). Post-hoc analysis showed ChatGPT-3.5 differed significantly from Gemini and Gemini Advanced. Spearman correlations between time points were positive but weak (ρ = 0.284 to 0.383, P < 0.001).

Conclusions

The study revealed statistically significant differences in repeatability among AI models. Despite high accuracy, some models exhibited limited consistency over time. These findings underscore the importance of evaluating both accuracy and temporal stability when integrating AI systems into clinical orthodontic communication.

El texto completo de este artículo está disponible en PDF.

Keywords : Large language models, Acquiring knowledge, Repeatable


Esquema


© 2025  CEO. Publicado por Elsevier Masson SAS. Todos los derechos reservados.
Añadir a mi biblioteca Eliminar de mi biblioteca Imprimir
Exportación

    Exportación citas

  • Fichero

  • Contenido

Vol 24 - N° 1

Artículo 101071- mars 2026 Regresar al número
Artículo precedente Artículo precedente
  • Predicting treatment pathways in Class II malocclusion patients using machine learning: A comparative study of four algorithms for classifying camouflage, growth modulation, and surgical decisions
  • Mukesh Kumar, Sumit Kumar, Malvika Agarwal, Ekta Yadav, Sougandhika Gandi
| Artículo siguiente Artículo siguiente
  • Comparison of differential and conventional rapid maxillary expansion on upper airway dimensions in children with bilateral complete cleft lip and palate: A CBCT-based secondary analysis of a clinical trial
  • Denise Caffer, Daniela Garib, Carolina Faber, Alexandre Meireles Borba, Luiz Volpato, Rita de Cássia Moura Carvalho Lauris, Araci Malagodi de Almeida, Rafael Guerra Lund

Bienvenido a EM-consulte, la referencia de los profesionales de la salud.
El acceso al texto completo de este artículo requiere una suscripción.

¿Ya suscrito a @@106933@@ revista ?

@@150455@@ Voir plus

Mi cuenta


Declaración CNIL

EM-CONSULTE.COM se declara a la CNIL, la declaración N º 1286925.

En virtud de la Ley N º 78-17 del 6 de enero de 1978, relativa a las computadoras, archivos y libertades, usted tiene el derecho de oposición (art.26 de la ley), el acceso (art.34 a 38 Ley), y correcta (artículo 36 de la ley) los datos que le conciernen. Por lo tanto, usted puede pedir que se corrija, complementado, clarificado, actualizado o suprimido información sobre usted que son inexactos, incompletos, engañosos, obsoletos o cuya recogida o de conservación o uso está prohibido.
La información personal sobre los visitantes de nuestro sitio, incluyendo su identidad, son confidenciales.
El jefe del sitio en el honor se compromete a respetar la confidencialidad de los requisitos legales aplicables en Francia y no de revelar dicha información a terceros.


Todo el contenido en este sitio: Copyright © 2026 Elsevier, sus licenciantes y colaboradores. Se reservan todos los derechos, incluidos los de minería de texto y datos, entrenamiento de IA y tecnologías similares. Para todo el contenido de acceso abierto, se aplican los términos de licencia de Creative Commons.