Eur J Dent Educ. 2026 Sep 3. doi: 10.1111/eje.70295. Online ahead of print.
ABSTRACT
AIM: Artificial intelligence (AI) has become an integral part of dental education and clinical practice. While several studies have assessed the performance of ChatGPT in medical exams, comparative analyses of various large language models (LLMs) in dentistry remain scarce. This study aimed to evaluate and compare the performance of five prominent AI models in answering clinical questions from Turkey’s dental specialty examination (DUS).
MATERIALS AND METHODS: A total of 1026 clinical science questions from 13 DUS examinations (2012-2021) were used to test the accuracy of five LLMs: ChatGPT-5, Gemini 3.5 Flash, Microsoft Copilot, DeepSeek-R1, and Grok 4. Questions were categorized by specialty and format (information-based vs. case-based). Each question was entered once, and responses were scored against official answer keys. Statistical analyses included chi-square and Z tests with Bonferroni correction.
RESULTS: Accuracy scores ranged from 79.2% to 90.6%. Gemini 3.5 Flash and ChatGPT-5 achieved the highest overall accuracy rates (90.6% and 88.2%, respectively). Microsoft Copilot, DeepSeek-R1, and Grok 4 demonstrated significantly lower performance (p < 0.001). Accuracy rates varied across examination years and dental specialties. The highest performance was observed in Oral and Maxillofacial Surgery, Paediatric Dentistry, Periodontology, and Restorative Dentistry, whereas Prosthetic Dentistry, Orthodontics, and Endodontics yielded comparatively lower accuracy rates. Significant differences among AI models were observed for both information-based (p < 0.001) and case-based questions (p = 0.029).
CONCLUSION: While LLMs show strong potential in supporting dental education through standardised exams, their performance varies by model and question type. Gemini 3.5 Flash and ChatGPT-5 achieved the highest and most consistent accuracy rates, whereas Microsoft Copilot, DeepSeek-R1, and Grok 4 showed lower performance. Advanced LLMs may represent useful supplementary tools for dental education and DUS preparation, although further improvements are needed to enhance reliability across different dental disciplines.
PMID:42691237 | DOI:10.1111/eje.70295