Baskaran, Ravanth, Sirikonda, Sai, Singh, Aditya, Srinivasan, Sripradha, Leveridge, Becky, Casals-Farre, Octavi, Kaur, Harmeena, Manivannan, Susruta, Elangovan, Kabilan, Wan Ning Quek, Chrystie, Fukutsu, Kanae, Cave, Judith, Somani, Khaskar Kumar, Shu Wei Ting, Daniel and Hassoulas, Athanasios ORCID: https://orcid.org/0000-0002-1029-1847
2026.
The temporal changes in GPT-4 performance on UKMLA practice questions: educational and clinical implications.
Frontiers in Medicine
13
, 1888765.
10.3389/fmed.2026.1888765
|
Preview |
PDF
- Published Version
Available under License Creative Commons Attribution. Download (502kB) | Preview |
Abstract
Background: Generative Artificial Intelligence (GAI) models, such as GPT-4, have been extensively studied for their integration into medical practice and education. GPT-4 has demonstrated excellent performance on medical licensing examinations, including the United Kingdom Medical Licensing Assessment (UKMLA). However, the field lacks longitudinal data on whether such performance is stable or varies over time. Theoretically, iterative improvements updated by the vendor should enhance performance, but empirical evidence of such longitudinal trends remains limited. Given that GPT-4 undergoes periodic vendor-side updates, we aimed to analyse the categorical, time-spaced performance of GPT-4 on the UKMLA to characterise how its performance changes over time in a medical context.Methods: Two publicly available UKMLA papers were fed into GPT-4 at two different time points, June 2023 and December 2024. 191 questions were provided with and without multiple-choice options to assess GPT-4’s clinical competence. McNemar’s test was performed to evaluate changes in GPT-4’s performance over time, comparing domain-specific questions.Results: GPT-4’s accuracy improved noticeably between the two rounds (MCQ: 88.0 to 93.7%, p = 0.027; non-MCQ: 68.1 to 81.7%, p = 1.00). Single-step accuracy rose from 73.1 to 82.3%, and multi-step from 57.4 to 80.3% without MCQ. GPT-4 showed improved accuracy from Round 1 to Round 2 for both single-step and multi-step questions, with MCQ-prompted responses consistently outperforming non-MCQ responses (up to 95.1% accuracy for multi-step MCQ questions in Round 2). GPT-4’s performance improved across all question categories from round one to round two, most notably in management questions without MCQ options (+23.30%), though these differences were not statistically significant.Discussion and conclusion: GPT-4’s performance on the UKMLA improved significantly over 18 months, suggesting that iterative vendor-side model updates enhance clinical reasoning capabilities. These findings indicate that GPT-4 may serve as a supplementary educational tool for medical students and clinicians; however, the underlying drivers of performance changes remain opaque, and such tools should be deployed with structured oversight to prevent overreliance.
| Item Type: | Article |
|---|---|
| Date Type: | Published Online |
| Status: | Published |
| Schools: | Schools > Medicine |
| Subjects: | L Education > LB Theory and practice of education > LB2300 Higher Education R Medicine > R Medicine (General) T Technology > T Technology (General) |
| Publisher: | Frontiers Media |
| ISSN: | 2296-858X |
| Date of First Compliant Deposit: | 23 July 2026 |
| Date of Acceptance: | 6 July 2026 |
| Last Modified: | 24 Jul 2026 08:24 |
| URI: | https://orca.cardiff.ac.uk/id/eprint/188432 |
Actions (repository staff only)
![]() |
Edit Item |





Dimensions
Dimensions