Espona, José, Roig, Elena, Garcia, Marc, Durán-Sindreu, Fernando, Abella, Francesc, Dummer, Paul M.H. ORCID: https://orcid.org/0000-0002-0726-7467, González, José Antonio, Elmsmari, Firas, Pineda, Kenneth, Figueras, Oscar and Roig, Miguel
2026.
Curated retrieval-augmented generation for dental traumatology.
Journal of Dentistry
175
, 106947.
10.1016/j.jdent.2026.106947
|
Preview |
PDF
- Accepted Post-Print Version
Available under License Creative Commons Attribution. Download (3MB) | Preview |
Abstract
Objectives: To validate DT-RAG, a curated retrieval-augmented generation system for dental traumatology decision support, against eight commercial large language models. Methods: A knowledge base of 250 curated text units (“chunks”) from five authoritative sources (IADT 2020, ESE 2021, Krastl 2021, AAE 2013, Cochrane) was built and deployed with Gemini 2.5 Flash as base model. In Study 1, 99 binary clinical questions were submitted in three runs to DT-RAG and eight LLMs; modal accuracy was compared by McNemar exact tests with Holm correction. In Study 2, seven blinded specialists scored DT-RAG against three frontier LLMs on ten clinical scenarios using a 92-point rubric; differences were estimated by linear mixed-effects regression. Results: DT-RAG achieved 96.0% modal accuracy (95% CI 90.1–98.4), significantly exceeding every commercial LLM (best comparator GPT-5.5 87.9%; paired difference +8.1 pp, 95% CI +2.9 to +14.9; p_Holm = 0.013). The curated knowledge base elevated the base model from 49.5% to 98.0% valid rationale rate, eliminating confabulations in this evaluation (0 vs 21). In Study 2, DT-RAG achieved the highest mean score (82.1/92; 89.3%), significantly exceeding Claude Opus 4.5 (72.6; paired difference +9.6 points, exact Wilcoxon p = 0.016), Gemini 2.5 Pro (60.3) and GPT-4.1 (50.4); all seven evaluators ranked DT-RAG first (Kendall’s W = 0.97). Conclusions: DT-RAG, a curated retrieval-augmented configuration, outperformed frontier-tier general-purpose LLMs on this dental traumatology benchmark, with no confabulated rationales observed among the responses assessed. This approach may be applicable to other well-defined clinical domains with authoritative guidelines. Clinical significance: Curated retrieval augmentation produced accurate, source-traceable and reproducible responses in dental traumatology under benchmark and simulated-scenario conditions, making every error auditable against its source. Clinical safety requires prospective evaluation.
| Item Type: | Article |
|---|---|
| Date Type: | Publication |
| Status: | Published |
| Schools: | Schools > Dentistry |
| Additional Information: | RRS policy applied |
| Publisher: | Elsevier |
| ISSN: | 0300-5712 |
| Date of First Compliant Deposit: | 12 August 2026 |
| Date of Acceptance: | 3 August 2026 |
| Last Modified: | 12 Aug 2026 08:45 |
| URI: | https://orca.cardiff.ac.uk/id/eprint/188887 |
Actions (repository staff only)
![]() |
Edit Item |





Dimensions
Dimensions