Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

GaelEval: Benchmarking LLM performance for Scottish Gaelic

Devine, Peter, Lamb, William, Alex, Beatrice, Ezeani, Ignatius, Knight, Dawn ORCID: https://orcid.org/0000-0002-4745-6502, Ó Meachair, Micheal J., Rayson, Paul and Wynne, Martin 2026. GaelEval: Benchmarking LLM performance for Scottish Gaelic. Presented at: LREC 2026, Palma, Mallorca, Spain, 11-16 May 2026. Published in: Montejo-Raez, Arturo, Grisot, Cristina and Blochowiak, Joanna eds. Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026 Workshop Proceedings. ELRA Language Resources Association, pp. 73-85.

[thumbnail of 2026.llms4ssh-1.0_.pdf]
Preview
PDF - Published Version
Available under License Creative Commons Attribution Non-commercial.

Download (304kB) | Preview

Abstract

Multilingual large language models (LLMs) often exhibit emergent ‘shadow’ capabilities in languages without official support, yet their performance on these languages remains uneven and under-measured. This is particularly acute for morphosyntactically rich minority languages such as Scottish Gaelic, where translation benchmarks fail to capture structural competence. We introduce GaelEval, the first multi-dimensional benchmark for Gaelic, comprising: (i) an expert-authored morphosyntactic MCQA task; (ii) a culturally grounded translation benchmark and (iii) a large-scale cultural knowledge Q&Atask. Evaluating 19LLMsagainstafluent-speakerhumanbaseline(n = 30), wefindthatGem ini 3 Pro Preview achieves 83.3% accuracy on the linguistic task, surpassing the human baseline (78.1%). Proprietary models consistently outperform open-weight systems, and in-language (Gaelic) prompting yields a small but stable advantage (+2.4%). On the cultural task, leading models exceed 90% accuracy, though most systems perform worse under Gaelic prompting and absolute scores are inflated relative to the manual benchmark. Overall, GaelEval reveals that frontier models achieve above-human performance on several dimensions of Gaelic grammar, demonstrates the effect of Gaelic prompting and showsa consistent performance gap favouring proprietary over open-weight models.

Item Type: Conference or Workshop Item - published (Paper)
Date Type: Publication
Status: Published
Schools: Schools > English, Communication and Philosophy
Publisher: ELRA Language Resources Association
ISBN: 9782493814852
Date of First Compliant Deposit: 13 April 2026
Last Modified: 05 Aug 2026 01:57
URI: https://orca.cardiff.ac.uk/id/eprint/186319

Actions (repository staff only)

Edit Item Edit Item

Downloads

Downloads per month over past year

View more statistics