Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Can GenAI be trained to mimic human markers of extended written assignments in Higher Education?

Kay, William, Rutherford, Stephen ORCID: https://orcid.org/0000-0002-5572-8854, Smith, David, Shore, Andrew ORCID: https://orcid.org/0000-0001-7115-2050 and Francis, Nigel ORCID: https://orcid.org/0000-0002-4706-4795 2026. Can GenAI be trained to mimic human markers of extended written assignments in Higher Education? Assessment & Evaluation in Higher Education 10.1080/02602938.2026.2710688
Item availability restricted.

[thumbnail of Kay et al_GenAI Marking_COMPOSITE_070826.pdf] PDF - Accepted Post-Print Version
Restricted to Repository staff only

Download (3MB)
[thumbnail of Provisional file] PDF (Provisional file) - Accepted Post-Print Version
Download (17kB)

Abstract

The rapid rise of Generative Artificial Intelligence (GenAI) has promoted interest in its use for marking student work in Higher Education. Harnessing the pattern-recognition capabilities of Large Language Models (LLMs) may potentially facilitate objective grading of students’ work at scale. This study investigated whether LLMs could mimic human marking of extended written work sufficiently to function as a formative quality benchmarking tool for students. Two versions of a popular GenAI platform, ChatGPT, were used to mark 50 undergraduate bioscience essays against seven assessment criteria under four different prompting conditions. LLM-assigned marks were compared to human marks in both their mean scores and mark variability using distributional modelling. While overall composite marks were often broadly similar between humans and LLMs, substantial discrepances emerged at the individual criterion and essay levels. LLMs generally awarded higher marks than humans, but with significantly reduced variability, resulting in systematic compression of marks toward the centre of the distribution. Lower-scoring essays tended to receive inflated marks, whereas higher-scoring essays received reduced marks relative to human assessment. These findings suggest that current LLMs do not reliably reproduce human judgement in the marking of extended written work, with important implications for assessment practice and GenAI-assisted benchmarking of marks.

Item Type: Article
Status: In Press
Schools: Schools > Biosciences
Subjects: L Education > L Education (General)
L Education > LB Theory and practice of education > LB2300 Higher Education
Additional Information: DOI not yet active 11/08/2026
Publisher: Taylor and Francis Group
ISSN: 0260-2938
Date of First Compliant Deposit: 11 August 2026
Date of Acceptance: 7 August 2026
Last Modified: 11 Aug 2026 08:00
URI: https://orca.cardiff.ac.uk/id/eprint/188794

Actions (repository staff only)

Edit Item Edit Item

Downloads

Downloads per month over past year

View more statistics