Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Knowing the facts but choosing the shortcut: understanding how large language models compare entities

Lehmann, Hans Hergen, Lee, Jae Hee, Schockaert, Steven ORCID: https://orcid.org/0000-0002-9256-2881 and Wermter, Stefan 2026. Knowing the facts but choosing the shortcut: understanding how large language models compare entities. Presented at: 19th Conference of the European Chapter of the Association for Computational Linguistics, Rabat, Morocco, 24-29 March 2026. Published in: Demberg, Vera, Inui, Kentaro and Marquez, Lluis eds. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics. , vol.1 Association for Computational Linguistics, pp. 4788-4821. 10.18653/v1/2026.eacl-long.222

[thumbnail of 2026.eacl-long.222.pdf] PDF - Published Version
Available under License Creative Commons Attribution.

Download (1MB)

Abstract

Large Language Models (LLMs) are increasingly used for knowledge-based reasoning tasks, yet understanding when they rely on genuine knowledge versus superficial heuristics remains challenging. We investigate this question through entity comparison tasks by asking models to compare entities along numerical attributes (e.g., “Which river is longer, the Danube or the Nile?”), which offer clear ground truth for systematic analysis. Despite having sufficient numerical knowledge to answer correctly, LLMs frequently make predictions which contradict this knowledge. We identify three heuristic biases that strongly influence model predictions: entity popularity, mention order, and semantic co-occurrence. For smaller models, a simple logistic regression using only these surface cues predicts model choices more accurately than the model’s own numerical predictions, suggesting heuristics largely override principled reasoning. Crucially, we find that larger models (32B parameters) selectively rely on numerical knowledge when it is more reliable, while smaller models (7-8B parameters) show no such discrimination, which explains why larger models outperform smaller ones even when the smaller models possess more accurate knowledge. Chain-of-thought prompting steers all models towards using the numerical features across all model sizes.

Item Type: Conference or Workshop Item - published (Paper)
Date Type: Publication
Status: Published
Schools: Schools > Computer Science & Informatics
Publisher: Association for Computational Linguistics
ISBN: 9798891763807
Date of First Compliant Deposit: 16 June 2026
Last Modified: 16 Jun 2026 10:15
URI: https://orca.cardiff.ac.uk/id/eprint/187576

Actions (repository staff only)

Edit Item Edit Item

Downloads

Downloads per month over past year

View more statistics