Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Understanding the limitations of logical reasoning in Language Models

Ariyani, Nurul 2026. Understanding the limitations of logical reasoning in Language Models. PhD Thesis, Cardiff University.
Item availability restricted.

[thumbnail of Thesis]
Preview
PDF (Thesis) - Accepted Post-Print Version
Available under License Creative Commons Attribution Non-commercial No Derivatives.

Download (5MB) | Preview
[thumbnail of Cardiff University Electronic Publication Form] PDF (Cardiff University Electronic Publication Form) - Supplemental Material
Restricted to Repository staff only

Download (207kB)

Abstract

Recent Language Models (LMs) achieve strong performance across a wide range of Natural Language Processing (NLP) tasks, yet they continue to struggle with tasks that require explicit logical reasoning. This gap between apparent general competence and weak deductive reasoning has motivated a growing body of NLP research, especially since many real-world tasks depend on accurate inference, even when they appear linguistically simple. This thesis analyses the deductive reasoning abilities of LMs in controlled settings, more specifically, propositional logic. Our study comprises three complementary strands. First, we analyse whether LMs can learn the representation of logic assertions and manipulate them to determine entailment. Second, given simple, more confined deductive reasoning tasks, we evaluate whether fine-tuned LMs perform consistently better on the tasks or whether their improvement remains limited to in-distribution instances. Third, although recent Large Reasoning Models (LRMs) excel at simple deductive reasoning, we evaluate their ability to handle belief revision, where consistency must be restored with minimal changes after adding conflicting information. To enable systematic evaluation, we formulate entailment checking and belief revision in propositional logic and express them in natural language. This strategy allows the generation of such structured instances and provides control over their difficulty. We evaluate a range of LMs, including embedding-based models, fine-tuned autoregressive models, and prompting-oriented reasoning models. Although LMs can achieve high accuracy in controlled settings, our results show that their reasoning behavior is often brittle and sensitive to changes in problem structures and representation. Learned assertion embeddings are weakly compositional, fine-tuned models generalize poorly on out-of-distribution, and even reasoning-oriented models with advanced reasoning skills struggle to perform principled belief revision. These findings highlight fundamental limitations in the deductive reasoning abilities of modern LMs and motivate further work on systematically evaluating their logical reasoning.

Item Type: Thesis (PhD)
Date Type: Completion
Status: Unpublished
Schools: Schools > Computer Science & Informatics
Subjects: Q Science > QA Mathematics > QA75 Electronic computers. Computer science
Date of First Compliant Deposit: 5 June 2026
Date of Acceptance: 4 June 2026
Last Modified: 05 Jun 2026 14:55
URI: https://orca.cardiff.ac.uk/id/eprint/187426

Actions (repository staff only)

Edit Item Edit Item

Downloads

Downloads per month over past year

View more statistics