Antypas, Dimosthenis
2025.
Language models for analysing social media discourse: resources and applications.
PhD Thesis,
Cardiff University.
Item availability restricted. |
Preview |
PDF
- Accepted Post-Print Version
Available under License Creative Commons Attribution Non-commercial No Derivatives. Download (20MB) | Preview |
|
PDF (Cardiff University Electronic Publication Form)
- Supplemental Material
Restricted to Repository staff only Download (451kB) | Request a copy |
Abstract
Social media has become central to daily life for millions worldwide, serving as a space where people exchange opinions, express concerns, and influence one another. From natural disasters to political debates, platforms such as X (formerly Twitter) and Reddit are now primary sources of news and public discourse. This activity generates vast amounts of largely unstructured data, creating major challenges for researchers seeking to study and analyse it effectively. Natural Language Processing (NLP) offers tools to handle this complexity, but social media poses unique difficulties compared to more formal or structured text. Rapidly evolving slang, trending topics, emojis, and the brevity of posts create hurdles that traditional NLP methods often struggle to overcome. Despite the impact of models such as BERT and RoBERTa, and more recently Large Language Models (LLMs) like GPT and LLaMA, social media data’s distinct characteristics continue to pose difficulties. While LLMs show impressive, broad capabilities, architectural differences mean that smaller, encoder-based models can be more efficient and effective for certain domain-specific applications. In this work, we present a collection of high-quality, domain-specific datasets sourced from social media to support the training and evaluation of targeted NLP models. Building on these resources, we develop small encoder-based models tailored to specific tasks, demonstrating that their performance can match, and in some cases exceed, that of larger decoder-based LLMs on domain-specific benchmarks. These tools are applied to case studies involving politically charged discourse, hate speech detection, and misinformation analysis. Collectively, this research highlights that specialised and efficient approaches can achieve competitive performance while supporting scalable and sustainable social media research.
| Item Type: | Thesis (PhD) |
|---|---|
| Date Type: | Completion |
| Status: | Unpublished |
| Schools: | Schools > Computer Science & Informatics |
| Subjects: | Q Science > QA Mathematics > QA75 Electronic computers. Computer science Q Science > QA Mathematics > QA76 Computer software |
| Date of First Compliant Deposit: | 2 June 2026 |
| Date of Acceptance: | 26 May 2026 |
| Last Modified: | 02 Jun 2026 14:16 |
| URI: | https://orca.cardiff.ac.uk/id/eprint/187354 |
Actions (repository staff only)
![]() |
Edit Item |




Download Statistics
Download Statistics