|
Campbell, René, Ołów, Edyta, Shin, Yen, Shin, Jisu and Colombo, Gualtiero
2025.
Evaluating social bias in large language models across different multilingual settings.
Presented at: International Conference on Cybersecurity and Intelligent Networks and Systems (ICCS 2025),
Cardiff, UK,
8 December 2025.
Item availability restricted. |
|
PDF
- Submitted Pre-Print Version
Restricted to Registered users only Download (535kB) | Request a copy |
Abstract
This study investigates how two recent large language models (LLMs), Gemini Flash and Llama-3.3, exhibit social bias in multilingual question-answer- ing (QA) settings. A culturally neutral English dataset was constructed across seven social dimensions by intersecting existing bias focused benchmarks: age, disability, gender identity, physical appearance, religion, socioeconomic status, and sexual orientation. The dataset was translated into German (DE) and Japanese (JA) and evaluated in two prompt conditions: ambiguous prompts (lacking detailed context) and disambiguated prompts (supplying full information to answer the question). Bias is quantified through associated scores, and accuracy is measured in parallel. A translation quality check using multilingual embedding similarity metrics confirms that translation does not contribute to the observed bias patterns. Across all languages, ambiguous prompts amplified bias: Llama-3.3 frequently chose stereotype-consistent answers, while Gemini Flash more often de- faulted to “Unknown.” Disambiguation reduced bias for both models, although Llama-3.3 retained higher bias in categories such as age and religion. These findings highlight how straightforward prompt engineering, given detailed context, can significantly reduce bias for multilingual LLMs, thus providing a scaffold for future multilingual bias benchmark extensions.
| Item Type: | Conference or Workshop Item - unpublished |
|---|---|
| Status: | Unpublished |
| Schools: | Schools > Computer Science & Informatics |
| Subjects: | Q Science > Q Science (General) |
| Uncontrolled Keywords: | Question-Answering, Large Language Models, Bias, Ambiguous, disambiguated |
| Additional Information: | the paper is currenly in production as the submission happened after the conference presentation |
| Last Modified: | 01 Apr 2026 09:00 |
| URI: | https://orca.cardiff.ac.uk/id/eprint/186117 |
Actions (repository staff only)
![]() |
Edit Item |




Download Statistics
Download Statistics