Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Large language models eroding science understanding: an empirical study of ‘malignment’

Collins, Harry, Grote, Hartmut ORCID: https://orcid.org/0000-0002-0797-3943, Newbury, Paul, Sutton, Patrick ORCID: https://orcid.org/0000-0003-1614-3922 and Thorne, Simon 2026. Large language models eroding science understanding: an empirical study of ‘malignment’. AI and Ethics 6 (5) , 556. 10.1007/s43681-026-01391-x

[thumbnail of s43681-026-01391-x.pdf] PDF - Published Version
Available under License Creative Commons Attribution.

Download (725kB)

Abstract

Large Language Models (LLMs) increasingly provide scientific information to non-specialist users, yet their linguistic fluency can create an impression of expertise that exceeds their capacity for scientific judgement. This paper experimentally investigates how readily an LLM can be configured to present fringe scientific claims as authoritative knowledge. We constructed two RAG-like Custom GPTs using GPT-5.1, one concerned with the fine-structure constant and the other with gravitational waves. In each case, ten papers selected by domain experts from the alternative preprint repository viXra were supplied as a domain-specific knowledge base, together with instructions directing the model to treat this material as authoritative. The underlying model was neither retrained nor fine-tuned. Expert-generated questions were then posed to the resulting “Fringe-LLMs”, to the unmodified model, and to domain experts. The Fringe-LLMs produced fluent, detailed and apparently authoritative answers that systematically reflected the fringe corpus, while expert assessment found many of these answers to be scientifically incorrect or seriously misleading. By contrast, the unmodified model generally produced answers substantially closer to mainstream scientific understanding. We describe this deliberate redirection of model output as malignment. The experiment demonstrates that scientifically misleading LLMs can be created with a low technical barrier and without altering model weights or poisoning training data. Because non-specialists may be unable to distinguish such outputs from genuine expertise, the results raise important concerns about the use of LLMs as intermediaries for scientific knowledge and reinforce the continuing importance of expert human judgement and provenance.

Item Type: Article
Date Type: Publication
Status: Published
Schools: Schools > Physical, Chemical & Environmental Sciences
Schools > Physics and Astronomy
Publisher: Springer
Date of First Compliant Deposit: 28 September 2026
Date of Acceptance: 8 September 2026
Last Modified: 28 Sep 2026 15:45
URI: https://orca.cardiff.ac.uk/id/eprint/189852

Actions (repository staff only)

Edit Item Edit Item

Downloads

Downloads per month over past year

View more statistics