Lansdall-Welfare, Thomas, Sudhahar, Saatviga, Thompson, James, Lewis, Justin ORCID: https://orcid.org/0000-0002-5300-9127, FindMyPast Newspaper Team and Cristianini, Nello 2017. Content analysis of 150 years of British periodicals. Proceedings of the National Academy of Sciences 114 (4) , E457-E465. 10.1073/pnas.1606380114 |
Preview |
PDF (Freely available online through the PNAS open access option)
- Published Version
Download (10MB) | Preview |
Abstract
Previous studies have shown that it is possible to detect macroscopic patterns of cultural change over periods of centuries by analyzing large textual time series, specifically digitized books. This method promises to empower scholars with a quantitative and data-driven tool to study culture and society, but its power has been limited by the use of data from books and simple analytics based essentially on word counts. This study addresses these problems by assembling a vast corpus of regional newspapers from the United Kingdom, incorporating very fine-grained geographical and temporal information that is not available for books. The corpus spans 150 years and is formed by millions of articles, representing 14% of all British regional outlets of the period. Simple content analysis of this corpus allowed us to detect specific events, like wars, epidemics, coronations, or conclaves, with high accuracy, whereas the use of more refined techniques from artificial intelligence enabled us to move beyond counting words by detecting references to named entities. These techniques allowed us to observe both a systematic underrepresentation and a steady increase of women in the news during the 20th century and the change of geographic focus for various concepts. We also estimate the dates when electricity overtook steam and trains overtook horses as a means of transportation, both around the year 1900, along with observing other cultural transitions. We believe that these data-driven approaches can complement the traditional method of close reading in detecting trends of continuity and change in historical corpora.
Item Type: | Article |
---|---|
Date Type: | Publication |
Status: | Published |
Schools: | Journalism, Media and Culture |
Uncontrolled Keywords: | artificial intelligence digital humanities computational history data science, Culturomics |
Additional Information: | Freely available online through the PNAS open access option. |
Publisher: | National Academy of Sciences |
ISSN: | 0027-8424 |
Date of First Compliant Deposit: | 26 July 2017 |
Date of Acceptance: | 30 November 2016 |
Last Modified: | 04 May 2023 12:49 |
URI: | https://orca.cardiff.ac.uk/id/eprint/103004 |
Citation Data
Cited 60 times in Scopus. View in Scopus. Powered By Scopus® Data
Actions (repository staff only)
Edit Item |