Accessibility navigation


How humans and machines identify discourse topics: a methodological triangulation

Gillings, M. and Jaworska, S. ORCID: https://orcid.org/0000-0001-7465-2245 (2025) How humans and machines identify discourse topics: a methodological triangulation. Applied Corpus Linguistics, 5 (1). 100121. ISSN 2666-7991

[img]
Preview
Text (Open Access) - Published Version
· Available under License Creative Commons Attribution.
· Please see our End User Agreement before downloading.

870kB
[img] Text - Accepted Version
· Restricted to Repository staff only

435kB

It is advisable to refer to the publisher's version if you intend to cite from this work. See Guidance on citing.

To link to this item DOI: 10.1016/j.acorp.2025.100121

Abstract/Summary

Identifying and exploring discursive topics in texts is of interest to not only linguists, but to researchers working across the full breadth of the social sciences. This paper reports on an exploratory study assessing the influence that analytical method has on the identification and labelling of topics, which might lead to varying interpretations of texts. Using a corpus of corporate sustainability reports, totalling 98,277 words, we asked 6 different researchers to interrogate the corpus and decide on its main ‘topics’ via four different methods: LLM-assisted analyses; topic modelling; concordance analysis; and close reading. These methods differ according to the amount of data that can be analysed at once, the amount of textual context available to the researcher, and the focus of the analysis (i.e., micro to macro). The paper explores how the identified topics differed both between analysts using the same method, and between methods. We conclude with a series of tentative observations regarding the benefits and limitations of each method, and offer recommendations for researchers in choosing which analytical technique to select.

Item Type:Article
Refereed:Yes
Divisions:Arts, Humanities and Social Science > School of Literature and Languages > English Language and Applied Linguistics
ID Code:120575
Uncontrolled Keywords:triangulation; topic modelling; CADS; concordance analysis; close reading, LLMs, ChatGPT, Claude, AI, Corpus Linguistics, Discourse Analysis
Publisher:Elsevier

Downloads

Downloads per month over past year

University Staff: Request a correction | Centaur Editors: Update this record

Page navigation