Skip to main content
Home
Snurblog — Axel Bruns

Main navigation

  • Home
  • Information
  • Blog
  • Research
  • Publications
  • Presentations
  • Press
  • Creative
  • Search Site

Towards a Systematic Test of LLM Codings of Populist Speeches

Snurb — Friday 11 September 2026 23:20
Politics | Elections | Polarisation | Artificial Intelligence | ECREA 2026 | Liveblog |

The post-lunch session on this last day at the ECREA 2026 conference in Brno is on computational methods for political discourse analysis, and starts with Felipe Barreto-Storandt. He begins by noting the growing adoption of LLMs in communication research, but also highlights that there are no standard methodologies for using them yet; this is experimental at this point, and not a planned feature by their designers.

Often they are simply run once over a dataset, though, and this creates only one data point. But because LLMs are stochastic, they will produce different scores every time; this means that multiple runs create a distribution. Further, setup decisions might be changed with each tun, producing a results grid, which also enables the bias of specific context settings to become measurable.

This paper tests this for a hard case: the assessment of populism. This is a thin, contested, contextual construct rather than a distinct ideology, which might make the detection of populism especially complicated; further, the idea of ‘the people’ that populism refers to is not static, and highly context-dependent, providing additional complications.

With this approach, we can test whether LLMs can replace human coders; whether their errors are systematic; and whether technical choices matter. The project applied this to a corpus of political speeches, used contextual variables from V-Dem and other sources, drew on V-Party and other datasets to assess the political positions of speakers, and engaged in language detection to assess the prevalence of languages in training datasets.

For some 670 speeches across 36 countries, this tested some 324 model prompt configurations in five runs across several models from the US, Europe, and China, at various levels of precision. Results were assessed for their difference from the human coding of each speech.

Overall, LLM reliability was below the human level, but validity was actually very high. Within-model variation was better for large models. Consistency was better for higher democratic tercile, and less reliable but more consistent for lower democratic terciles. The LLM is better in identifying higher-populism cases, and few-shot prompting increases rather than decreases error rates.

These are somewhat confounding results, and raise various concerns. New models will of course change these results again, too; this means that we need to start thinking about LLMs in new ways: we shouldn’t constantly try to prove that their results match human coding, but humans are themselves also biased – rather, we should systematically test which features of speeches trigger classifications as populist. These results might then be able to shared throughout the field.

  • 5 views
INFORMATION
BLOG
RESEARCH
PUBLICATIONS
PRESENTATIONS
PRESS
CREATIVE

Recent Work

Presentations and Talks

Revisiting ‘the’ Public Sphere and Its Algorithmically Shaped Publics (ZeMKI ComAI 2026)

» more

Books, Papers, Articles

Untangling the Furball: A Practice Mapping Approach to the Analysis of Multimodal Interactions in Social Networks (Social Media + Society)

» more

Opinion and Press

Breaking through Infoglut: The Anger-Information Overload Cycle (360info)

» more

Creative Work

Brightest before Dawn (CD, 2011)

» more

Lecture Series


Gatewatching and News Curation: The Lecture Series

Bluesky profile

Mastodon profile

Queensland University of Technology (QUT) profile

Google Scholar profile

Mixcloud profile

[Creative Commons Attribution-NonCommercial-ShareAlike 4.0 Licence]

Except where otherwise noted, this work is licensed under a Creative Commons BY-NC-SA 4.0 Licence.