The next speaker in this session at the SEASON 2026 conference in Hamburg is my excellent colleague Ashwin Nagappa, who further highlights the challenges of doing research on search engines. Search engine research has long been considered an extension of the information retrieval disciplines, but this has been limiting: various other fields also have a stake in this, and in contrast to the situation, for instance, in social media research the research tools and data access provisions for search engine research are still a great deal more rudimentary.
Our work is in the ARC Centre of Excellence for Automated Decision-Making and Society (ADM+S), through two iterations of the Australian Search Experience project. This started by building on past work by Algorithm Watch, which used a browser plugin to enable ordinary users to donate their search data for generic keyword searches set by the project (names of politicians and parties; major topics in public debate; etc.), and like other studies showed that there was no personalisation of search results except for some localisation for searches that had a distinct local Information need.
Phase two of the project moved away from data donations on a small number of set search terms, therefore, and instead explored the consequences of variations in the search queries themselves; this is therefore free to work with simulated searches by virtual agents and personas, but must face the difficult challenge of formulating a wide range of searches that realistically represent the variety of queries that actual users might pose.
The use of virtual agents is itself somewhat complicated by the fact that search engines regularly block virtual agents scraping search results if they detect them, however; this approach now requires substantial anti-detection measures. Some other tools for this have also emerged in the meantime; these include the SERP API tool, as well as the Result Assessment Tool (RAT) developed in Hamburg.
RAT works as a browser extension operated by the researcher; once installed, it enables researchers to systematically run a variety of queries and capture the results. As it runs, it mimics user behaviour in the browser window, enters queries, navigates the search results page, and captures its contents; this works well overall, but is also time-consuming. What it does manage to do, though, is to capture the experience of an actual user navigating the search interface. However, Google still often locks out users of this tool: it may block users altogether, or require them (repeatedly) to complete a Captcha challenge.
How feasible is the systematic observation of search engine results, then? There clearly is a pressing public interest in this research, given the importance of search engines in society; search engine providers systematically frustrate such research, however, by blocking these research efforts. This makes our observations more difficult, and less reliable, and the operations of search engines are thereby hidden from public scrutiny even while they themselves prolifically gather data on their users search practices.
These anti-scrutiny efforts also frustrate the development and maintenance of public-interest, open-source research tools, of course; and if such tools are not available or their operation is not reliable, then this also means that the quality, reproducibility, and comparability of the research that is done with them becomes more limited. Efforts such as the EU Digital Services Act may help with this eventually, but there’s still a great deal of work to be done here.












