The next session at the SEASON 2026 conference in Hamburg has started with my great colleague Kateryna Kasianenko from the ARC Centre of Excellence for Automated Decision-Making and Society, who is presenting our work on exploring the impact of query variations on the results that search engines return. But how do we study such query variations systematically? How do we even determine what range of queries people might use as they search for a specific topic?
One answer to this wicked problem is via hackathons: but such hackathons tend to involve only interested participants, and people who may be more similar to each other than would be desirable. How can such issues be addressed? This project implemented a judging approach which specifically rewarded diverse perspectives, and drew on PhD students and early-career scholars from various disciplines; their task was to develop a methodological approach to wicked problems in search. This was done on day one, and on day two these approaches were evaluated.
Problems formulated by these teams were diverse: how to find the right information during local disasters; how to address domestic violence in culturally diverse contexts; how to capture query formulation differences between experts and non-experts; how to address risky queries by children; etc.
Approaches to query development involved developing search personas, based on interviews and other inputs; LLM-based query generation with human validation; surveys of experts and non-experts; and self-formulated queries by the group. For these query approaches, search results were then captured via our in-house Australian Search Experience search query infrastructure or via screenshotting; these were then evaluated.
Such queries could then be used to audit Google, ChatGPT, and YouTube results, and these were often very different from each other, and differently problematic. Kids’ health results were sometimes confronting or misleading, for instance; domestic violence queries sometimes returned results relevant to the US rather than Australia; local disaster queries did not always return local sources, but also national and global sources with less specific information.
Overall, though, the hackathon approach, does offer some pointers to new opportunities, but it does not solve the wicked problem of query formulation by itself; some approaches are more productive than others, but are not necessarily scalable for further studies.












