Spotlights on Workshops @ACL 2026
General Observations
- First things first: I have been regularly attending *CL Conferences since 2022. Since then, I have never seen such high-quality workshops as this year in San Diego. And even now, after the first day of the main conference, this impression persists. I am very grateful for the many excellent presentations that I could attend in the two workshops that I participated in.
- Reasoning chains enter in focus. The reasoning chains generated by so-called reasoning LLMs are increasingly becoming a subject of study: What are their properties, how is their quality related to the quality of the final predictions, can they be improved, can they be exploited for malicious purposes, etc.
- The Capriciousness of LLM(-Agents). The third recurring theme that I picked up was the fact of the unpredictability of LLM behavior and attempts to get it under control. Several papers have examined aspects of LLM behavior that makes them hard to predict or even outright unreliable, unless used with competence and a view towards their limitations.
Individual Papers/Presentations
- “A Three-Level Audit of LLM Alignment for Argument Quality Assessment” (Link to paper) (ArgMining Workshop). The paper examines to what extent the reasoning chains generated by reasoning LLMs adhere to the instructions given to them (e.g., a rubric for grading arguments), and it assesses whether the reasoning chains actually support the labels predicted at the end by the reasoning LLM.
- Keynote Wachsmuth (ArgMining). Emphasized how argument mining as a task connects to millennia of theorizing about argumentation, and he emphasized the relevance of argumentation for the era of reasoning LLMs: Good reasoning contains good argumentation, and the field of argument mining is well-positioned to contribute here. His work on developing fine-tuning and reinforcement routines to improve argumentative abilities of LLMs fits this window.
- “Are LLM Benchmarks Already Contaminated? A Systematic Review of Contamination Detection Methods” (Link to paper) (GEM Workshop). This paper summarizes current evidence on contamination of benchmarks. This refers to the fact that current frontier LLMs might already have been trained, perhaps deliberately optimized, on some of the benchmarks on which they are then evaluated afterwards. This research is important because it indicates that the performance of models at the benchmark tasks might not be indicative of their abilities in the wild if they have been deliberately optimized for said benchmarks.
- “Autorubric” (Link to Paper) (GEM). A highly sophisticated framework to use LLMs in open-ended evaluation settings (such as short answer grading). What I found impressive here is the elegant integration of technical, pedagogical, and institutional aspects. The framework, which is available as a python library, makes it easy to do use LLMs responsibly in such open-ended tasks.
Postscript: UZH@ArgMining
UZH was well-represented at the ArgMining Workshop. A team from UZH, lead by Yingquian Gao and including myself, ran the shared task on “Reconstructing the Reasoning in United Nations Resolutions” (Link to paper). Also, I was participating in a panel on “Argument Mining Meets Reasoning: Understanding and Evaluating Arguments in Both Human and Machine Reasoning”, which was fun.
Enjoy Reading This Article?
Here are some more articles you might like to read next: