Harder Than It Looks: Methodological Challenges in Experiments Comparing Simultaneous and Machine Interpreting
Date & Time: 11/10/2026 (12:30-14:00)
Location: Teaching Room 1 - Ionian University Building
Tomasz Korybski (University of Warsaw), Małgorzata Tryuk (University of Warsaw), Wojciech Figiel (University of Warsaw)

As technology providers announce rapid — and at times market-ready — advances in machine simultaneous interpreting, and as the underlying ASR+MT+TTS pipelines continue to evolve at pace while entirely new architectures emerge, including large language model-augmented and end-to-end speech-to-speech systems that bypass the traditional cascaded paradigm altogether, independent experimental research has never been more urgent.

Mounting a live experiment that places human interpreters and machine interpreting (MI) systems in direct, simultaneous comparison is considerably more complex than it might initially appear. Decisions that seem purely logistical — who interprets, in which direction, from what source, for which audience, and evaluated by what means — carry substantial theoretical weight, and the choices made at each point shape not only what can be measured but what ultimately counts as evidence. This paper draws on recent experimental work in the field (Korybski et al., 2026) to propose a set of general methodological guidelines for researchers designing comparative live studies of human and machine simultaneous interpreting.

The guidelines address five recurring challenge areas. First, interpreter cohort design: the choice on a continuum between professional and novice human interpreters is rarely neutral, as audience attributions of quality may attach to individual performers rather than to modalities, confounding modality-level conclusions. Second, directionality: bidirectional experimental designs increase ecological validity but introduce asymmetric quality effects that require careful analytical separation – particularly considering the present-day dominance of English in investigated language pairs. Third, audience profile and expertise: expert listeners (e.g. interpreters, trainers, and students of interpreting) may  apply more stringent evaluative criteria to MI output than non-expert audiences, raising the question of whose perception should constitute the primary validity criterion. Fourth, live data collection instruments: managing the audio streams and multi-track recordings requires meticulous preparation, and embedding surveys within event environments offers immediacy but risks response quality trade-offs that must be anticipated in the design phase. Fifth, the issue of commercial provider involvement or non-commercial and tailor-made pipelines for MI: the participation of proprietary MI systems introduces uncertainty as to the underlying technology, asymmetric conditions of engagement, and reproducibility concerns that remain under-theorised in the methodological literature.

Taken together, these considerations point toward a broader challenge articulated by Pöchhacker (2024): whether MI evaluation requires a framework genuinely distinct from those developed for assessing human performance. We argue that this question is not merely theoretical — rather, it is a pertinent design decision that must be confronted before data collection begins.

Tomasz Korybski (University of Warsaw)

Dr Tomasz Korybski is Assistant Professor at the Institute of Applied Linguistics, University of Warsaw, and a freelance conference interpreter with over 20 years of experience. His current research focuses on interpreting technology testing, quality evaluation and benchmarking of both human and machine interpreting, and technology-assisted interpreting workflows. Dr Korybski is co-editor of the Routledge Handbook of Interpreting, Technology and AI (2025) and serves as a reviewer for leading peer-reviewed journals in interpreting studies and translation technology.

Małgorzata Tryuk (University of Warsaw)

n/a

Wojciech Figiel (University of Warsaw)

n/a


Back