The rapid development of large language models (LLMs) and the ongoing evolution of machine translation systems present challenges and opportunities for institutional translation workflows. In this context, it is important to assess how emerging technologies, such as LLMs, compare with established neural machine translation (NMT) engines in terms of translation quality and practical usefulness.
We present a study conducted by the Language Support & Innovation Section of the European Central Bank, where we compared the quality of different MT models and LLMs.
The study assessed the ECB’s current machine translation engine, eTranslation (Finance domain), alongside other tools, including DeepL and RWS Language Weaver, as well as ChatGPT 5.1. The ultimate aim was to determine which tools best support professional translation work in different text types.
The study involved 77 expert linguists from the ECB and National Central Banks. The output of the aforementioned tools was evaluated in translations from English into 23 target languages. The texts included ECB standard publications such as the Annual Accounts and Projections, as well as excerpts from a blog post and a speech. Each text was at least 15 segments. The assessment data were then collected and analysed to identify the best-performing models by language and text type.
The study showed that neural machine translation remains a useful support tool for standard publications, outperforming ChatGPT5.1 on expert texts, particularly when the engine is trained on relevant material. It also showed that Large Language Models are a promising emerging technology, particularly for less standardised text types.
Cristina Farroni holds a PhD in Multilingual Communication and Translation from the University of Macerata (Italy). She currently works as a Language Technologist and Terminologist in the Language Support & Innovation Section of the Language Services Division at the European Central Bank.