Artificial Intelligence (AI), a powerful catalyst for change, has significantly impacted every facet of contemporary society, reshaping how people live, work, and communicate. Even though it seems to reshape both academic and professional landscapes globally, its role in the field of accessibility has not yet been extensively investigated. When it comes to audio description (AD) for people who are blind or visually impaired (BVIP), as “intersemiotic, intermodal or cross-modal translation or mediation (e.g., Benecke 2007, Bourne & Jiménez Hurtado 2007, Braun 2007, Orero 2005)” in Braun (2008:3), AI tools and tools integrating AI parameters have recently started to be developed, with more and more researchers also trying to approach this topic (e.g., Bergin & Oppegaard 2024, Cheema et. al 2025, and Pacurar 2025). In the case of Greek, no research has been conducted to date on this aspect, making it a challenging and interesting pathway to study to increase the AD output for the primary target audience, namely BVIP, as Mazur (2020) has put it, and to accelerate its workflows.
Our research focuses on the comparison of a human-created AD with an AI-generated version based on the award-winning short film titled “Φωνές” [“Whispers”] by Lefteris Keteoglou-Poulias. In particular, an audio describer created and timed the AD script with the use of Subtitle Edit, assigned a voice talent, based on specific qualitative criteria (Karantzi 2023), to do the voicing in a professional studio, and synchronized it with the film in REAPER. After running several tests with Gemini 3.0 Pro, ChatGPT 5.2, and Claude Sonnet 4.5, it was evident that Claude Sonnet 4.5 outperformed Gemini and ChatGPT 5.2. Hence, it was chosen to create a time-coded AD script in Greek, which was post-edited, and ElevenLabs was assigned to do the voicing to provide consistent, natural-sounding narration. The analysis focused on 5 dimensions derived from existing AD guidelines and research (synchronization, visual element recognition, information prioritization, accuracy, and linguistic relevance) and was conducted by the authors. Findings suggest that even though the AI tools chosen are of good quality and provide a supportive framework for the audio-described film, they show a lack of proper wording, information selection, prioritization, synchronization, consistency, and visual recognition. However, in terms of voicing, ElevenLabs has strong capabilities. Feedback is hoped to be received from members of the Greek BVIP community and the director to validate our results and to ensure the optimal quality for the final deliverable.
Ismini Karantzi (Dr) is a postdoctoral researcher at the Department of Foreign Languages, Translation and Interpreting (Ionian University, Greece), where she previously worked as an Academic Scholar teaching Audiovisual Translation, Technical Translation and Humour in Translation. She has presented her research at international conferences and published papers in international academic journals. She is an official translator and a member of the Panhellenic Association of Professional Translators Graduates of the Ionian University/PEEMPIP), where she also serves as a core mentor in its mentoring programme. She began her professional career in 2015 and has run her own business since 2017, working in translation, editing, and subtitling. Moreover, she has worked as a Translation Project Manager in the Life Sciences and taught Braille at the Regional Association of the Blind of Western Greece. Her academic interests include audiovisual translation, accessibility (AD, SDH, multisensory approach, and artificial intelligence), and readability.
Vilelmini Sosoni holds a PhD in Translation and Text Linguistics from the University of Surrey and is Associate Professor at the Department of Foreign Languages, Translation and Interpreting at the Ionian University where she also serves as Associate Head. She is the Director of the Research Laboratory for Specialised Translation and Language Technologies (STraLaT) and a member of the University’s Gender Equality and Anti-Discrimination Committee. She also has extensive professional experience having worked as a professional translator, editor, content creator and audiovisual services provider. She has participated in several EU-funded projects and has published extensively in the areas of the Translation of Institutional Texts, Machine Translation and Audiovisual Translation and Accessibility, including Audio Description and SDH.
Thanasis Epitidios is a researcher, sound artist, and PhD candidate at the Ionian University's Department of Audio and Visual Arts, holding an MA in Sound Arts and Technologies and a BA in Audio Technology. His research and artistic practice focus on synthetic soundscapes, specializing in real-time procedural audio for game engines and VR. Professionally, he works in sound design and foley for films and immersive media. He serves as a teaching assistant at the Ionian University and presents algorithmic sound installations and electroacoustic performances internationally, including at Évora_27 (Portugal), Audio Art Festival (Poland), and MusicAcoustica (China). He is a member of HELMCA.