In our blog series Meet the Minds Inventing the Future of Video, we’ve been going behind the scenes to find out more about some of Ateme’s brightest minds and what they’ve been working on. In this new edition, we introduce Robin Herin, Director of Standardization. He spoke with us about AI-Based Speech-to-Text and Speech-to-Speech translation.
What is Your Role at Ateme?
I am a Director of Standardization within the CTO Office.
What Have You Been Working on at Ateme?
Content localization using AI-Based Speech-to-Text and Speech-to-Speech translation.
What is Speech-to-Text and Speech-to-Speech translation?
Essentially, it allows to add additional languages to a piece of content, both as a live or a file media asset.
Let’s start with Speech-to-Text first. In order to add the desired languages, the first step is transcription : we use an ASR model (Automated Speech Recognition) to ingest the incoming audio and output text. Said text can be used immediately as basic subtitles (without translation) or it can be passed onto an LLM (Gemini for instance) for translation in additional languages. The additional text is then conditioned in the chosen subtitling format (DVB Sub or TTML for instance) and re-added into the original container for the media (MPEG2-TS over SRT, etc…).
In the case of Speech-to-Speech, we reproduce the steps used for Speech-to-Text but add also another component which is dubbing. It uses another LLM that can ingest text and create audio based on that, either through a pre-existing digital voice or by creating a synthetic digital replica that matches the tone & emotions of the original speaker (what is commonly called as Voice Cloning). Similarly to the previous process, the output audio is then conditioned in the chosen audio format (AAC being the most frequent) and re-added into the original container for the media.
What Industry Challenges Does it Address?
Two aspects : accessibility & distribution
Being able to add additional languages in subtitling for instance allows for the content to be more accessible, including for instance for people with impaired hearing. The European Accessibility Act was just ratified last year in Europe and incudes regulations around this topic, making sure that content can be enjoyed by everyone regardless of their situation.
It is also an opportunity for rights owners to distribute existing content to additional countries or regions. This is particularly important in terms of preservation in places where the native language is not the same as the national language. For instance, Spain uses Spanish but also Galician, Basque or Catalan. In New Zealand, English is commonly spoken but Māori is still very much present.
What Have You Achieved in This Field at Ateme?
Not only we have achieved functionality using both Live & VOD assets, we were also able to work with some partners on the topic and achieve production quality during Live events. Specifically, our work with Syncwords and the integration between our TITAN Live & NEA Live OTT headend and their Syncwords live service leveraged our common CMAF Ingest implementation.
What Does it Change for Viewers?
Audience will benefit from an improved experience for everyone as well as a more fair experience to those who need accessibility features. Content is now more global than ever and viewers can enjoy the same shows across every almost every timezone and languages there is.