DittoDub logo

Introducing Native 6

Published by Ditto Team · 4 min read · 4 days ago

DittoDub Native 6 in blue and gold

There is a moment in a good video when you forget you are watching someone on a screen. You hear the excitement in their voice, or the way it softens when a story becomes personal. You feel what they mean.

That feeling should survive translation.

Today, we are introducing Native 6, our new dubbing model. It carries the original speaker's voice and emotion into another language with extraordinary fidelity. We believe it is the best dubbing model ever made.

Available now in 124 languages Starting today, every new video you upload to DittoDub uses Native 6 automatically. Upload your video and choose your languages. The new model is already part of the workflow.

The emotion stays with the speaker

The same words can sound playful, frustrated, reassuring, or uncertain. A good translation needs to carry that difference. Otherwise, the viewer understands the sentence but misses the person behind it.

Native 6 follows the emotion in the original performance. When a speaker gets excited, the translated voice carries that excitement. When they lower their voice for a serious moment, the delivery follows. The goal is for someone watching in another language to feel as close to the speaker as someone watching the original.

This is what makes Native 6 feel so different. The voice belongs to the person on screen, and the feeling belongs to the moment.

A closer voice match

The translated speaker sounds more like the person your audience already knows.

Emotion that carries through

The energy and expression of the original performance stay with the translated voice.

Clearer pronunciation

Stronger pronunciation in low resource languages helps more audiences hear a natural delivery.

A difference audiences can hear

In a blind test of international audiences, 86% preferred Native 6 over human dubbing.

86%

preferred Native 6 over human dubbing

In a blind test of international audiences.

A step forward from Speak 5

Compared with our previous model, Speak 5, Native 6 improves speaker similarity by 12%, emotional similarity by 21%, and pronunciation in low resource languages by 38%.

These are the qualities that make a dub feel familiar: how closely the voice matches, how faithfully the emotion comes through, and how naturally the words are pronounced.

Speak 5 is set to 100% for each measure. The chart shows relative performance, not absolute accuracy. Native 6 improvements are compared with Speak 5.

More of the world should hear you clearly

Native 6 is available in 124 languages, with a particularly strong improvement in pronunciation for low resource languages. A creator's voice should feel just as considered in a language with fewer available training resources as it does in a widely spoken one.

For creators reaching new audiences, that means more of the original performance can travel with the video. Viewers hear the explanation, the story, or the joke in a language they understand, with a voice that still feels like you.

Ready for your next upload

Native 6 is now the default for new video uploads in DittoDub. There is no model to switch on and no new process to learn. Create a project, upload your video, and choose the languages you want to reach.

Start with a video where the delivery matters. An excited explanation, a personal story, a moment that makes people laugh. Listen to the original, then listen to the dub. That is where Native 6 makes itself heard.

Common Questions

What is Native 6?

Native 6 is DittoDub's new dubbing model. It translates video into 124 languages while closely preserving the original speaker's voice and emotional delivery.

How do I use Native 6?

Native 6 is the default for new video uploads. Create a project in DittoDub, upload a video, and choose your languages. There is no separate model setting to enable.

What has improved since Speak 5?

Native 6 improves speaker similarity, emotional similarity, and pronunciation in low resource languages. The result is a translated voice that stays closer to the original performance.