AI Dubbing From an Editable Transcript — 280+ Languages, 35 AI Voices

Smartcat transcribes your video, lets you edit every line, translates it into any of 280+ languages, and generates a timed AI voice track in 35 voice-over locales — with the original music and effects preserved underneath.

Translate any file. In seconds.
Images, docs, video, audio, courses — pick a format and drop your file.
Drop images here
or browse files · Supports .png, .jpg, .webp, .svg and more

1,000+ enterprise brands choose Smartcat for AI dubbing

What AI Dubbing Actually Does

AI dubbing replaces a video's spoken audio with a synthetic voice in another language. Smartcat generates the dub from an editable transcript: fix a name or rewrite a sentence, and the voice says exactly that. AI voice-over is available in 35 locales — powered by ElevenLabs [1] — and it works best on training, product, and marketing video.

35

Voice-over locales

Translation covers 280+ languages, but AI voice-over is a shorter list: 35 locales powered by ElevenLabs, including Japanese, Chinese, German, Hindi, Korean, French (France and Canada), and English in US, UK, Australian, and Canadian variants. Where no voice exists for your target language, export translated subtitles instead.

5

Steps from upload to export

Upload the video, fix the transcript once, pick languages and voices, review the dub against the picture, export. Speech is separated from background audio, so music and sound effects carry over under the new voice instead of disappearing.

10

Smartwords per source word

Dubbing is metered by transcript word count, not video minutes — so a long, quiet video costs less than a short, dense one. Standard translation is 1 Smartword per source word, so budget dubbing at roughly ten times the text-translation cost of the same script. Details: pricing.

Why Do Most Dubbed Videos Never Ship?

Studio dubbing is priced per minute, per language. A 10-minute training video into eight languages is a five-figure quote, so the dubbing line gets cut and seven markets read subtitles — or nothing.

Each language adds weeks: casting, recording, review, re-recording. By the time the German dub is approved, the product in the video has changed.

And teams fear the robot voice. Everyone has heard bad AI dubbing, so brands quietly decide "not for us" without testing whether the failure they remember is one the tooling has since fixed.

1,000+

enterprise brands on Smartcat

280+

translation languages

16

video formats supported (WebM is not)

6 GB

maximum upload size per file

Three of the Four Reasons Are Fixable

Mispronounced names, a voice rushing through a long translation, music that vanishes — each one has a control in the editor. Set the pronunciation rule once, retime the cue, separate the background audio. The fourth reason is a real limit, and it is what decides when to ship subtitles instead.

How Does AI Dubbing Work?

AI dubbing in Smartcat runs from an editable transcript rather than straight from audio: the spoken track is transcribed into timed cues, you correct names and product terms once, and every language follows from the corrected text. Here is the whole flow.

  • Transcribe — the Media Translation Agent turns the video's spoken audio into timed, editable cues. Dubbing works from speech, not from text baked into the picture.

  • Correct — you fix the transcript once (names, acronyms, product terms) before anything is translated.

  • Translate — the AI translates the cues into your target languages, reusing your glossary and translation memory.

  • Voice — an AI voice track is generated, timed to the original cues, in one of 35 locales and powered by ElevenLabs; editing a subtitle line changes what the AI voice says.

  • Review and export — any language can be assigned to a professional human reviewer in the same workflow, and you can re-generate individual lines, including pronunciation fixes, without redoing the whole video.

Speech is separated from background audio, so music and sound effects carry over under the new voice instead of disappearing. Dubbing is metered by transcript word count rather than video length, at 10 Smartwords per source word.

From upload to export in five steps — you fix the transcript once, and every language follows

1

Upload the video

Sixteen video formats are supported: 3GP, 3G2, AVI, FLV, M2V, M4V, MKV, MOV, MP4, MPEG, MPG, OGV, QT, WMV, TS and VOB. WebM is not supported, so convert it first.

  • Files up to 6 GB; split anything over 1 GB for faster processing.
  • No published maximum video length — the constraint is file size, not running time.
  • No cap on how many files go into one project, on any plan.

2

Fix the transcript once

Correct names, acronyms, and product terms in the timed transcript before anything is translated. Every correction flows into every language.

3

Pick languages and voices

Translate into any of 280+ languages, and choose an AI voice per language across the 35 voice-over locales. Where a locale has no voice, export translated subtitles instead.

4

Review the dub against the picture

Edit any cue in the subtitle editor and the voice for that line re-renders. Hand a language to a professional human reviewer in the same workflow if the stakes warrant it.

5

Export

Download the dubbed video, or the voice track on its own.

See It Dub a Video

Upload a video, fix the script once, and hear the difference an editable transcript makes. Free for 15 days with 15,000 Smartwords — full access to translation capabilities, no credit card.

What Teams Ship With Smartcat Video Localization

2–3 days

Instead of Ten, at Smith+Nephew

Smith+Nephew cut eLearning translation turnaround from an average of ten days with their previous providers to two to three days for the same course length on life-science and medical content.

30%

More Output With the Same Team

Wunderman Thompson increased localized content output by 30% with the same team and resources.

50%

Lower Costs, Higher Productivity

expondo improved productivity by 50% and cut localization costs by half.

Rated by Smartcat users on G2

Dub the Videos Your Budget Said You Couldn't

Upload a video, fix the script once, and ship a natural voice track in every market that matters — with subtitles for the content AI shouldn't voice. Free for 15 days, no credit card.

AI dubbing — the fine print

How is this different from YouTube's auto-dubbing?

YouTube auto-dub is automatic and out of your hands — you can't edit the script, fix a pronunciation, or review a language before viewers hear it (you can disable it per video in YouTube Studio's audio settings). Smartcat is the opposite trade: you control the transcript, voices, and review step, and you own the exported file for any channel.

What video formats can I upload, and how long can a video be?

Sixteen video formats are supported: 3GP, 3G2, AVI, FLV, M2V, M4V, MKV, MOV, MP4, MPEG, MPG, OGV, QT, WMV, TS and VOB. WebM is not supported, so convert it first.

  • Size — the upload ceiling is 6 GB, and Smartcat recommends splitting files over 1 GB for faster processing.
  • Length — Smartcat publishes no maximum video duration; the documented constraint is file size, not running time, so a 90-minute recording is a size question rather than a length one.
  • Count — there is no cap on how many videos go into one project, and batch uploads of hundreds of files are supported.

None of these limits vary by plan.

Does dubbing translate the text on screen?

No. Video translation works from spoken audio. Text embedded in the picture — titles, lower thirds, graphics, overlays — cannot be extracted or translated automatically; you would need to edit the original video source files to change it.

How many languages have AI voices?

Translation covers 280+ languages, but AI voice-over is a shorter list: 35 locales, powered by ElevenLabs, including Japanese, Chinese, German, Hindi, Korean, French (France and Canada), and English in US, UK, Australian, and Canadian variants. Where a voice isn't available for your target language, export translated subtitles instead.

Can I edit the dub after it's generated?

Yes — edit any cue in the subtitle editor and the voice for that line re-renders. You don't regenerate the whole video to fix one sentence.

Can a human review the dub before it ships?

Yes. Assign any language to a vetted reviewer from Smartcat's Marketplace inside the same project; their edits update the cues, and the voice re-renders from the corrected text.

Is my video confidential?

Smartcat is SOC 2 Type II compliant; files are encrypted in transit and at rest, and workspaces are isolated with role-based access. Details: security.

Question not answered here? Book a demo — a 1:1 consultation, no commitment.

Why does AI dubbing sound bad — and when won't it?

Most bad AI dubbing fails for one of four reasons, and three of them are fixable.

  • Mispronounced names — fixed by pronunciation rules applied once, everywhere.
  • A voice that rushes because the translation ran long — fixed by editing the cue in the subtitle editor.
  • Music and effects vanishing — fixed by background-sound separation.
  • Synthetic voices don't act — the fourth reason is a real limit: they read scripts well, but for emotional or comedic material, use subtitles or human voice talent.

Two more limits before you upload: lip-sync is time-matched, not mouth-matched, so tight close-ups on a speaker's face are dubbing's worst case and subtitles' best; and a noisy source degrades everything downstream, which is why the editable-transcript step exists.

What's the difference between dubbing, voice-over translation and subtitles?

Dubbing replaces the voice. Voice-over translation lays a translated narration over the original, which stays audible underneath. Subtitle translation leaves the audio untouched and puts the translation on screen. All three come from the same editable transcript, so you can pick per video — or export more than one — without re-translating.

Sources

1. Downie, A., & Hayes, M. (2025, April 17). What is AI voice? IBM.