Smartcat transcribes your video, lets you edit every line, translates it into any of 280+ languages, and generates a timed AI voice track in 35 voice-over locales — with the original music and effects preserved underneath.
1,000+ enterprise brands choose Smartcat for AI dubbing
AI dubbing replaces a video's spoken audio with a synthetic voice in another language. Smartcat generates the dub from an editable transcript: fix a name or rewrite a sentence, and the voice says exactly that. AI voice-over is available in 35 locales — powered by ElevenLabs [1] — and it works best on training, product, and marketing video.
35
Voice-over locales
Translation covers 280+ languages, but AI voice-over is a shorter list: 35 locales powered by ElevenLabs, including Japanese, Chinese, German, Hindi, Korean, French (France and Canada), and English in US, UK, Australian, and Canadian variants. Where no voice exists for your target language, export translated subtitles instead.
5
Steps from upload to export
Upload the video, fix the transcript once, pick languages and voices, review the dub against the picture, export. Speech is separated from background audio, so music and sound effects carry over under the new voice instead of disappearing.
10
Smartwords per source word
Dubbing is metered by transcript word count, not video minutes — so a long, quiet video costs less than a short, dense one. Standard translation is 1 Smartword per source word, so budget dubbing at roughly ten times the text-translation cost of the same script. Details: pricing.
Studio dubbing is priced per minute, per language. A 10-minute training video into eight languages is a five-figure quote, so the dubbing line gets cut and seven markets read subtitles — or nothing.
Each language adds weeks: casting, recording, review, re-recording. By the time the German dub is approved, the product in the video has changed.
And teams fear the robot voice. Everyone has heard bad AI dubbing, so brands quietly decide "not for us" without testing whether the failure they remember is one the tooling has since fixed.
enterprise brands on Smartcat
translation languages
video formats supported (WebM is not)
maximum upload size per file
Mispronounced names, a voice rushing through a long translation, music that vanishes — each one has a control in the editor. Set the pronunciation rule once, retime the cue, separate the background audio. The fourth reason is a real limit, and it is what decides when to ship subtitles instead.
AI dubbing in Smartcat runs from an editable transcript rather than straight from audio: the spoken track is transcribed into timed cues, you correct names and product terms once, and every language follows from the corrected text. Here is the whole flow.
Transcribe — the Media Translation Agent turns the video's spoken audio into timed, editable cues. Dubbing works from speech, not from text baked into the picture.
Correct — you fix the transcript once (names, acronyms, product terms) before anything is translated.
Translate — the AI translates the cues into your target languages, reusing your glossary and translation memory.
Voice — an AI voice track is generated, timed to the original cues, in one of 35 locales and powered by ElevenLabs; editing a subtitle line changes what the AI voice says.
Review and export — any language can be assigned to a professional human reviewer in the same workflow, and you can re-generate individual lines, including pronunciation fixes, without redoing the whole video.
Speech is separated from background audio, so music and sound effects carry over under the new voice instead of disappearing. Dubbing is metered by transcript word count rather than video length, at 10 Smartwords per source word.
From upload to export in five steps — you fix the transcript once, and every language follows
1
Upload the video
Sixteen video formats are supported: 3GP, 3G2, AVI, FLV, M2V, M4V, MKV, MOV, MP4, MPEG, MPG, OGV, QT, WMV, TS and VOB. WebM is not supported, so convert it first.
2
Fix the transcript once
Correct names, acronyms, and product terms in the timed transcript before anything is translated. Every correction flows into every language.
3
Pick languages and voices
Translate into any of 280+ languages, and choose an AI voice per language across the 35 voice-over locales. Where a locale has no voice, export translated subtitles instead.
4
Review the dub against the picture
Edit any cue in the subtitle editor and the voice for that line re-renders. Hand a language to a professional human reviewer in the same workflow if the stakes warrant it.
5
Export
Download the dubbed video, or the voice track on its own.
Upload a video, fix the script once, and hear the difference an editable transcript makes. Free for 15 days with 15,000 Smartwords — full access to translation capabilities, no credit card.
2–3 days
Instead of Ten, at Smith+Nephew
Smith+Nephew cut eLearning translation turnaround from an average of ten days with their previous providers to two to three days for the same course length on life-science and medical content.
30%
More Output With the Same Team
Wunderman Thompson increased localized content output by 30% with the same team and resources.
50%
Lower Costs, Higher Productivity
expondo improved productivity by 50% and cut localization costs by half.
Upload a video, fix the script once, and ship a natural voice track in every market that matters — with subtitles for the content AI shouldn't voice. Free for 15 days, no credit card.
YouTube auto-dub is automatic and out of your hands — you can't edit the script, fix a pronunciation, or review a language before viewers hear it (you can disable it per video in YouTube Studio's audio settings). Smartcat is the opposite trade: you control the transcript, voices, and review step, and you own the exported file for any channel.
Sixteen video formats are supported: 3GP, 3G2, AVI, FLV, M2V, M4V, MKV, MOV, MP4, MPEG, MPG, OGV, QT, WMV, TS and VOB. WebM is not supported, so convert it first.
None of these limits vary by plan.
No. Video translation works from spoken audio. Text embedded in the picture — titles, lower thirds, graphics, overlays — cannot be extracted or translated automatically; you would need to edit the original video source files to change it.
Translation covers 280+ languages, but AI voice-over is a shorter list: 35 locales, powered by ElevenLabs, including Japanese, Chinese, German, Hindi, Korean, French (France and Canada), and English in US, UK, Australian, and Canadian variants. Where a voice isn't available for your target language, export translated subtitles instead.
Yes — edit any cue in the subtitle editor and the voice for that line re-renders. You don't regenerate the whole video to fix one sentence.
Yes. Assign any language to a vetted reviewer from Smartcat's Marketplace inside the same project; their edits update the cues, and the voice re-renders from the corrected text.
Smartcat is SOC 2 Type II compliant; files are encrypted in transit and at rest, and workspaces are isolated with role-based access. Details: security.
Question not answered here? Book a demo — a 1:1 consultation, no commitment.
Most bad AI dubbing fails for one of four reasons, and three of them are fixable.
Two more limits before you upload: lip-sync is time-matched, not mouth-matched, so tight close-ups on a speaker's face are dubbing's worst case and subtitles' best; and a noisy source degrades everything downstream, which is why the editable-transcript step exists.
Dubbing replaces the voice. Voice-over translation lays a translated narration over the original, which stays audible underneath. Subtitle translation leaves the audio untouched and puts the translation on screen. All three come from the same editable transcript, so you can pick per video — or export more than one — without re-translating.
1. Downie, A., & Hayes, M. (2025, April 17). What is AI voice? IBM.