AI Video Translator: Subtitles, Dubbing, or Voice-over From One Upload

Smartcat transcribes what’s said, keeps the timing, and translates it into any of 280+ languages — then renders it back as subtitles, or as a dubbed or voice-over track in 35 voice locales. Your choice, same file.

Translate any file. In seconds.
Images, docs, video, audio, courses — pick a format and drop your file.
Drop images here
or browse files · Supports .png, .jpg, .webp, .svg and more

1,000+ enterprise brands localize their video in Smartcat 

How to Translate a Video With Smartcat

Smartcat’s AI video translator transcribes a video’s speech into a timed transcript, then translates that transcript into 280+ languages using your glossary and translation memory. The result renders as translated subtitles, an AI-dubbed audio track, or a voice-over. The transcript is editable before translation, and any language can be reviewed by a human before export.

Upload the video

Drop in the video or audio file — or an existing .srt or .vtt subtitle file if you already have one. Smartcat detects the source language.

Check the transcript

Pick languages and output

Review only what matters

Export

Upload the video

Drop in the video or audio file — or an existing .srt or .vtt subtitle file if you already have one. Smartcat detects the source language.

Three Outputs From One Upload

The same transcript feeds subtitles, dubbing, and voice-over, so you can subtitle a video for eight markets and dub it for the two that matter most — without re-uploading or re-translating.

Pick per language, not per project

Subtitles, dubbing, and voice-over all run from the same transcript. Go deeper on each: subtitle and SRT translation, AI dubbing, or voice-over translation.

Fix the transcript once, and it’s right in every language

Translation runs from an editable timed transcript rather than straight from audio, so a misheard product name or acronym is corrected a single time — before any language is generated.

Subtitles re-timed for the language they’re in

When a translation runs longer than the source — normal going into German, Finnish, or Russian — cues are adjusted so lines stay readable and stay inside the shot. Every cue is visible and editable.

Your terminology carries across every video

Approved terms and past translations are stored and reused, so the product name, the tagline, and the disclaimer come out identically in video forty as in video one — and words you already paid to translate aren’t charged again.

Why Translated Video Usually Needs a Second Pass by Hand

Automatic video translation fails in three predictable places, and none of them is the translation itself.

1

The transcript was already wrong

Speech-to-text mishears your product name once — and that single error is then translated faithfully into all twelve languages. You find it when a regional team asks what the product in the video is called.

2

The timing stops fitting

German runs noticeably longer than English, Finnish longer still. Subtitles that sat comfortably under a shot in English overrun the cut in German, so someone opens every cue and re-times it.

3

The picture doesn’t get translated

Translation is speech-only. A slide behind the speaker, or an English menu in a screen recording, stays in the source language unless someone edits the video itself.

Rated for Setup and Everyday Use

9.6/10

ease of setup, on G2

9.3/10

ease of use, on G2

1,000+

enterprise brands

Top AI Choice for Employee Training

Training Industry recognized Smartcat as a leading AI provider, offering AI-powered video training creation tools that enhance learning and streamline content development, delivery, and analytics to meet organizational needs.

Talk to a Media Localization Specialist

A 1:1 consultation with a media localization specialist — no commitment, no form queue.

Video Translation Guides & Resources

How to Add Translated Subtitles to a Video

Where subtitle translation fits into an existing video editing workflow, step by step.

How to Translate an SRT File

Translate .srt and .vtt subtitle files, or generate them from the video.

What Are AI Agents?

Discover intelligent, task-driven AI agents that collaborate with your human workforce.

Teams Already Shipping Multilingual Video

2–3 days

Down From Ten Days

See how Smith+Nephew cut the turnaround on multilingual content from ten days to two or three.

30%

More Translation Output

Discover how Wunderman Thompson uses Smartcat to generate 30% more global content on the same budget.

50%

Higher Productivity at Half the Cost

Find out how expondo cut their global content production costs in half while making 50% productivity gains.

Video and Audio Tools That Run From the Same Transcript

Smartcat reads speech from the common video and audio containers — .mp4, .mov, .mkv, .mp3, .wav — plus .srt and .vtt subtitle files. WebM is not supported.

Video Subtitle Translator

Translate and re-time video subtitles across as many languages as you need.

AI Video Dubbing

Replace the original audio with an AI-dubbed track, timed to the translated transcript.

Training Video Creator

Build training videos from scratch with AI, then translate them in the same workspace.

AI Audio Translator

Translate audio-only files with AI transcription and translation, metered by transcript word count.

Voiceover Translator

Generate a translated narration track that plays over the original audio.

AI Voice Dubbing

Choose from 35 voice locales for a dubbed track that replaces the original audio.

One Upload. Every Market You Sell In.

Drop in the video, check the transcript once, and export subtitles, dubbing, or a voice-over in as many languages as you need. Free for 15 days with 15,000 Smartwords — full access to translation capabilities, no credit card.

Start From a Video, an Audio File, or a Subtitle File

Upload video, audio, or subtitle files and get translated subtitles, an AI-dubbed track, or a voice-over back.

Any Language Pair, Including the Hard Ones

Translate from and into any of 280+ languages, in any direction — including Chinese ↔ English, Japanese ↔ English, and Arabic ↔ English, where script direction and character density break tools built for European languages.

Rated by Smartcat users on G2

Video Translation — the Fine Print

How much does it cost to translate a video?

Every plan starts with a 15-day free trial with 15,000 Smartwords and full access to translation capabilities, no card required. Paid plans start at $1,200/year (Adapt) and scale to enterprise tiers.

Video is metered by transcript word count, not by video length. A 20-minute product walkthrough with sparse narration costs less than a 5-minute densely scripted explainer, because you pay for words, not minutes. Those words are charged in Smartwords, the usage credits included with every plan. Full breakdown on the pricing page.

Does it translate text that appears on screen?

No — and it is the most common surprise, which is why it has its own section above. Anything painted into the picture needs an edit in the video itself, or a re-render from source. If your training videos are slide-heavy, translate the deck in Smartcat first and re-record, rather than trying to fix it downstream.

Does the dubbed voice sound like a person?

For narration, explainers, training, and product walkthroughs — yes, close enough that most viewers don’t flag it. Where it still reads as synthetic is emotional performance: comedy timing, raised voices, sarcasm, and overlapping banter. If the video’s value is in the delivery rather than the information, use subtitles or a human voice-over from the Marketplace instead of AI dubbing.

What video and audio formats can I upload?

The container list is above, under “Video and Audio Tools That Run From the Same Transcript”. Two things it does not say there: dubbing and voice-over are available in 35 voice locales, while subtitles are available in all 280+ languages; and an .srt or .vtt file can be uploaded on its own, with no video at all.

Do I get subtitles burned into the video, or as a separate file?

Both — you choose at export. The trade-off is covered above; the short rule is sidecar for anything you may re-edit, burned-in for platforms that ignore caption files.

How many speakers can it handle in one video?

Smartcat separates audio by speaker and labels each one in the transcript, so a two-person interview or a panel comes back attributed rather than as one continuous block. Accuracy drops when speakers talk over each other, so panels usually need a pass in the editor.

Can I re-use a video’s translation after I edit the original?

Yes. Because the translation is stored against the transcript, re-uploading an edited cut only charges for speech that actually changed — a new intro, a corrected sentence — rather than the whole file again. Existing cues and their approved translations carry over.

Is my video confidential?

Smartcat is SOC 2 Type II compliant. Files are encrypted in transit and at rest, workspaces are isolated so no other account can reach your media, and access is per user with role-based rights. SSO through Azure AD, Okta and ADFS is available on higher tiers. Details on the security page.

Where does video translation fall short?

Three places.

  • Overlapping dialogue and crosstalk confuse speaker separation, so a panel discussion needs manual cleanup that a single-presenter video doesn’t.
  • Heavy background music or a strong accent lowers transcription accuracy, and a wrong transcript produces a confidently wrong translation — which is why the transcript-check step exists.
  • Dubbing is time-matched, not lip-matched: mouth movements won’t line up frame by frame, so tight close-ups on a speaker’s face are the worst case for dubbing and the best case for subtitles.

Is there an API for video translation?

Yes, on the Anticipate and Autonomous plans — the API is not part of the free trial. It accepts media files, returns translated subtitle files and audio tracks, and syncs translation memories, so video localization can run inside your own publishing pipeline rather than through the web interface.

My question isn’t answered here.

Book a demo — a 1:1 consultation with a media localization specialist, no commitment and no form queue.

Sources

  1. Martin, E. J. (2024, February 9). 2024 state of AI in the speech technology industry: AI is revolutionizing translation, dubbing, and subtitling. Speech Technology Magazine.

  1. MIT Technology Review Staff. "A New AI Translation System for Headphones Clones Multiple Voices Simultaneously." MIT Technology Review, 9 May 2025.

  2. Slator. "Translating Podcasts & Videos in Your Own Voice Is Possible Thanks to Translated." Slator, 2023.

  3. Abukins, S. (2024, December 18). Eight key insights from “Ai and the future of translation and interpretation.” Middlebury Institute of International Studies at Monterey.