Gains:
- Ability to fit while preserving the meaning by applying the technical rules of the subtitle (line, character, CPS)
- Ability to verify proper name and number errors by voice while speeding up transcription and automatic timecoding with ASR
- Ability to manage the limits of artificial intelligence by taking into account visual context, tonal difference and lip sync in dubbing
Translating a movie, TV series, corporate video or online course is radically different from translating plain text: time, reading speed, on-screen image, voice and the lip movement of the characters are also constraints. In this unit, you will learn audiovisual translation (subtitling, dubbing, voice-over), the technical rules of subtitling, AI-supported transcription and time coding, and the quality pitfalls of this field. The aim is to work like an audiovisual translation expert who "fits the audience's eyes and ears" rather than "translating the text".
Basic concepts
Audiovisual translation (AVT) is the translation of content containing audio and video; The main forms are subtitling, dubbing and audio description. Subtitling is the reflection of the speech in written form on the screen. Dubbing is the re-dubbing of the original voice in the target language. Voice-over is when the original sound is muted and a translation is read over it (common in documentaries).
Two critical technical terms of subtitling: CPS (Characters Per Second) limits the speed of subtitling so that the viewer can read it comfortably; The generally accepted upper limit is around 17 characters per second (lower for children's content). Timecode is the start-end time that indicates when the subtitle will appear and disappear on the screen (such as 00:01:23,400).
Why is it difficult? Because it means fitting translation into subtitles. Two lines, ~42 characters per line, a minimum of ~1 second and a maximum of ~6 seconds of screen time, and the CPS limit — you need to comply with all of these while maintaining meaning and tone. That's why subtitle translation is often a matter of compression and rephrasing, not word-for-word translation.
Tip: Avoid the reflex of "translating every word" in subtitles. The viewer both watches and reads the image; Too long subtitles cannot be read. Finding the shortest natural expression that preserves meaning is the true mastery of the captioner.
The role of AI in audiovisual translation
AI significantly speeds up several steps in this area:
- Transcription: With ASR (Automatic Speech Recognition), the voice is automatically transcribed. This extracts the raw source of the subtitle in minutes.
- Automatic time coding: Tools divide the conversation into segments and produce draft time codes.
- Translation draft: The transcript text is translated.
- CPS/length control: The AI can mark which subtitle exceeds the reading speed and suggest shortening.
But beware: ASR errs on noise, accent, overlapping speech, proper names, and technical terms; cannot provide "audio description"; Who is talking may get confused. The AI translation doesn't see the image — it can't know when an object shown on the screen is called "that." Therefore, the human verifies the visual context and deciphering accuracy.
Caution: The most common mistake is to think that the ASR output is "decoded". ASR makes frequent mistakes, especially in proper names, numbers, brand names and technical terms; an incorrect transcription carries over to translation and subtitling. Be sure to verify critical areas by listening to the audio.
Subtitle format rules
Common rules (varies by platform):
- ~37-42 characters per line, 2 lines maximum.
- Duration: ~1-6 seconds; Avoid very short "flash" subtitles.
- CPS not to exceed ~17 (adult content).
- Break the sentence where it makes sense (do not leave "and" or "but" at the end of the line).
- If possible, do not skip the scene cut (shot change).
- Follow the platform rules for items such as swearing, songs, screen writing.
three mini cases
Case 1 — ASR + post-edit speeded up by 60%. A team first transcribed and time-coded a 45-minute documentary with ASR, then translated and post-edited it. Time reduced by 60% compared to transcribing + time coding from scratch. But all proper names and numbers were corrected by sound comparison.
Case 2 — CPS check saved readability. The first translation of an instructional video with subtitles was accurate, but the CPS averaged 24 — the audience couldn't keep up. A round of AI-powered shortening shortened subtitles by 30%, reducing CPS to 16; meaning was preserved, readability came.
Case 3 — Visual context error. In the ASR-based translation, a character pointed to a sign on the screen and said "read that"; The AI translated it as "read it", but the text on the sign was not translated, so the viewer could not see what to read. The captioner added a separate caption for the screen caption; The person who saw the image made up the difference.
Four copyable templates
1) Transcription (ASR) verification list:
Below is an automatic transcription of a video. Mark possible errors in the transcription: proper names, brand names, numbers, technical terms, ambiguous/incomplete sentences. List these as "verify by voice." Don't translate. Transcribe: [...]
2) Subtitle translation (CPS/length limited):
Translate the following transcription into [target language] subtitles. CONSTRAINT: each subtitle must be maximum 2 lines, line ~42 characters, CPS not to exceed ~17. SHORTEN and naturalize while preserving the meaning; do not translate word by word. Preserve time codes. Transcription + time codes: [...]
3) CPS/readability check:
Calculate approximate CPS based on duration and character count for each of the following subtitles. Mark any that exceed 17 and suggest a shorter alternative that preserves the meaning. Subtitles (text | duration seconds): [...]
4) Screen text / visual context check:
In the following dialogue, mark the places that may refer to the image (screen text, pointed object, sign). Say whether additional subtitles / explanations are required for these, assuming that the viewer cannot see the image. Dialogue: [...]
Weak prompt / Strong prompt
Weak: "Translate these subtitles." (No length, CPS, line rules; output will not fit on the screen, unreadable.)
Güçlü: "Translate this dialogue transcript into Turkish subtitles. Each subtitle must be a maximum of 2 lines, 42 characters per line; shorten it while preserving the meaning when necessary for reading speed. Give the spoken language in natural Turkish, soften the slang. Do not touch the time codes."
Difference: strong prompt gives the technical constraints of the subtitle (line, character, CPS, tone); the output actually approaches usable subtitles.
AVT formats comparison chart
format
What does
AI contribution
human critical point
subtitle
Transcribes the speech
ASR, translation, CPS
Fit, visual context
dubbing
Voiceover in target language
Translation draft
lip sync, acting
voice-over
Read translation above
Translation, timing
Natural flow, tone
Audio description
Describes the image with sound
draft text
Visual interpretation (human)
Common mistakes
- Translating the ASR transcript without verifying it. Proper name/number errors are carried over to the subtitle.
- Translating word by word and getting around CPS. The audience cannot read; fit must be made.
- Ignoring visual context. Screen text and pointed objects remain untranslated.
- Making line splitting pointless. It spoils readability.
- Flattening the tone/register. The characters' dialects and slang disappear.
Dubbing and lip sync
While the constraint of subtitling is reading speed, the constraint of dubbing is time and lips. There are three distinct types of sync in voice-over translation: lip-sync — where the translation matches the character's mouth movement, especially open vowels in close-ups; isochrony — the translation line being the same length as the original line, neither too short nor too long; gesture synchrony — matching what is said with the character's hand and arm movements. That's why the dubbing translator often writes a "catchy" line that preserves the meaning but with completely different words; Word fidelity is even less than subtitle here.
The AI can provide a raw translation draft and a time warning for dubbing ("this line is 40% longer than the original, shorten it"). New technologies can even clone sound and produce automatic dubbing; but acting, emotion, fine-tuning the character's breathing and lip-syncing are still the job of the human translator and voice actor. Automatic dubbing provides an outline; The artistic and synchronous decision is human's. Additionally, the consent of the person whose voice is cloned is an important ethical and legal issue.
In summary
Audiovisual translation is the art of fitting text within the constraints of the viewer's eye and ear: line/character/CPS boundaries in subtitling, lip-syncing in dubbing, visual context is the determining factor in all of them. AI speeds up transcription, timecoding, translation drafting, and CPS checking with ASR; But ASR is mistaken in proper names and noise, it cannot see the image and the fit-tone decisions are made by the human. The master subtitler does not translate word for word; finds the shortest, most natural expression that preserves the meaning.
Application task
Choose a short (2-3 minute) video clip. Transcribe with an ASR tool and compare and correct risky areas with the audio with a “transcription verification” template. Then translate with the CPS ~17 constraint with the "subtitle translation" template and shorten the subtitles that exceed the limit with the "CPS control". If there is a screen text or sign, apply a "visual context check".
checklist
- [ ] I verified the ASR decryption by voice for the specific name/number.
- [ ] I made the subtitles fit within the line/character/CPS rules.
- [ ] I shortened and naturalized it, not word for word, preserving the meaning.
- [ ] I took the visual context (onscreen text, sign) into account.
- [ ] I translated it to preserve the differences in tone and character.