How to Add Subtitles That Actually Get Watched
Most short-form video is watched on mute. A practical guide to subtitle styling, timing, placement and accessibility — and the mistakes that make captions unreadable.

Captions stopped being an accessibility add-on years ago and became the primary way short-form video is consumed. A large share of feed viewing happens with sound off, which means for many viewers your captions are the video.
That reframing changes what good captions look like. They're not a transcript pasted over footage; they're a design element with a job to do in the first second.
The three mistakes that kill readability
Almost every unwatchable caption falls into one of three traps.
Too many words on screen
A full sentence at once forces the viewer to read ahead instead of watching. Two to four words per beat, timed to the speech, keeps the eye moving with the audio rather than racing it.
No contrast
White text on a bright, busy background disappears. An outline, a shadow or a semi-opaque background box costs nothing and makes captions legible over any footage. If you only fix one thing, fix this.
Wrong position
Captions placed at the very bottom of a vertical frame sit behind the platform's own interface. Roughly the lower third to lower quarter is unusable on most feeds.
Styling that earns attention
Beyond legibility, a few choices reliably improve watch time.
- Per-word highlighting — colouring the word currently being spoken pulls the eye along and measurably holds attention
- Bold, heavy weights — thin fonts read as text; heavy fonts read as design
- Sentence case, not ALL CAPS — caps are harder to scan at speed, especially over four words
- One accent colour — a single highlight colour that matches your brand, used consistently, does more for recognition than varied styling
Timing is the invisible half
Well-styled captions with bad timing still fail. The fix is mostly about where the breaks land: split on natural speech pauses rather than at a fixed word count, so a phrase doesn't get cut across two cards.
Watch for captions that appear a beat before the word is spoken. Slightly late reads as natural; early reads as a spoiler and breaks the rhythm of a punchline.
Accessibility isn't a separate task
Captions produced for reach also serve viewers who are deaf or hard of hearing — but only if you keep a couple of things in mind that reach-oriented captions sometimes drop.
- Don't omit words for style. Paraphrasing to fit a design breaks captions for anyone relying on them
- Include meaningful non-speech audio where it carries information (laughter, a door, a phone)
- Keep contrast high — it's the same fix that helps everyone else
- Provide a real subtitle file (SRT/VTT) alongside burned-in captions where the platform supports it
Consistency beats per-clip polish
The single highest-leverage thing you can do is stop restyling captions per clip. A feed of clips that each look slightly different reads as amateur even when every individual clip is well made.
Set the font, the highlight colour, the outline and the position once, then apply the same treatment to everything in a series. The consistency is what makes a month of clips look like one body of work — and it removes a decision from every future clip.
Translated captions open a second audience
Once a clip is transcribed, translating the captions is close to free — and it reaches an audience that was never going to watch the original. For creators whose subject is language-independent (anything visual, technical or numeric), this is one of the cheapest reach multipliers available.
Two practical notes. Translate the captions, not the on-screen graphics, unless you're willing to maintain several versions of each clip. And post translated versions as separate uploads rather than stacking two languages on one clip — two sets of captions competing for the same frame is unreadable in both.
- Prioritise by where your analytics already show unexplained international viewing
- Keep the same caption styling across languages so the series still reads as one
- Check line lengths after translation — German and Spanish routinely run 20–30% longer than English and will overflow a box sized for the original
A workflow that scales
For volume, the order matters. Transcribe and caption as part of producing the clip rather than as a job you take a finished clip to — that way caption placement is chosen against the actual vertical framing, and there's no export-caption-reimport round trip.
Review captions by reading them rather than watching them. If the caption text reads well as text, it almost always plays well, and reading is far faster than scrubbing.
Font choice, concretely
Caption fonts are judged under conditions nothing else is: small, moving, over unpredictable backgrounds, read in under a second. That rules out most of what looks good in a document.
- Heavy sans-serif weights — bold or black. Regular weights read as body text and disappear over footage
- Wide letterforms — condensed fonts save space and cost legibility, which is the wrong trade at four words per card
- No thin serifs — the thin strokes vanish against busy frames, and outlines do not rescue them
- Consistent size — resist shrinking text to fit a long phrase. Break the phrase across two cards instead
Test against your worst frame
Pick the busiest, brightest frame in a typical clip and check the captions against that, not against a dark title card. If they read there, they read everywhere. This one habit removes most caption problems permanently.
Names, jargon and numbers are where transcription breaks
Automatic transcription is very good at ordinary speech and predictably weak at exactly the words that matter most: people names, company names, product names, technical terms and figures. Those are also the words a viewer will notice being wrong.
The efficient fix is not to proofread everything. It is to check only those categories.
- Scan for proper nouns — names of people, companies and products
- Check every number, especially prices, percentages and dates
- Check industry terms your audience would recognise as misspelled
- Leave ordinary sentences alone — that is where transcription is already reliable
Emphasis without clutter
Highlighting works because it is selective. Colour every important word and nothing stands out; the clip just looks noisy.
A workable rule: one emphasised word per caption card at most, and only when the word genuinely carries the point — a number, a name, the surprising word in the sentence. Per-word highlighting that follows the speech is different and can run throughout, because it is tracking rather than emphasising.
Avoid stacking effects. Outline plus shadow plus background box plus highlight colour on the same text is four treatments competing, and the result reads as less legible than any one of them alone.
Different footage needs different placement
One caption position does not suit every clip, and the exceptions are predictable.
Talking head
Captions in the lower-middle third, above the platform interface. This is the default and it works because there is nothing else competing in the frame.
Screen recording or slides
Move captions off the content — usually to the very top or into a band below the screen area. Captions over a spreadsheet make both unreadable.
B-roll and cutaways
Centre the captions. There is no face to avoid, and centred text over footage reads as deliberate rather than as an afterthought.
A two-minute QA pass
Before publishing, in this order. It is quick because each step checks one thing.
- Read the caption text as text, ignoring the video. If it reads well, the timing is almost certainly fine
- Check proper nouns and numbers — the categories transcription gets wrong
- Scrub to the brightest frame and confirm contrast holds
- Watch the first two seconds muted — that is the part that decides whether the clip gets watched at all
- Confirm nothing sits in the bottom quarter where the interface will cover it
Burned-in captions and caption files do different jobs
These get treated as alternatives and they are not. They solve different problems, and the right answer for most publishing is both.
Burned-in captions are drawn into the video frames. They cannot be turned off, they look exactly as you designed them everywhere, and they work on every platform including ones with no caption support. That reliability is why they are the default for short-form.
Caption files — SRT or VTT uploaded alongside the video — are text rather than pixels. They can be toggled, translated automatically by the platform, read by screen readers, and indexed by search. What they cannot do is guarantee how they look, because each platform renders them in its own style.
What to do in practice
Burn in captions for anything going to a vertical feed, because styling is part of the content there and you cannot rely on the viewer enabling anything. Additionally upload a caption file wherever the platform accepts one — YouTube in particular, where the file feeds search and automatic translation.
The one thing to avoid is both at once on a platform that renders its own captions over your burned-in ones. Two sets of text stacked in the same part of the frame is unreadable, and it happens more often than you would expect when a caption file is uploaded to a feed that displays captions by default.
Keep the transcript
Whatever you publish, keep the transcript itself. It is the source for translated captions later, for the written post that accompanies the clip, and for finding the clip again in six months. Discarding it after captioning is the most common avoidable waste in this whole workflow.
Key takeaways
- Two to four words per beat, not full sentences.
- Always use an outline, shadow or background box — contrast is the number one fix.
- Keep captions out of the bottom quarter of a vertical frame.
- Set caption styling once per series rather than per clip.
Frequently asked questions
Do subtitles actually improve reach?
They improve completion, which is what most feed algorithms reward. A large share of short-form viewing happens with sound off, so a clip without captions is unintelligible to a big part of its potential audience — those viewers scroll, and the completion rate falls.
Should captions be burned in or uploaded as a file?
Both, where you can. Burned-in captions guarantee the styling and work everywhere; an uploaded SRT or VTT is searchable, translatable and better for screen readers. Burned-in alone is fine for feeds that don't support caption files.
How many words should be on screen at once?
Two to four for short-form. Enough to form a phrase, few enough that the viewer reads at the pace of the speech rather than ahead of it.
Is ALL CAPS better for captions?
Generally no. Caps are harder to scan quickly, and at four-plus words the readability cost outweighs the emphasis. A heavy font weight gives you the impact without the penalty.
What font works best for video captions?
A heavy sans-serif in bold or black weight, with wide rather than condensed letterforms. Regular weights and thin serifs disappear over footage, and outlines do not rescue them. Test against the busiest, brightest frame in a typical clip rather than a dark title card.
How do I stop auto-transcription getting names wrong?
Do not proofread everything — check only the categories that break: proper nouns, numbers, and industry terms. Ordinary sentences are where automatic transcription is already reliable, so scanning for those four things catches nearly every error a viewer would notice.
Where should captions go on a screen recording?
Off the content — usually the very top of the frame or a band below the screen area. The default lower-middle position works for talking heads because nothing competes with it, but over a spreadsheet or slide it makes both the captions and the content unreadable.
Keep reading
- AI Subtitles: animated captions in your brand style
- Brand Kit: set your caption style once
- Instagram Reels workflow
- Clipo for educators
Turn one recording into a week of content
Clipo finds the clips, reframes them vertical, captions them, and drafts the posts and carousels that go with them — from a single upload.
Try Clipo Free