Captions are the cheapest upgrade available to your videos. Not the cheapest in effort — the cheapest in cost, because most of the tools that do the heavy lifting are free. Adding accurate captions to a video takes about 15 to 30 minutes once you have a repeatable workflow, and the payoff is measurable: longer watch time, more completion, more searchable text, and an audience that includes the hundreds of millions of people who watch with the sound off or can't hear your audio at all.
The problem is that most creators treat captions as a finishing touch they'll get to later. They export, upload, and let the platform's auto-captions do whatever they do — which is usually 70 to 85 percent accurate, unpunctuated, and full of brand names spelled wrong. That's worse than no captions in some cases, because it's the version of your video that gets indexed, translated, and quoted back to you by viewers.
TL;DR: Auto-generated captions are a starting point, not a finished product — expect 70–85% accuracy and always clean them up. Stick to 2 lines per caption, 32–42 characters per line, 1–6 seconds on screen, and a reading speed under 180 words per minute. Upload an SRT file when the platform accepts one (YouTube, Facebook, LinkedIn, X) and burn captions into the frame for TikTok and Reels. Caption style matters less than accuracy and consistency — contrast, weight, and safe zones do the visual work.
Why captions are a growth lever, not an accessibility checkbox
The accessibility case for captions is the most important one, and it's non-negotiable if you're publishing to a website covered by accessibility standards. Roughly 430 million people worldwide have disabling hearing loss, and if your content only works with sound, you've drawn a hard line around who can consume it. Captions bring in viewers who can't hear your audio, viewers in noisy places, viewers in quiet places, and viewers who don't speak your language fluently enough to follow speech at full speed.
The growth case is what gets most creators to actually do the work:
| What captions change | Why it happens |
|---|---|
| Watch time on silent autoplay | Viewers keep scrolling when there's no text to read — captions give a reason to stop |
| Completion rate | Text holds attention through moments where audio alone would lose it |
| Comprehension of accents and jargon | Readers parse unclear words faster than listeners do |
| Discoverability | Spoken words become indexed text on YouTube, TikTok, Instagram, and search engines |
| Reach into non-native audiences | Machine-translated captions open your video to new markets |
| Repurposing value | A clean transcript is a blog post, a newsletter section, a carousel, and a quote bank |
Autoplay is the key thing to sit with. Every major feed plays video silently and waits for the viewer to decide whether it earns sound. Your first two seconds are competing against a thumb. On-screen text is the only part of your video that works with zero audio — which makes it the only part that's guaranteed to be seen.
Captions vs subtitles: the distinction that matters
People use the terms interchangeably. Platforms do too, which is why you'll find "subtitle" buttons that actually turn on closed captions. Knowing the difference helps you build the right file.
| Closed captions | Subtitles | |
|---|---|---|
| Assumes the viewer | Cannot hear the audio | Can hear the audio but doesn't understand the language |
| Includes non-speech audio | Yes — [door slams], [upbeat music], speaker labels |
No — dialogue and narration only |
| Toggleable | Yes, on by default where legally required | Usually yes |
| Typical formats | SRT, VTT, TTML, SCC | SRT, VTT |
| Required for | Accessibility compliance (WCAG, and broadcast rules) | Localisation |
In practice, the file you build is nearly the same. If you're captioning properly, you include non-speech sound information, and if you're also providing translations, you export additional files from the same timed transcript. Build one good transcript and you get both.
The three ways to caption a video
| Method | Speed | Typical accuracy | Cost | Best for |
|---|---|---|---|---|
| Platform auto-captions | Instant | 70–85% | Free | A first draft you'll edit |
| AI transcription tools (Descript, Rev, Veed, Whisper-based apps) | 1–5 minutes | 85–95% | Free to ~$20/month | Most creators, most of the time |
| Human captioner | 12–48 hours | 99%+ | $1–$5 per minute | Client work, ads, regulated content |
The mistake is assuming you have to choose one. The workflow that works is a hybrid: let an AI tool generate the transcript, then spend ten minutes as an editor. AI tools are excellent at getting the words down and mediocre at the things that matter most — punctuation, proper nouns, technical vocabulary, and the difference between "their," "there," and "they're" when someone is talking fast.
The highest-value ten minutes you'll spend on any video is the one where you fix the auto-transcript. Not because viewers will consciously notice a misspelling, but because every downstream system — translation, search indexing, quote cards, transcripts on your site — inherits the errors.
Platform-by-platform caption specs
Caption requirements are close enough that one well-built SRT file covers most platforms, but the delivery mechanism differs. Some platforms want a sidecar file, some want burned-in text, and some only accept auto-captions you edit in-app.
| Platform | Sidecar file upload | Burned-in captions | Notes |
|---|---|---|---|
| YouTube | Yes (.srt, .vtt, .sbv) | Optional | Auto-captions can be edited line by line in Studio; captions are indexed for search |
| TikTok | No | Yes (in-app styles) | Auto-captions are editable before posting; in-app caption tool with style presets |
| Instagram Reels | No | Yes (in-app) | Auto-captions editable per line; text stickers work but aren't indexed the same way |
| Yes (.srt) | Yes | SRT upload from desktop is the cleanest route; auto-captions otherwise | |
| Yes (.srt) | Yes | Native upload is supported in the post editor — most creators skip it | |
| X / Twitter | Yes (.srt) | Yes | SRT upload on video posts; accuracy on platform auto-captions is weak |
| Your own website / embeds | Yes (.vtt, .srt) | Both | .vtt is the web standard; required for accessible HTML5 video players |
| Podcast video (Spotify) | Yes | No | Video podcasts accept caption files |
Two practical takeaways. First, always export both .srt and .vtt from your transcript tool — the second is needed for web embeds and podcast video, and it's the same content with a different header. Second, on vertical short-form platforms, caption style is part of the product: keep the in-app style consistent across your posts so your videos read as yours.
The professional workflow, step by step
Here's the sequence that produces broadcast-quality captions in under 30 minutes per video, assuming you recorded raw footage and nothing else.
Step 1: Generate a first-draft transcript. Run your footage through your tool of choice — Descript, Veed, Whisper-based apps, or the platform's own auto-caption. You want a timestamped transcript, not just a text dump, because timestamps are what let you export a real caption file later.
Step 2: Fix the words. Read the transcript against the audio at 1.5x speed and correct names, brands, jargon, numbers, and homophones. This is the step creators skip and the step that matters most. If your video mentions a tool, a client, or a price, that word needs to be right.
Step 3: Add punctuation and split into caption events. Auto-transcripts typically arrive as long unpunctuated blocks. Break them into caption events — the units that appear on screen one at a time. Each event should be a phrase or short sentence that can be read in one glance.
Step 4: Apply the formatting rules. These numbers come from broadcast captioning standards and they scale well to social video:
| Rule | Target | Why |
|---|---|---|
| Lines per caption | 1–2 | Three lines covers too much of the frame |
| Characters per line | 32–42 | Comfortable single-glance reading width |
| Duration on screen | 1–6 seconds | Shorter is unreadable, longer is distracting |
| Reading speed | 140–180 wpm | Beyond 180, viewers can't keep up |
| Gap between captions | ~2 frames | Prevents visual smearing between events |
| Minimum duration | 1 second | Sub-second captions flash and get missed |
Step 5: Export the files. Save an .srt for sidecar uploads and a .vtt for web embeds. If you have a proofreading pass to do, do it before export — regenerating after manual fixes means redoing the work.
Step 6: Quality check on the smallest screen you can find. Watch the video at phone size with the sound off. If a caption is unreadable at that scale, your line length is too long, your font is too thin, or you've placed it somewhere the platform's UI covers.
Styling captions so they help instead of hurt
Caption styling is where creators either elevate a video or make it look like a template. A few principles carry across every tool:
- Weight over decoration. A bold or semibold sans-serif beats a light font in every scenario. Thick strokes survive compression and read at small sizes.
- Contrast is the whole game. White text with a solid outline or a semi-transparent background plate works on any footage. White text alone disappears against bright skies and white walls.
- Respect safe zones. Vertical platforms reserve the bottom third for captions, usernames, and CTAs. Keep your text above that band, and centred horizontally.
- Stay consistent. Same font, same size, same position across your videos. Consistency is what makes captions feel like part of your brand rather than an afterthought.
- Karaoke styles for short-form, blocks for long-form. Word-by-word highlighting keeps attention on fast-paced vertical videos. For YouTube and interviews, clean two-line blocks are easier to follow over 20 minutes.
- Don't fight the platform UI. If a platform draws its own progress bar and captions, turn off your burned-in version or move it up. Two caption layers on screen is unreadable.
Captions are SEO: how spoken words become findable text
This is the part most creators never connect. Every platform that hosts video is also a search engine, and none of them can listen. They read.
On YouTube, your caption file is indexed and used as a ranking input alongside your title and description — a video where you say "beginner-friendly drone setup for under $300" can rank for that phrase only if the phrase exists as text. On TikTok and Instagram, on-screen text and captions feed the platforms' content understanding, which affects who sees your post in search and suggested feeds. On your own website, captions in a .vtt file give search engines text to associate with your embedded video, and they give visitors a way to scan content without pressing play.
There's a compounding effect worth noticing: a clean transcript is also the fastest raw material for everything else you publish. Pull three quotes for a carousel, lift a 400-word section into a newsletter, turn the same section into a blog post, and drop a timed version into a short video. Creators who caption consistently often find they've accidentally built a content library.
If you're driving viewers from that captioned video to a bio link, put the transcript, the translations, and the related posts behind a single link rather than five separate ones — a link-in-bio page like Biolinky keeps that tap simple and lets you see which one people actually open.
Accessibility: doing it properly
Captioning for growth is good. Captioning for accessibility is right, and a few extra habits make your captions genuinely usable rather than technically present.
Include non-speech audio in brackets — [laughter], [phone ringing], [soft music playing]. Identify speakers when more than one person talks, using either a name label or consistent colour coding. Never use all caps for emphasis; it's harder to read and can break screen readers. Avoid captions that appear for less than a second. If your video includes text on screen that isn't spoken, mention it in the audio or add it to the caption track, because otherwise that information exists for only one group of viewers.
Also worth checking: does your video player on your own site support captions at all? Many creator sites embed video with captions disabled by default or missing the track element entirely, which means the work you did never reaches the people who need it.
The mistakes that make captions worse than nothing
| Mistake | What it costs you |
|---|---|
| Shipping raw auto-captions | Wrong names and numbers that get indexed and translated as-is |
| Blocking the bottom third | Platform UI covers your text on TikTok and Reels |
| Three or more lines per caption | Viewers stop reading and watch the distraction instead |
| Captions under 1 second | Flashes nobody can read |
| No punctuation | Subtitle files become unreadable walls of words |
| One style across every format | Karaoke captions feel wrong on a 30-minute interview, block captions get lost on vertical video |
A repeatable 30-minute caption routine
Build this into your publishing checklist and it stops being a project. Export your transcript as soon as the edit locks. Fix names and numbers while the footage is still fresh in your head. Apply the two-line, 42-character, 180-words-per-minute rules mechanically rather than by feel. Export both file types, then upload the sidecar file on the platforms that accept one and burn captions in on the platforms that don't.
Fifteen videos in, you'll have a consistent look across every platform, a searchable archive of everything you've ever said, and an audience that includes people who would never have been able to watch you otherwise. That's a disproportionate return for the time it takes.
Captions aren't a detail you add after the video is finished. They're the version of your video most people will actually experience.
