← Notes

What automatic captions get wrong

Automatic captions are good enough to start from and rarely good enough to publish unchecked. Here are the mistakes speech recognition makes most often, why, and a fifteen-minute review routine that catches them.

Automatic captions have become genuinely good. For clear speech in a quiet room, most words come out right. That is exactly what makes the remaining errors dangerous: you skim the first few lines, they look fine, and you publish a video in which your product's name is spelled three different ways.

Speech recognition fails in predictable places. Knowing them makes checking fast.

The usual errors

Names. People, companies, products, places. Recognition works by predicting likely words, and a name it has not seen gets replaced by a common word that sounds similar. Your product name is the word most likely to be wrong — and the one viewers notice.

Jargon and acronyms. Technical terms, internal abbreviations, model numbers. API becomes a pea eye, SQL becomes sequel, HEVC becomes anything.

Homophones. their and there, your and you're, to, too and two. The recogniser chooses by context and sometimes chooses wrong.

Numbers. fifteen and fifty, 1.5 and 15, prices and dates. A wrong number in a caption can be worse than no caption.

Punctuation and sentence breaks. Spoken language has no full stops. Captions often run sentences together or break them in odd places, which changes how they read.

Overlapping speakers. When two people talk at once, recognition usually follows one and loses the other, or merges them.

Music and noise. Background music, especially with vocals, and room noise lower accuracy considerably. Sections with music under speech need extra checking.

Accents and fast speech. Accuracy drops for accents underrepresented in training data and for rapid delivery.

Filler words. Recognisers either transcribe every "um" and "you know", or drop them inconsistently.

Timing. Captions that appear slightly early or late, or stay on screen after the speaker has moved on. Most noticeable at cuts.

Why it matters

Captions are not decoration. Many people watch with the sound off, and for deaf and hard-of-hearing viewers they are the video. A wrong word is a wrong statement. A wrong number can be a wrong price.

Reading speed and line length

Beyond accuracy, captions must be readable:

  • Two lines at most on screen at once.
  • Roughly 32–42 characters per line is a common guideline, depending on the screen and platform.
  • A reading speed that viewers can keep up with — commonly cited guidelines are in the range of 15–20 characters per second for adult viewers.
  • Break lines at natural phrase boundaries, not in the middle of a name or between an adjective and its noun.
  • Keep captions clear of on-screen text and important parts of the picture.

Automatic captions often produce lines that are too long or change too fast. Splitting and merging caption segments is part of the review.

A fifteen-minute review routine

For a typical five-minute video:

  1. Make a word list first. Names, products, technical terms and numbers that appear in the video. Before reading anything, search the captions for each and fix it everywhere.
  2. Watch with sound, reading along. At normal speed, pause at each error. Do not skim the text without audio — you will read what you expect, not what is there.
  3. Check every number against what was actually said.
  4. Fix punctuation so each caption reads as a sentence or a clear phrase.
  5. Check timing at cuts, where captions most often drift.
  6. Watch once more without sound, as a caption-only viewer would. If anything is confusing, it needs rewording.

Reducing errors at the source

  • Record clean audio. A microphone close to the speaker, a quiet room, no music under the speech while recording.
  • Speak clearly and a little slower than conversation.
  • Say names clearly the first time they appear.
  • Add music after, not during, recording.

Privacy: where the recognition happens

Captioning tools send audio to different places. Some run speech recognition on your device; some upload the audio to a server. For a video of a private meeting or unreleased product, it is worth knowing which.

What Garfi does

Garfi Video Editor creates automatic captions using Apple speech recognition, and the captions are editable — you correct words, and the routine above works directly in the editor. Garfi does not use generative AI and runs no developer cloud backend. Apple's speech recognition works on device where supported; where on-device recognition is unavailable for a language or device, it may require Apple processing, which is stated plainly in the product's privacy notes.

Captions sit alongside text, clip audio, music and voice-over on a precise timeline, and the canvas formats — landscape, portrait, square and 4:5 — let you check caption placement in the shape you will publish.

Export is MOV or MP4 in H.264 or HEVC, up to 4K when supported, with no watermark. The download is free with an export trial you start yourself; a one-time purchase unlocks export permanently.

Keep reading