← Notes

Adding captions to a video on a Mac

Most people watch with the sound off. Here is how captions actually get onto a video on macOS, what the automatic ones reliably get wrong, and when burning them in is the right call.

A large share of video on social platforms is watched muted, at least at first. Captions are not an accessibility afterthought any more; they are how most of your audience reads the first three seconds and decides whether to turn the sound on.

There are two genuinely different things people mean by "adding captions", and picking the wrong one wastes an afternoon.

Sidecar files versus burned-in text

A sidecar file — usually .srt or .vtt — sits next to the video. The player draws the text. YouTube, Vimeo and most broadcast workflows want this. It is searchable, the viewer can switch it off, and you can fix a typo without re-exporting.

Burned-in captions are pixels. They are part of the picture, they cannot be switched off, and fixing a typo means exporting again. But they survive everywhere: Instagram, TikTok, an autoplaying embed, a file someone drags into Keynote.

If you are publishing to a platform that accepts a subtitle file, use one. If you are posting to a feed, burn them in. Plenty of people do both from the same edit.

What macOS gives you out of the box

Less than you would hope. QuickTime Player will display a subtitle track if the file already has one, but it will not create captions. Preview is not in this business at all. iMovie has titles, which are not captions — laying out a full transcript as a stack of title cards is a well-known afternoon of misery.

So on a stock Mac the honest options are: write an SRT by hand in TextEdit, upload to a platform that transcribes for you and download its output, or use an editor that transcribes locally.

Writing SRT by hand is more reasonable than it sounds for short clips. The format is plain text:

1
00:00:01,400 --> 00:00:04,120
The first line of what somebody says.

2
00:00:04,300 --> 00:00:06,900
Then the next.

Numbering, a time range with a comma before the milliseconds, the text, a blank line. For a 30-second clip that is ten minutes of work. For anything longer you want a machine to take the first pass.

Automatic captions and their specific failures

Speech recognition on modern devices is good. It is not good at the things that matter most in your video, and the failures are predictable:

  • Proper nouns. Product names, your company, people's surnames. These are exactly the words you cannot afford to get wrong, and they are the first to go.
  • Numbers and units. "Two hundred and fifty megabytes" versus "250MB" versus "250 mb" — recognisers pick one, rarely the one you would write.
  • Jargon and acronyms. Anything domain-specific gets mapped to the nearest common word.
  • Overlapping speech and cross-talk. Two people on a call become one confident, wrong sentence.
  • Punctuation. A run-on sentence with no commas reads fine as audio and badly as text.

The practical consequence: treat automatic captions as a draft to correct, never as output. Budget a proofread pass at roughly a quarter of the video's runtime. If a tool does not let you edit the recognised text afterwards, it is not saving you work — it is committing you to its mistakes.

A workflow that holds up

  1. Lock the edit first. Captions are timed to the cut. If you move a clip afterwards, everything downstream drifts.
  2. Generate the first pass with whatever transcription you have.
  3. Read it against the audio, fixing names, numbers and sentence breaks. This is the step nobody skips twice.
  4. Check line length. Two lines, roughly forty characters each, is the comfortable ceiling. Longer than that and viewers stop reading mid-line.
  5. Check position. Platforms overlay their own interface along the bottom and right. Keep captions inside the middle two-thirds vertically, or they land under a username.
  6. Export. Sidecar for platforms that take one, burned-in for feeds.

On privacy, briefly

Transcription has to happen somewhere. On recent Apple hardware, speech recognition often runs on the device itself. When on-device recognition is not available for a language or a device, macOS may send audio to Apple for processing instead. That is a reasonable trade for most work and a poor one for a confidential interview. If the content is sensitive, check what your tool does before you feed it the file, and prefer a transcript you typed yourself.

Doing this in Garfi

Garfi Video Editor generates captions on the timeline and — this is the part that matters — lets you edit the recognised text in place, so the proofread pass happens where the video is rather than in a separate file you have to re-import. Captions can be styled and burned into the export, and the same timeline handles the trim, the music bed and the aspect ratio you need for a portrait feed.

It uses Apple's speech recognition, with the caveat above: recognition may require Apple processing when the on-device path is unavailable. There is no generative AI in the app, no account, and no watermark on what you export.

Keep reading