Captions are the styled, animated text creators add to hold attention — usually large, centred, and popping in word by word. They’re also the most awkward kind of text to remove, for a reason specific to how they’re designed.
First: can you avoid removing them?
Two checks worth doing before any processing.
Do you have the project file? If the video is yours, open the project, hide the caption layer, and re-export. Instant, lossless, perfect — nothing is reconstructed. This is by far the best outcome and people skip it constantly.
Are they soft-coded? If captions can be toggled off in a player, they’re a separate track and the frames were never altered. Switch them off. Only captions burned in before export need real removal. See remove subtitles from video.
If neither applies, you’re reconstructing.
Why creator captions are hard
Subtitles are easy: consistent position, consistent size, a predictable band near the bottom. Define the region once and it covers the whole clip.
Modern creator captions are the opposite:
- Word-by-word animation. Each word appears individually, so the text region changes several times a second.
- Scale changes. Words pop in larger and settle, or emphasise on the beat.
- Position drift. Captions move around the frame to avoid covering the subject.
- Colour highlighting. Karaoke-style tracking changes which word is emphasised.
- Heavy strokes and shadows that extend well past the letterforms.
A single fixed removal region misses most of this. You end up with a clip where some words vanished and others didn’t.
The shortcut that makes it manageable
Don’t track individual words.
Define one removal region covering the full area the caption ever occupies across a segment, then apply it to that whole segment.
You’re reconstructing more of the frame than strictly necessary. Where that extra area is ordinary background — a wall, a body, a room — it costs you nothing, because the reconstruction there is just as good.
This turns an intricate tracking problem into a simple one, and it’s the single most useful technique for this kind of text. Split into segments only where the caption zone genuinely moves to a different part of the frame.
When the shortcut doesn’t apply: if the caption travels across something that reconstructs badly — a face, a sign, a screen — then covering the full travel area means reconstructing all of that badly. In that case tighter per-segment regions are worth the extra effort.
What to expect
| Caption sits over… | Result |
|---|---|
| A person’s body or clothing | Invisible |
| Blurred background | Invisible |
| Wall, sky, room interior | Invisible |
| Busy scene with camera motion | Very good |
| A face | Risky — uncanny |
| A sign, screen or document | Gibberish |
| Static locked-off shot | Good, flaws more visible |
Captions usually sit over a person or a background rather than over other text, which is why they generally come off well despite being fiddly to select.
Don’t forget the stroke
Captions almost always carry a heavy outline or drop shadow so they stay readable over any footage. Those extend noticeably past the letterforms — often several pixels.
Select only the visible characters and you’ll leave a faint ghost outline of every word. Include the stroke, and add margin beyond it.
Steps
- Try re-exporting from the project file first.
- Check whether they’re soft-coded.
- Otherwise, define one region covering the full caption zone for each segment.
- Include strokes and shadows.
- Process, then review full screen at normal speed — look for shimmer and for ghosted words the region missed.
- Export at maximum bitrate.
The cross-posting angle
A common reason for removing captions: they were positioned for one platform’s layout and you’re posting to another.
TikTok, Reels and Shorts each place their interface differently. Captions positioned to clear TikTok’s UI can end up behind Instagram’s buttons — or sitting exactly where Instagram’s own caption goes, which reads as obviously recycled.
Removing and re-adding is one approach. Exporting a separate cut per platform from your editor is better, and avoids reconstruction entirely. See repost TikTok to Instagram Reels.
On iPhone
MarkOff handles captions with per-segment removal areas, so a caption zone that moves through the clip is covered throughout. Tap-to-select picks up the full text extent including strokes and shadows.
Video is processed on-device, and there’s no blur step — a blurred band where captions were is arguably worse than the captions. See remove watermark without blur.
Whose captions are they?
Removing captions from your own video is uncomplicated.
On someone else’s, captions are sometimes the work itself — translation, accessibility captioning, or commentary that took real effort. Stripping them to repost the clip removes that contribution along with the text.