Text burned into a video is a slightly different problem from a watermark, and mostly an easier one. Understanding why explains both the good results and the one failure that catches everybody.
Why text is easier than a watermark
Sharp edges. Text has crisp, high-contrast boundaries. A tool can determine precisely where it starts and stops. A faint semi-transparent watermark spread diagonally across a frame is genuinely harder to isolate, even though it destroys less of the image.
High contrast. Captions are usually white on darker footage, or come with an outline or drop shadow to stay readable. That contrast makes detection reliable.
Predictable placement. Subtitles sit in a consistent band. Titles sit in thirds. Captions cluster in the lower-middle. Consistency makes the region easy to define once and reuse.
Compact. Text occupies a small share of the frame, leaving abundant surrounding context to reconstruct from — and coverage is the main constraint on inpainting quality.
The result: text over ordinary footage — a person talking, a room, a street, a landscape — usually comes off invisibly.
The one thing that never works
Text over other text.
If a caption sits across a shop sign, a document, a screen, a number plate or a subtitle in another language, the reconstruction will produce something that looks like writing and reads as nonsense.
The reason is structural, not a tool limitation. Inpainting reconstructs texture by predicting how visual content continues. Letterforms aren’t texture — there’s exactly one correct arrangement of pixels for a particular word, and no way to derive it from the surrounding image. The model generates something word-shaped, confidently, and it’s wrong.
No tool solves this. If the text underneath matters, the information is gone. See how AI watermark removal works.
The test that predicts every result: if you can’t read what’s under the text by looking at the frame yourself, the model can’t either. It has no more information than you do.
What to expect
| Text sits over… | Result |
|---|---|
| Sky, wall, grass, water | Invisible |
| Blurred background | Invisible |
| A person’s body or clothing | Very good |
| Busy street or scene | Very good |
| Moving footage with camera motion | Good |
| A face | Risky — uncanny |
| A sign, document or screen | Gibberish |
| Another subtitle track | Gibberish |
| Static locked-off shot | Good, but flaws are more visible |
That last row is worth knowing. On a static shot the patched region sits in the same place for the whole clip, so any imperfection has the full duration to be noticed. Camera motion hides small errors well.
Text that moves
Plenty of on-screen text doesn’t hold still:
- Animated captions that pop, slide or scale word by word — the TikTok and Reels house style
- Scrolling credits or tickers
- Karaoke-style highlighting where words change colour in sequence
- Titles that animate in and out
A single fixed removal region won’t cover any of these. What works is segmenting the timeline — split at each change and define a removal area per segment.
Animated word-by-word captions are the most laborious case, because the text region changes shape constantly. The practical shortcut is defining one region covering the full area the text ever occupies, rather than tracking each word. You reconstruct more of the frame than strictly necessary, and if that area is over ordinary background it costs nothing. Detail in remove moving text from video.
Captions, subtitles, or titles?
Different jobs, and they’re worth separating:
- Captions — the styled, animated text creators add for engagement. Usually centred, often large. See remove captions from video.
- Subtitles — translation or accessibility text, usually a consistent band near the bottom. Hardcoded ones can’t be switched off. See remove subtitles from video.
- Titles and lower thirds — names, locations, branding. Usually static within a segment, so among the easiest.
- Watermark text — a username or URL, typically semi-transparent. Handled as a watermark: see remove watermark from video.
Steps
- Start from the highest-quality source. Reconstruction quality depends on surrounding detail.
- Import and select the text, plus a small margin — text usually has an outline or drop shadow that extends past the visible letterforms, and missing it leaves a ghost.
- Segment if the text changes or moves.
- Process, then review full screen at normal speed. Check for shimmer in the patched region.
- Export at maximum bitrate.
The shimmer problem
The main thing separating good video tools from bad ones, and it’s invisible in a screenshot.
Each frame gets its own reconstruction, and each is plausible in a slightly different way. Pause and every frame looks perfect. Play it and the patched area crawls or pulses, because the generated fill shifts frame to frame.
Tools that handle video properly condition each frame on its neighbours so the fill stays stable. Tools that run a photo inpainter across every frame produce the shimmer.
Always judge in motion, full screen.
Is there a better route?
Sometimes the text doesn’t need removing at all.
If it’s your own video, re-export from your editor without the caption layer. Instant, lossless, perfect. Worth checking before anything else.
If the subtitles are soft-coded, they’re a separate track rather than burned in — just switch them off. Only hardcoded subtitles need removal.
If it’s a TikTok video, the share link fetches the master, though note that captions burned in by the creator before upload will still be there. Only the platform watermark comes off this way.
On iPhone
MarkOff handles text the same way as watermarks: tap it, the app finds the region boundary including any outline or shadow, and AI reconstructs the area. Text that moves is covered by defining separate removal areas per segment.
Video is processed on-device, and there’s no blur step — see remove watermark without blur.
All text-removal guides
| Topic | Guide |
|---|---|
| Free options | Remove text from video free |
| How the AI works | AI text remover |
| Photos and stills | Remove text from photo |
| Burned-in captions | Remove captions from video |
| Hardcoded subtitles | Remove subtitles from video |
| Animated or travelling text | Remove moving text from video |
Whose video is it?
Removing text from your own footage is uncomplicated.
Removing it from someone else’s is worth a thought — text is often the attribution. A creator’s username, a photographer’s credit, a translator’s subtitle work. Removing that and republishing strips the credit as well as the text.