Skip to content

Repository files navigation

VideoWipe logo

VideoWipe

Remove hardcoded (burned-in) subtitles, watermarks, and logos from video locally.
Preview the detection first, then wipe. CLI, Docker, or a local web UI.

Auto-detect burn-in text · review each track · inpaint the background · keep the original audio.
Files stay on your machine. No cloud account.

GitHub stars Latest release License: GPL-3.0 Python 3.10+ Runs locally Docker CPU and GPU

中文 · Site · Quick start · FAQ · Docker


Development status

v0.11.1 improves subtitle recovery during white flashes and cross-fades, preserves failed-frame safety, and verifies macOS operation without requiring a GUI-free build.

The reviewed white-flash, cross-fade, Chinese complex-background, and synthetic outlined-subtitle clips passed user acceptance at normal playback speed. This does not promise pixel-perfect reconstruction or complete removal on every video. The original user screenshot case remains unverified because its source and output videos are unavailable. macOS clean installation and operational checks passed; provider identity and native RECORD integrity checks remain enforced. Strong contrast recovery may require more detection work. See the release notes for verification results and metric tradeoffs.

See the delivery status and diagnostic tools. See v0.11.1 release notes for verification results and known limitations.

Remove hardcoded subtitles, watermarks, and logos locally

VideoWipe is a self-hosted hardcoded subtitle remover and video delogo tool. It finds burned-in / burn-in text, watermarks, logos, and on-screen timestamps, lets you review each track, then inpaints only the pixels you chose to erase.

People usually land here looking for:

  • remove hardcoded subtitles from video / remove burned-in subtitles
  • remove watermark from video / remove logo from video (delogo)
  • local, offline, self-hosted video text removal (no upload)
  • a scriptable alternative to online removers and one-click desktop apps

The cleaned MP4 keeps the original audio track.

If you want a Windows one-click desktop app, video-subtitle-remover (VSR) is the established option. Use VideoWipe when you need to see the detection before it erases, run it headless (CLI, Docker, worker), or keep source files off the cloud.

Not soft subtitles: VideoWipe does not strip .srt / .ass tracks. It removes text that is burned into the picture.

Quick start

Requirements: Python 3.10+, and either ONNX Runtime or PyTorch. Model weights download automatically on first run to ~/.videowipe/weights/.

VideoWipe is not on PyPI yet — install from source:

git clone https://github.com/KKenny0/videowipe.git
cd videowipe
pip install -e ".[onnx]"

# Auto-detect and remove hardcoded text overlays
videowipe clean input.mp4 -o result/

Prefer a browser UI (still local):

# Apple Silicon (M1/M2/M3/M4): automatic MPS
pip install -e ".[web,torch]"
# Other CPU environments
pip install -e ".[web,onnx]"
videowipe serve
# Open http://127.0.0.1:8000 — upload, preview tracks, download cleaned MP4

No Python? Use Docker.

Optional extras: .[torch] (PyTorch), .[ocr] (OCR text recognition), .[propainter] (adapter deps only; model not bundled).

Try a short cleanup first

Choose targets → try 3 seconds → review the full result.

Cleanup workspace: video and processing area on the left, removable targets and protection controls on the right

Recorded development UI (Chinese interface). Orange marks the processing area. This screenshot shows the workflow, not a visual-quality acceptance result.

  1. Choose what to remove. Inspect target cards; leave anything you want to keep unchecked.
  2. Try a short section. Use the suggested three-second window and compare the synchronized original and cleaned video. Adjust the targets or protect an area if needed.
  3. Process and review the whole video. Check the suggested review windows and download the result. A good trial does not guarantee a good full video.

Edits are saved, and unchanged trials can be reused. See workspace details for recovery, protection, cache limits, and SDK options.

Features

Hardcoded subtitle removal Burn-in / burned-in text, including multi-language clips
Watermark, logo, timestamp cleanup Persistent corner marks and on-screen clocks
Auto detection No hand-drawn mask required (you can still supply one)
Preview before erase Review each track: remove or keep, with time ranges
Local-first CLI, web UI, Docker, Python SDK — no cloud account
Original audio preserved Cleaned video keeps the source soundtrack
Multilingual detection Chinese, English, Korean, and more out of the box
Pluggable quality Default STTN; optional ProPainter or any external inpainting command
Batch / embeddable Reuse one engine across many videos in your own worker

Demo

Subtitle removal

Before After
Before: video frame with hardcoded Korean subtitle After: VideoWipe removed hardcoded Korean subtitle

Watch sample video

Auto-detection (no manual mask)

Built-in detector finds text regions across multilingual content:

VideoWipe auto-detecting Chinese hardcoded subtitles VideoWipe auto-detecting English hardcoded subtitles VideoWipe auto-detecting multilingual subtitles and watermark

Video Candidates Selected Types
Chinese drama 4 2 top subtitle, bottom subtitle
English clip 2 2 bottom subtitle
Music video (Korean + Burmese) 7 5 top watermark, bottom multilingual subtitles

Tested with --detect-mode balanced (50 sampled frames). Green boxes show regions selected for cleanup.

Local web UI

The workspace keeps video, target cards, and the next action together. Choose a file, inspect an automatically recommended trial, then play and download the full result in place. Target cards use the original frame where each target was observed; the player also supports a paused original/result comparison.

Who is it for?

VideoWipe is built for people who already know they want the burn-in gone, and do not want a black box to decide where.

  • Editors and archivists cleaning private or archive footage before a recut (respect copyright and platform terms)
  • Self-hosters who will not upload source video to an online remover
  • Developers who need detect → review → inpaint as a local CLI, Docker job, or Python worker

It is a weaker fit if you want a Windows .exe with no install steps. That is VSR's job.

Common use cases

  • Remove burned-in bottom subtitles from drama, lecture, or language-study clips
  • Clean a corner logo / watermark (delogo) before reuse
  • Strip on-screen timestamps or platform chrome
  • On multilingual videos, preview first, then remove only the tracks you care about
  • Run batch jobs on a worker without a desktop display

How it works

Three stages:

  1. Detection — Sample frames, find text regions, group them into stable tracks over time (multilingual, no manual mask required).
  2. Planning — Build a reviewable WipePlan: each track has type (subtitle / watermark / logo / timestamp), remove|keep action, time segments, and a precise mask.
  3. Inpainting — Only remove-track masks are applied per frame; default STTN fills from neighboring frames. Optional external models (e.g. ProPainter) plug in for higher quality.

For example, videowipe clean input.mp4 -o result/ runs this path through WipeEngine:

Input video → detect text → build WipePlan → apply per-frame removal masks
            → inpaint with STTN → write cleaned MP4 with original audio

The default CLI command continues straight to inpainting; it does not pause for review. Choose how to review or supply the removal areas:

Input / option Execution path
--preview Detect and save the plan and preview artifacts, then stop without loading the inpainting model. Edit the plan JSON before executing it.
--plan plan/wipe_plan.json Load and validate the reviewed plan, then inpaint without rerunning automatic detection.
--mask mask.png Skip automatic detection and planning; inpaint using the supplied mask. Mutually exclusive with --plan.

For an interactive CLI selection before inpainting, use --confirm. The local web UI provides the choose → trial → full-result workflow shown above.

VideoWipe vs alternatives

Online removers VSR Hand masks (AE / Resolve) VideoWipe
Privacy Upload required Local Local Local
Auto-detect burn-in text Sometimes Yes Manual Yes
Review before erase Rare Limited Manual Yes (WipePlan / web UI)
Headless / Docker / worker No Desktop GUI Desktop CLI, Docker, Python
Windows one-click .exe Sometimes Yes N/A No (source or Docker)
Embed in your pipeline Hard Hard Hard Python SDK + CLI

What VideoWipe is not

  • Not a soft subtitle editor (.srt / .ass)
  • Not a translation or re-caption product
  • Not a promise of perfect pixels on every shot — fast motion and complex textures are harder; preview first
  • Not a cloud SaaS — you run it yourself

CLI usage

# Recommended: auto-detect and remove overlays
videowipe clean input.mp4 -o result/

# Only subtitles, only bottom of frame
videowipe clean input.mp4 --target subtitle --region bottom -o result/

# Natural-language intent
videowipe clean input.mp4 --intent "remove bottom Chinese subtitles" -o result/

# Preview detection only (no inpainting)
videowipe clean input.mp4 --preview -o plan/

# Execute a reviewed plan
videowipe clean input.mp4 --plan plan/wipe_plan.json -o result/

# Manual mask when you want full control
videowipe clean input.mp4 -m mask.png -o result/

clean options

Flag Description Default
--target Target type (repeatable): subtitle, timestamp, watermark, logo auto-detect all
--region Screen region (repeatable): top, bottom, top-left, top-right, bottom-left, bottom-right, center all regions
--intent Natural-language cleanup intent —
--preview Write detection artifacts only (no inpainting) off
--plan Execute an existing wipe_plan.json (mutually exclusive with -m, --mask) —
--confirm Show detected targets and confirm before processing off
--detect-mode fast (24 frames); balanced (50) / sensitive (80) densely recheck detector-backed remove segments balanced
--ocr OCR: auto, off, rapidocr auto
--agent Local LLM CLI for intent-based selection (e.g. claude, codex) —
--external-command External inpainting command (bypasses built-in STTN) —
-g, --gap Frames per inpainting segment. 25 balances performance and quality; larger values add temporal context but grow compute and memory superlinearly 25
-d, --dual Side-by-side original in the output off
-m, --mask Mask image path auto
Legacy: detext command

detext auto-detects subtitles only. Prefer clean for new usage.

videowipe detext -v input.mp4 -o result/
videowipe detext -v input.mp4 -m mask.png -o result/
Flag Description Default
-v, --video Input video path required
-m, --mask Mask image path auto
-o, --output Output directory result/
-w, --weight Model weight path (PyTorch .pth/.pt, or ONNX prefix) auto
-g, --gap Frames per inpainting segment; larger values cost more compute and memory 25
-d, --dual Side-by-side original in the output off
--external-command External inpainting command —

Docker

CPU:

docker pull ghcr.io/kkenny0/videowipe:latest
docker run --rm -v "$(pwd)":/data ghcr.io/kkenny0/videowipe clean /data/input.mp4 -o /data/result/

GPU:

docker pull ghcr.io/kkenny0/videowipe:gpu
docker run --rm --gpus all -v "$(pwd)":/data ghcr.io/kkenny0/videowipe:gpu clean /data/input.mp4 -o /data/result/

Or use the wrapper (auto-picks CPU/GPU):

./scripts/docker-videowipe.sh clean input.mp4 -o result/
Image Size GPU Notes
ghcr.io/kkenny0/videowipe:latest ~480 MB No CPU only, smallest
ghcr.io/kkenny0/videowipe:gpu ~1.4 GB Yes Prebuilt GPU image
videowipe:gpu ~1.4 GB Yes Local build tag

Floating latest / gpu track the current main build. Pin versioned tags from GHCR after a successful release.

Build from source
docker build --target runtime-cpu -t videowipe:latest .
docker build --target runtime-gpu --build-arg VARIANT=gpu -t videowipe:gpu .

GPU image needs NVIDIA runtime for CUDA; otherwise ONNX Runtime falls back to CPU.

docker run --rm -v "$(pwd)":/data videowipe:latest clean /data/input.mp4 -o /data/result/
docker run --rm --gpus all -v "$(pwd)":/data videowipe:gpu clean /data/input.mp4 -o /data/result/

Higher-quality inpainting (optional)

Default backend is STTN (works on CPU via ONNX). For tougher shots, plug in an external model.

ProPainter is validated as a higher-quality option:

git clone https://github.com/sczhou/ProPainter.git ../models/ProPainter

videowipe clean input.mp4 --model propainter --propainter-dir ../models/ProPainter
# equivalent:
videowipe clean input.mp4 --external-command "python scripts/propainter_wipe.py"

Note: ProPainter needs a GPU with ~16GB VRAM for 480p, and uses NTU S-Lab License 1.0 (non-commercial).

Quality comparison: ProPainter vs STTN

Multilingual music video (Korean + Burmese subtitles, 852×480, 10s). Same mask for both.

Original ProPainter (GPU fp16) STTN (CPU ONNX)
Original frame with hardcoded subtitles ProPainter inpainting result STTN inpainting result

Python API (for pipelines)

On supported Macs, Torch automatically uses MPS (CUDA → MPS → CPU). Explicit device="cpu" remains available through the Python API. Temporal plans skip crop bands with no active removal pixels in the emitted segment while retaining full context for active bands.

The same engine powers CLI, web, and Docker. Use it when you want batch jobs or a custom worker.

from videowipe import remove_text

# Mask optional — regions auto-detected if omitted
remove_text(video="input.mp4", output="result/")

Full clean pipeline with target selection:

from videowipe import WipeEngine

engine = WipeEngine(task="clean", detect_mode="balanced", ocr="auto")
engine.process(
    video="input.mp4",
    targets=["subtitle", "watermark"],
    regions=["bottom"],
    intent="remove Chinese subtitles and logo watermark",
    output="result/",
)
engine.cleanup()

Batch with a long-lived engine (model stays loaded):

from videowipe import CancellationToken, WipeEngine, WipeRequest

with WipeEngine(task="detext") as engine:
    result = engine.run(
        WipeRequest(video="clip1.mp4", mask="mask.png", output_dir="result/clip1"),
        cancellation=CancellationToken(),
    )
    print(result.output_path, result.backend, result.timings)

See examples/batch_worker.py and examples/custom_inpainter.py.

Review and edit the WipePlan

Generate a plan without loading the inpainting model, edit track action / segments in JSON (do not edit the sidecar .npz masks), then execute:

from videowipe import CancellationToken, WipeEngine, WipeRequest

engine = WipeEngine(task="clean")
plan = engine.plan(
    WipeRequest(video="input.mp4", output_dir="plan/"),
    on_candidates=lambda snapshot: print(snapshot["default_selected_ids"]),
    cancellation=CancellationToken(),
)
engine.cleanup()

with WipeEngine(task="clean") as engine:
    result = engine.run(WipeRequest(
        video="input.mp4",
        output_dir="result/",
        plan="plan/wipe_plan.json",
    ))

Optional on_candidates receives an independent metadata snapshot after candidate evidence is saved, before temporal refinement. Use it to show an early review; wait for plan() to return before execution. Exceptions from the callback propagate.

videowipe clean input.mp4 --preview -o plan/
videowipe clean input.mp4 --plan plan/wipe_plan.json -o result/

The plan is ordinary JSON — no LLM or cloud required. A plan is bound to its source video.

FAQ

Does VideoWipe remove soft subtitles (.srt / .ass)?
No. Soft tracks are separate files or streams. VideoWipe removes hardcoded / burn-in text painted into the video frames.

Do my videos leave my computer?
Default path is fully local (CLI, web UI on 127.0.0.1, Docker with a bind mount). Nothing is uploaded to a VideoWipe cloud.

Is the original audio kept?
Yes. Downloaded / output MP4 keeps the source audio track.

Do I need a GPU?
No. CPU + ONNX Runtime works. GPU images and PyTorch/ProPainter paths are faster or higher quality when available.

Is it on PyPI?
Not yet. Install from this repo or pull the Docker image.

Hardcoded subtitle vs watermark — can I choose?
Yes. Use --target subtitle / --target watermark / --region bottom, natural-language --intent, or the web UI track toggles.

STTN or ProPainter?
STTN is the default (lighter, CPU-friendly). ProPainter often looks better on hard regions but needs more VRAM and a non-commercial license for that model.

Will cleanup always look perfect?
No tool can guarantee that. Fast motion, thin textures, and semi-transparent marks are harder. Use preview / confirm before a long run.

Can I plug this into my own product or worker?
Yes. Treat VideoWipe as an embeddable engine: one WipeEngine, stable request/result types, optional custom inpainters. Product boundary is detect → plan → inpaint — not a full NLE or cloud studio.

How is this different from video-subtitle-remover (VSR)?
VSR is a desktop GUI with a Windows package. VideoWipe is preview-first and built to run from CLI, Docker, a local web UI, or your own worker. Same job (hardcoded subtitle / watermark removal), different way to operate it.

Support

If VideoWipe saves you time on subtitle, watermark, or overlay cleanup:

https://kkenny0.github.io/support/

Support helps maintain model packaging, Docker images, detection tuning, and docs.

Related projects

Project Relationship
Video-Auto-Wipe Ancestor this repo derives from
video-subtitle-remover (VSR) Popular desktop GUI for the same job
STTN Default inpainting model
OnnxOCR Built-in text detection
ProPainter Optional higher-quality inpainter (not bundled)
InpaintDelogo AviSynth+ delogo plugin, different stack

Credits

Built on STTN and the original Video-Auto-Wipe. Built-in text detection from OnnxOCR.

License

GNU General Public License v3.0 — see LICENSE.

This repository derives from GPL-3.0-licensed Video-Auto-Wipe. If you distribute VideoWipe or a combined work, review the GPL-3.0 obligations for your distribution model.

Star History

Star History Chart

About

Preview-first local remover for hardcoded subtitles, watermarks & logos. CLI, Docker, web UI. / 本地擦硬字幕和水印:先预览再擦,支持命令行、Docker、网页。

Topics

Resources

Stars

55 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages