Upload a portrait and a script; get a talking avatar with synchronised lips and speech. Four lip-sync engines: an instant in-browser preview, Wav2Lip for the fastest local render, MuseTalk for photoreal rendering on your own machine, and HeyGen for cloud rendering that also moves the head. Pick one in the Renderer dropdown, press Render video.
uv venv --python 3.11 python/.venv
source python/.venv/bin/activate
uv pip install -r pyproject.toml # or: uv sync
./python/download_models.sh # ~3.5 GB of MuseTalk weights, once
PORT=8137 python python/app.py # → http://localhost:8137Backend is Python (FastAPI); the frontend is plain HTML/CSS/JS because that's what
browsers run. There is no build step and no node_modules.
Live preview — in-browser, instant, free. The script is converted to phonemes,
phonemes to visemes (15 mouth shapes), and that track is stretched onto the real audio's
duration. An AnalyserNode then reads the audio's loudness every frame and gates the jaw:
the viseme decides what shape the mouth makes, the waveform decides how far into it the
mouth actually is — so pauses genuinely close the mouth. The photo's lower face is
displaced in slices to fake a jaw, and a mouth cavity is composited in multiply so it
darkens the real skin instead of pasting over it. The avatar does not blink: a painted
eyelid over a photograph of an open eye reads as a flash, not a lid.
It is honest about what it is: geometry, not a face. Good enough to time a script and choose a voice. It will not fool anyone.
⚡ Wav2Lip — local, ~16s, the fastest real render. A 2020 GAN that synthesises the
mouth region at 96×96 and composites it back, so the timing is excellent — it is still the
sync benchmark others are measured against — but the mouth is soft once you scale it into a
1080p frame. It is pinned to old torch/librosa versions that fight with MuseTalk's, so it runs
in its own virtualenv when one exists (WAV2LIP_DIR / WAV2LIP_PY override the paths).
✨ Render photoreal — MuseTalk, local, ~1 min. A neural model that inpaints the mouth region in the VAE's latent space, conditioned on Whisper audio features. It generates real mouth pixels — teeth, lips, shadow — matched to the subject's skin and lighting. Runs on Apple Silicon (MPS); ~120 frames in ~75s on an M4.
HeyGen — cloud, ~40s, costs credits. The only engine that also animates the head and
blinks. Uses the v3 API (POST /v3/videos); v1/v2 are on the sunset path and reject
current keys. Note v3 renders only avatars you own — stock avatars are rejected by the
API, so the picker lists just your own avatars and photo avatars.
Note on the MuseTalk install. The official pipeline gets its face box from DWPose/
mmpose, which drags inmmcv— no Apple Silicon wheels, painful source build.musetalk_engine.pycomputes the same box from MediaPipe instead (the sameface_landmarker.taskthe browser uses), sommcv,mmposeandtensorfloware never installed. The model receives the crop geometry it was trained on.
docker build -t digitalhuman . # 877 MB, CPU
docker run -p 8137:8137 -e GEMINI_API_KEY=... -e HEYGEN_API_KEY=... digitalhumanIn Coolify: point it at the repo, use the Dockerfile (or docker-compose.yml) build
pack, and set the keys as environment variables — the app reads os.environ
exactly like it reads .env, so no secrets need to be committed. Mount a volume at
/app/python/renders or rendered MP4s vanish on redeploy.
MuseTalk does not ship in that image, and that is deliberate. It needs a GPU: ~75s
per clip on Apple MPS or CUDA, but many minutes on a plain CPU VPS — unusable, not
merely slow. And its weights are 3.7 GB, which have no business inside a deploy image.
So the CPU image reports musetalk: false from /api/config and the UI disables the
Render photoreal button, instead of offering one that fails a minute later.
| Host | What you get |
|---|---|
| CPU VPS (typical Coolify box) | Everything except MuseTalk. Photoreal video via HeyGen (cloud). |
| GPU host (NVIDIA) | Build Dockerfile.gpu, mount a volume at /app/python/models, run download_models.sh once. Full local MuseTalk. |
The browser preview, face detection, all TTS engines, voice cloning and HeyGen work
fine on a plain CPU box. Ollama and Piper run outside the container — point at them
with OLLAMA_URL / PIPER_URL (they default to localhost, which inside a container is
the container itself).
| Engine | Offline | Recordable | Cloning | Needs |
|---|---|---|---|---|
| Browser TTS | yes | no | no | nothing |
| Piper | yes | yes | no | a local Piper server |
| Gemini | no | yes | no | API key (24 voices, M/F) |
| ElevenLabs | no | yes | yes | API key |
| OpenAI | no | yes | no | API key |
| Your own audio | yes | yes | — | an MP3/WAV |
Browser TTS speaks through the OS, so the page never receives the samples: it can preview but cannot be exported or sent to MuseTalk. Use any other engine for those.
Voice cloning is ElevenLabs-only. Drop a ~1-minute clip — video is fine: ffmpeg
strips the audio locally and only the speech is uploaded, never your video. Gemini and
OpenAI don't offer cloning.
None are needed for the core tool — the browser preview, face detection and MuseTalk
all run locally with zero keys. Put any optional keys in .env and the server uses them
automatically; the browser is told only which services are configured, never their values.
GEMINI_API_KEY=... # or GOOGLE_API_KEY — TTS
ELEVENLABS_API_KEY=... # TTS + the only voice cloning
OPENAI_API_KEY=... # TTS
HEYGEN_API_KEY=... # cloud talking-photo renderer
No. Ollama has no audio output — Gemma, Qwen, Llama, any size, are text models and
cannot produce speech. What they are good for is writing the script: ✦ Draft with
Ollama sends your notes to your local model (default gemma4; any installed model works,
e.g. qwen3.5). Pair it with Piper for offline speech and MuseTalk for the render, and
the entire pipeline runs on your machine with no keys at all.
Face detection is automatic: MediaPipe's face_landmarker runs in-browser on WASM
(vendored under vendor/, so it works offline). Because a detector fed a whole
1792×2400 photo cannot see a face sixty pixels tall, the search sweeps progressively
smaller overlapping tiles — which is what makes full-body and group shots work.
If the lips still look off, Adjust mouth & eyes lets you drag the four rig points (scroll to zoom, drag to pan). Frame to head crops the stage to a portrait around the rig. MuseTalk ignores all of this and finds the face itself.
The server sends Cache-Control: no-cache for js/ and css/ (and versions the script
URLs), because a browser silently serving a stale app.js makes a fixed bug look unfixed
— you end up debugging code that isn't running. vendor/ and samples/ still cache hard;
they never change.
| Path | Role |
|---|---|
| python/app.py | FastAPI: static host, TTS/HeyGen proxies, cloning, Ollama, render jobs |
| python/musetalk_engine.py | MuseTalk inference, MediaPipe instead of mmpose |
| js/visemes.js | Grapheme→phoneme→viseme, timeline building |
| js/face.js | The four-point rig, MediaPipe detection, tile search, editor |
| js/renderer.js | Jaw warp, mouth compositing, head motion |
| js/heygen.js | HeyGen v3 client (avatars, voices, render jobs) |
| js/tts.js | The speech engines + cloning |
| js/app.js | Playback loop, audio graph, recording, render jobs, UI |
- The browser renderer is a talking-photo effect, not a face. Use Render photoreal for anything you'd actually show someone.
- MuseTalk animates the mouth only — the head does not move or blink in its output. For head motion too, look at SadTalker or Hallo.
- The grapheme-to-phoneme pass (browser preview) is rule-based English, not a forced aligner. It mispronounces names; loudness gating hides most of it. MuseTalk doesn't use it at all — it listens to the audio directly.
- Neither local engine blinks. A painted eyelid over a photo of an open eye reads as a flash, not a lid — HeyGen is the only engine here that blinks for real.
- Browser export is WebM (that's what
MediaRecordergives you). MuseTalk outputs MP4. - Only clone a voice, or animate a face, that you own or have explicit permission to use.
MuseTalk (~3.5 GB) and Wav2Lip (~440 MB) are far too big to commit, so a fresh clone shows
those engines greyed out. Rather than send you to a shell, pick the engine and the app offers
a ⬇ Download weights button: it pulls the code from the projects' own GitHub repos and the
model files from their published Hugging Face repos, reports real progress, and enables the
engine without a restart. python/download_models.sh still works if you prefer a terminal.
Wav2Lip and MuseTalk report a genuine percentage (frames done). HeyGen does not — its API only ever says processing or completed — so its bar is an elapsed-time estimate that eases toward 95% and stops there, reaching 100% only when the video actually arrives. A bar that sits at 100% while still spinning is a lie.
