Skip to content

Repository files navigation

Digital Human Studio

Upload a portrait and a script; get a talking avatar with synchronised lips and speech. Four lip-sync engines: an instant in-browser preview, Wav2Lip for the fastest local render, MuseTalk for photoreal rendering on your own machine, and HeyGen for cloud rendering that also moves the head. Pick one in the Renderer dropdown, press Render video.

Digital Human Studio — avatar, TTS and lip sync

uv venv --python 3.11 python/.venv
source python/.venv/bin/activate
uv pip install -r pyproject.toml       # or: uv sync
./python/download_models.sh            # ~3.5 GB of MuseTalk weights, once
PORT=8137 python python/app.py         # → http://localhost:8137

Backend is Python (FastAPI); the frontend is plain HTML/CSS/JS because that's what browsers run. There is no build step and no node_modules.

The three renderers

Live preview — in-browser, instant, free. The script is converted to phonemes, phonemes to visemes (15 mouth shapes), and that track is stretched onto the real audio's duration. An AnalyserNode then reads the audio's loudness every frame and gates the jaw: the viseme decides what shape the mouth makes, the waveform decides how far into it the mouth actually is — so pauses genuinely close the mouth. The photo's lower face is displaced in slices to fake a jaw, and a mouth cavity is composited in multiply so it darkens the real skin instead of pasting over it. The avatar does not blink: a painted eyelid over a photograph of an open eye reads as a flash, not a lid.

It is honest about what it is: geometry, not a face. Good enough to time a script and choose a voice. It will not fool anyone.

⚡ Wav2Lip — local, ~16s, the fastest real render. A 2020 GAN that synthesises the mouth region at 96×96 and composites it back, so the timing is excellent — it is still the sync benchmark others are measured against — but the mouth is soft once you scale it into a 1080p frame. It is pinned to old torch/librosa versions that fight with MuseTalk's, so it runs in its own virtualenv when one exists (WAV2LIP_DIR / WAV2LIP_PY override the paths).

✨ Render photoreal — MuseTalk, local, ~1 min. A neural model that inpaints the mouth region in the VAE's latent space, conditioned on Whisper audio features. It generates real mouth pixels — teeth, lips, shadow — matched to the subject's skin and lighting. Runs on Apple Silicon (MPS); ~120 frames in ~75s on an M4.

HeyGen — cloud, ~40s, costs credits. The only engine that also animates the head and blinks. Uses the v3 API (POST /v3/videos); v1/v2 are on the sunset path and reject current keys. Note v3 renders only avatars you own — stock avatars are rejected by the API, so the picker lists just your own avatars and photo avatars.

Note on the MuseTalk install. The official pipeline gets its face box from DWPose/mmpose, which drags in mmcv — no Apple Silicon wheels, painful source build. musetalk_engine.py computes the same box from MediaPipe instead (the same face_landmarker.task the browser uses), so mmcv, mmpose and tensorflow are never installed. The model receives the crop geometry it was trained on.

Deploying (Coolify, or any Docker host)

docker build -t digitalhuman .          # 877 MB, CPU
docker run -p 8137:8137 -e GEMINI_API_KEY=... -e HEYGEN_API_KEY=... digitalhuman

In Coolify: point it at the repo, use the Dockerfile (or docker-compose.yml) build pack, and set the keys as environment variables — the app reads os.environ exactly like it reads .env, so no secrets need to be committed. Mount a volume at /app/python/renders or rendered MP4s vanish on redeploy.

MuseTalk does not ship in that image, and that is deliberate. It needs a GPU: ~75s per clip on Apple MPS or CUDA, but many minutes on a plain CPU VPS — unusable, not merely slow. And its weights are 3.7 GB, which have no business inside a deploy image. So the CPU image reports musetalk: false from /api/config and the UI disables the Render photoreal button, instead of offering one that fails a minute later.

Host What you get
CPU VPS (typical Coolify box) Everything except MuseTalk. Photoreal video via HeyGen (cloud).
GPU host (NVIDIA) Build Dockerfile.gpu, mount a volume at /app/python/models, run download_models.sh once. Full local MuseTalk.

The browser preview, face detection, all TTS engines, voice cloning and HeyGen work fine on a plain CPU box. Ollama and Piper run outside the container — point at them with OLLAMA_URL / PIPER_URL (they default to localhost, which inside a container is the container itself).

Voices

Engine Offline Recordable Cloning Needs
Browser TTS yes no no nothing
Piper yes yes no a local Piper server
Gemini no yes no API key (24 voices, M/F)
ElevenLabs no yes yes API key
OpenAI no yes no API key
Your own audio yes yes — an MP3/WAV

Browser TTS speaks through the OS, so the page never receives the samples: it can preview but cannot be exported or sent to MuseTalk. Use any other engine for those.

Voice cloning is ElevenLabs-only. Drop a ~1-minute clip — video is fine: ffmpeg strips the audio locally and only the speech is uploaded, never your video. Gemini and OpenAI don't offer cloning.

API keys

None are needed for the core tool — the browser preview, face detection and MuseTalk all run locally with zero keys. Put any optional keys in .env and the server uses them automatically; the browser is told only which services are configured, never their values.

GEMINI_API_KEY=...        # or GOOGLE_API_KEY — TTS
ELEVENLABS_API_KEY=...    # TTS + the only voice cloning
OPENAI_API_KEY=...        # TTS
HEYGEN_API_KEY=...        # cloud talking-photo renderer

Can a local LLM do the voice?

No. Ollama has no audio output — Gemma, Qwen, Llama, any size, are text models and cannot produce speech. What they are good for is writing the script: ✦ Draft with Ollama sends your notes to your local model (default gemma4; any installed model works, e.g. qwen3.5). Pair it with Piper for offline speech and MuseTalk for the render, and the entire pipeline runs on your machine with no keys at all.

The rig (browser preview only)

Face detection is automatic: MediaPipe's face_landmarker runs in-browser on WASM (vendored under vendor/, so it works offline). Because a detector fed a whole 1792×2400 photo cannot see a face sixty pixels tall, the search sweeps progressively smaller overlapping tiles — which is what makes full-body and group shots work.

If the lips still look off, Adjust mouth & eyes lets you drag the four rig points (scroll to zoom, drag to pan). Frame to head crops the stage to a portrait around the rig. MuseTalk ignores all of this and finds the face itself.

Notes for hacking on it

The server sends Cache-Control: no-cache for js/ and css/ (and versions the script URLs), because a browser silently serving a stale app.js makes a fixed bug look unfixed — you end up debugging code that isn't running. vendor/ and samples/ still cache hard; they never change.

Files

Path Role
python/app.py FastAPI: static host, TTS/HeyGen proxies, cloning, Ollama, render jobs
python/musetalk_engine.py MuseTalk inference, MediaPipe instead of mmpose
js/visemes.js Grapheme→phoneme→viseme, timeline building
js/face.js The four-point rig, MediaPipe detection, tile search, editor
js/renderer.js Jaw warp, mouth compositing, head motion
js/heygen.js HeyGen v3 client (avatars, voices, render jobs)
js/tts.js The speech engines + cloning
js/app.js Playback loop, audio graph, recording, render jobs, UI

Known limits

  • The browser renderer is a talking-photo effect, not a face. Use Render photoreal for anything you'd actually show someone.
  • MuseTalk animates the mouth only — the head does not move or blink in its output. For head motion too, look at SadTalker or Hallo.
  • The grapheme-to-phoneme pass (browser preview) is rule-based English, not a forced aligner. It mispronounces names; loudness gating hides most of it. MuseTalk doesn't use it at all — it listens to the audio directly.
  • Neither local engine blinks. A painted eyelid over a photo of an open eye reads as a flash, not a lid — HeyGen is the only engine here that blinks for real.
  • Browser export is WebM (that's what MediaRecorder gives you). MuseTalk outputs MP4.
  • Only clone a voice, or animate a face, that you own or have explicit permission to use.

Missing weights? The app fetches them

MuseTalk (~3.5 GB) and Wav2Lip (~440 MB) are far too big to commit, so a fresh clone shows those engines greyed out. Rather than send you to a shell, pick the engine and the app offers a ⬇ Download weights button: it pulls the code from the projects' own GitHub repos and the model files from their published Hugging Face repos, reports real progress, and enables the engine without a restart. python/download_models.sh still works if you prefer a terminal.

A note on the progress bar

Wav2Lip and MuseTalk report a genuine percentage (frames done). HeyGen does not — its API only ever says processing or completed — so its bar is an elapsed-time estimate that eases toward 95% and stops there, reaching 100% only when the video actually arrives. A bar that sits at 100% while still spinning is a lie.

About

Turn a portrait and a script into a talking avatar. Three lip-sync engines: an instant in-browser preview, MuseTalk for photoreal rendering on your own machine (Apple MPS/CUDA), and HeyGen v3 for cloud renders that also move the head. TTS via Gemini, ElevenLabs (with voice cloning from a video clip), OpenAI or Piper. FastAPI backend, no build step.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages