Extract, isolate, clean and enhance any single speaker's voice from any source — a YouTube URL, a podcast, an interview — in one command. Chains a full open-source voice stack into a single zero-intervention pipeline.
Status: built for studio use — isolates reference voices feeding downstream lipsync and dubbing work.
yt-dlp → Demucs → pyannote → Whisper → MossFormer2 → Resemblyzer → SpeechScore
download stem diarize transcribe enhance speaker-match quality
- Download audio from a URL (or take a local file).
- Diarize with
pyannote— who speaks when. - Transcribe with Whisper and align transcript segments to speakers.
- Pick a speaker — interactively, by ID, or by matching a reference voice clip (
resemblyzer). - Cut & export just that speaker's audio.
- Enhance (optional) — vocal separation (Demucs) + speech enhancement (MossFormer2 / ClearerVoice).
- Score (optional) — objective speech-quality metrics.
# YouTube → interactive speaker selection → clean clips
python scripts/rhea_pipeline.py "https://youtube.com/watch?v=..." -o output/
# Local file, auto-select a speaker, full enhancement
python scripts/rhea_pipeline.py podcast.wav --speaker SPEAKER_01 --enhance
# Match a speaker by a reference voice clip, enhance + upscale
python scripts/rhea_pipeline.py interview.wav --reference my_voice.wav --enhance --upscale
# Batch several URLs, extract every speaker
python scripts/rhea_pipeline.py URL1 URL2 --all-speakers --enhance --score
# Quick mode — diarize + cut only, skip AI processing
python scripts/rhea_pipeline.py audio.wav --speaker SPEAKER_01 --quick| File | Role |
|---|---|
scripts/rhea_pipeline.py |
The full orchestrator (all stages, all flags) |
scripts/isolate_speaker.py |
Lean version: download → diarize → transcribe → cut |
scripts/vocals_extractor.py |
Demucs vocal-stem separation |
pip install -r requirements.txt
cp .env.example .env # add your HuggingFace token (pyannote gating)Needs ffmpeg on PATH. Speech enhancement uses the bundled ClearerVoice-Studio tools under tools/.