|
| 1 | +# Model Manager |
| 2 | + |
| 3 | +`tools/model_manager.py` downloads or assembles supported model packages into the |
| 4 | +framework's expected `models/` layout. |
| 5 | + |
| 6 | +This tool is still useful for safetensors-based packages and a few composite model |
| 7 | +layouts, but it is gradually becoming a legacy path as audio.cpp moves toward |
| 8 | +standalone GGUF packages. |
| 9 | + |
| 10 | +If a model has a ready-to-use GGUF package, prefer that route first. |
| 11 | + |
| 12 | +## GGUF Downloads |
| 13 | + |
| 14 | +Ready-to-use GGUF packages are published here: |
| 15 | + |
| 16 | +- Core released models: [audio-cpp/audio.cpp-gguf](https://huggingface.co/audio-cpp/audio.cpp-gguf) |
| 17 | +- Community OuteTTS package: [mirek190/audio.cpp](https://huggingface.co/mirek190/audio.cpp/tree/main/Text%20to%20audio%20(TTS)) |
| 18 | + |
| 19 | +For support status and tested precision coverage, see the [GGUF guide](gguf.md). |
| 20 | +For measured 16-bit vs Q8 speed and peak-VRAM results, see the |
| 21 | +[Q8 performance report](reports/gguf_q8_performance.md). |
| 22 | + |
| 23 | +## Dependencies |
| 24 | + |
| 25 | +- Python 3 |
| 26 | +- `torch` |
| 27 | +- `safetensors` |
| 28 | +- `PyYAML` |
| 29 | +- Network access to the upstream model source |
| 30 | + |
| 31 | +## Commands |
| 32 | + |
| 33 | +- `list` shows the available package ids |
| 34 | +- `list --json` prints a machine-readable package catalog |
| 35 | +- `info` shows the target layout, required files, and install source for one package |
| 36 | +- `info <package> --json` prints machine-readable package details |
| 37 | +- `install` downloads or converts one package into a models root |
| 38 | + |
| 39 | +The runtime loader catalog is also available from: |
| 40 | + |
| 41 | +```bash |
| 42 | +audiocpp_cli --list-loaders --json |
| 43 | +``` |
| 44 | + |
| 45 | +## Quick Start |
| 46 | + |
| 47 | +List installable packages: |
| 48 | + |
| 49 | +```bash |
| 50 | +python3 tools/model_manager.py list |
| 51 | +``` |
| 52 | + |
| 53 | +Inspect one package: |
| 54 | + |
| 55 | +```bash |
| 56 | +python3 tools/model_manager.py info qwen3_tts_1_7b_base |
| 57 | +``` |
| 58 | + |
| 59 | +Install into the default `models/` directory: |
| 60 | + |
| 61 | +```bash |
| 62 | +python3 tools/model_manager.py install qwen3_tts_1_7b_base |
| 63 | +``` |
| 64 | + |
| 65 | +Install into a custom models root: |
| 66 | + |
| 67 | +```bash |
| 68 | +python3 tools/model_manager.py install vevo2 --models-root /path/to/models |
| 69 | +``` |
| 70 | + |
| 71 | +Overwrite an existing install: |
| 72 | + |
| 73 | +```bash |
| 74 | +python3 tools/model_manager.py install pocket_tts --overwrite |
| 75 | +``` |
| 76 | + |
| 77 | +Install a converter-style package that needs a source file: |
| 78 | + |
| 79 | +```bash |
| 80 | +python3 tools/model_manager.py info voxcpm2_audiovae |
| 81 | +python3 tools/model_manager.py install voxcpm2_audiovae --source-file models/VoxCPM2/audiovae.pth --models-root models --overwrite |
| 82 | +``` |
| 83 | + |
| 84 | +## Package Notes |
| 85 | + |
| 86 | +For shared audio.cpp GGUF packages, the model manager installs the default `q8_0` |
| 87 | +GGUF. Other precision variants can be downloaded directly from |
| 88 | +[audio-cpp/audio.cpp-gguf](https://huggingface.co/audio-cpp/audio.cpp-gguf). |
| 89 | + |
| 90 | +`Yes` means Hugging Face has a ready-to-use repo that the framework can download |
| 91 | +as-is. `No` means the tool must assemble, convert, or post-process files before the |
| 92 | +framework can use them. |
| 93 | + |
| 94 | +Packages whose loaders are not registered in the current release tree are listed as |
| 95 | +**Unavailable**; see [loader/catalog sync notes](maintainers/loader_and_catalog.md). |
| 96 | + |
| 97 | +| Package id | Model | HF ready-to-use repo | |
| 98 | +|---|---|---| |
| 99 | +| `ace_step` | ACE-Step 1.5 Turbo/Base | No | |
| 100 | +| `chatterbox` | Chatterbox | **Yes** | |
| 101 | +| `citrinet_asr` | Citrinet ASR converted layout | No | |
| 102 | +| `fish_audio_s2_pro` | Fish Audio S2 Pro GGUF Q8_0 | **Yes** | |
| 103 | +| `heartmula` | HeartMuLa | No | |
| 104 | +| `higgs_audio_stt` | Higgs Audio STT | No | |
| 105 | +| `higgs_audio_v3_tts_4b` | Higgs Audio v3 TTS 4B GGUF Q8_0 | **Yes** | |
| 106 | +| `htdemucs` | HTDemucs | No | |
| 107 | +| `hviske_asr` | Hviske ASR | **Yes** | |
| 108 | +| `irodori_tts_500m_v3` | Irodori-TTS 500M v3 | No | |
| 109 | +| `irodori_tts_600m_v3_voice_design` | Irodori-TTS 600M v3 VoiceDesign | No | |
| 110 | +| `index_tts2` | IndexTTS-2 | **Yes** | |
| 111 | +| `mel_band_roformer` | Mel-Band RoFormer MLX | **Yes** | |
| 112 | +| `miocodec_25hz_44k_v2` | MioCodec 25Hz 44.1kHz v2 | No | |
| 113 | +| `miotts_1_7b` | MioTTS 1.7B | No | |
| 114 | +| `moss_audio_tokenizer_nano` | MOSS Audio Tokenizer Nano | No | |
| 115 | +| `moss_audio_tokenizer_v2` | MOSS Audio Tokenizer v2 | No | |
| 116 | +| `moss_tts_nano_100m` | MOSS-TTS-Nano 100M | No | |
| 117 | +| `moss_tts_nano_100m_model` | MOSS-TTS-Nano 100M model subcomponent | No | |
| 118 | +| `moss_tts_local_v1_5` | MOSS-TTS-Local Transformer v1.5 | No | |
| 119 | +| `nemotron_asr` | Nemotron ASR | **Yes** | |
| 120 | +| `omnivoice` | OmniVoice | **Yes** | |
| 121 | +| `outetts_1_0_1b` | OuteTTS 1.0 1B with IBM DAC codec and Qwen3-aligned voice cloning | No | |
| 122 | +| `pocket_tts` | PocketTTS | **Yes** | |
| 123 | +| `qwen3_asr_0_6b` | Qwen3 ASR 0.6B | **Yes** | |
| 124 | +| `qwen3_asr_1_7b_hf` | Qwen3 ASR 1.7B HF | **Yes** | |
| 125 | +| `qwen3_forced_aligner_0_6b` | Qwen3 Forced Aligner 0.6B | **Yes** | |
| 126 | +| `qwen3_tts_0_6b_base` | Qwen3 TTS 12Hz 0.6B Base | **Yes** | |
| 127 | +| `qwen3_tts_1_7b_base` | Qwen3 TTS 12Hz 1.7B Base | **Yes** | |
| 128 | +| `qwen3_tts_1_7b_custom_voice` | Qwen3 TTS 12Hz 1.7B Custom Voice | **Yes** | |
| 129 | +| `qwen3_tts_1_7b_voice_design` | Qwen3 TTS 12Hz 1.7B Voice Design | **Yes** | |
| 130 | +| `seed_vc` | SeedVC-MLX | **Yes** | |
| 131 | +| `sortformer_diar_4spk_v1` | Sortformer diarization 4 speaker v1 | **Yes** | |
| 132 | +| `stable_audio_3_medium` | Stable Audio 3 Medium | **Yes** | |
| 133 | +| `stable_audio_3_small_music` | Stable Audio 3 Small Music | **Yes** | |
| 134 | +| `stable_audio_3_small_sfx` | Stable Audio 3 Small SFX | **Yes** | |
| 135 | +| `supertonic_3` | Supertonic 3 | **Yes** | |
| 136 | +| `vevo2` | VeVo2 | No | |
| 137 | +| `vietneu_tts_v3_turbo` | VieNeu-TTS v3 Turbo | **Yes** | |
| 138 | +| `vibevoice_1_5b` | VibeVoice 1.5B | **Yes** | |
| 139 | +| `vibevoice_7b` | VibeVoice 7B | **Yes** | |
| 140 | +| `vibevoice_asr` | VibeVoice ASR | **Yes** | |
| 141 | +| `voxcpm2` | VoxCPM2 | No | |
| 142 | +| `voxtral_realtime` | Voxtral Mini 4B Realtime GGUF Q8_0 | **Yes** | |
0 commit comments