Skip to content

Commit 3578567

Browse files
committed
Release 0.4
1 parent 20c57bd commit 3578567

8 files changed

Lines changed: 302 additions & 269 deletions

File tree

‎CONTRIBUTING.md‎

Lines changed: 7 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -39,7 +39,9 @@ Good follow-up work for existing model families includes:
3939

4040
## New Model PRs
4141

42-
New model PRs should include enough evidence for maintainers and users to understand exactly what was tested. Follow the validation style shown in [PR #19](https://github.com/0xShug0/audio.cpp/pull/19).
42+
New standalone model ports should normally start under `community_models/`. This keeps ownership clear and lets useful model ports land with a lighter review bar than core framework models. Models can graduate into the core model tree later after they are validated, polished, and maintained as part of the main release surface.
43+
44+
Even for community models, PRs should include enough evidence for maintainers and users to understand exactly what was tested. Follow the validation style shown in [PR #19](https://github.com/0xShug0/audio.cpp/pull/19) and [PR #63](https://github.com/0xShug0/audio.cpp/pull/63).
4345

4446
Please include:
4547

@@ -52,6 +54,8 @@ Please include:
5254
- Relevant timing, RTF, RSS, VRAM, or resident-memory notes
5355
- Known limitations
5456

57+
For TTS-style community models, the most useful validation includes a long-lived session with multiple requests, a long-form request through the framework text chunker, cache/graph reuse logs when relevant, and peak VRAM under repeated requests. These measurements do not need to be perfect on the first PR, but they make review much faster and help avoid a large cleanup pass after merge.
58+
5559
For models with Python references, include parity evidence when practical. For models without a clean reference path, include reproducible generated outputs and enough setup detail for another contributor to repeat the run.
5660

5761
Please be prepared to help maintain new model contributions as framework APIs evolve. Keeping model code aligned with the shared framework surface is part of making the implementation useful long term.
@@ -80,11 +84,12 @@ audio.cpp is moving faster because people keep showing up with real fixes, caref
8084
- [@lapy](https://github.com/lapy) for the machine-readable loader/package catalog exports and the loader-catalog sync checks that keep package metadata honest in [#74](https://github.com/0xShug0/audio.cpp/pull/74) and [#86](https://github.com/0xShug0/audio.cpp/pull/86).
8185
- [@fedeizzo](https://github.com/fedeizzo) for the cross-platform Nix flake and follow-up Nix documentation polish in [#82](https://github.com/0xShug0/audio.cpp/pull/82) and [#83](https://github.com/0xShug0/audio.cpp/pull/83).
8286
- [@phuocnguyen90](https://github.com/phuocnguyen90) for bringing VieNeu-TTS v3 Turbo into the community model surface in [#80](https://github.com/0xShug0/audio.cpp/pull/80).
87+
- [@mosujiba](https://github.com/mosujiba) for adding configurable CORS handling to the server path in [#85](https://github.com/0xShug0/audio.cpp/pull/85).
8388
- [@adambenhassen](https://github.com/adambenhassen) for PocketTTS runtime fixes and upstream-aligned English defaults in [#76](https://github.com/0xShug0/audio.cpp/pull/76) and [#77](https://github.com/0xShug0/audio.cpp/pull/77).
8489
- [@vicenteliu](https://github.com/vicenteliu) for hardening the server against client disconnects by ignoring `SIGPIPE` in [#78](https://github.com/0xShug0/audio.cpp/pull/78).
8590
- [@Cr4xy](https://github.com/Cr4xy) for improving multipart upload handling and removing temporary-file writes from that path in [#61](https://github.com/0xShug0/audio.cpp/pull/61).
8691
- [@kevin-ho](https://github.com/kevin-ho) for making single-model server voice discovery work cleanly when the model parameter is omitted in [#64](https://github.com/0xShug0/audio.cpp/pull/64).
8792
- [@xashr](https://github.com/xashr) for Dockerfiles, Docker examples, Docker documentation, and CI workflow polish in [#30](https://github.com/0xShug0/audio.cpp/pull/30), [#51](https://github.com/0xShug0/audio.cpp/pull/51), and [#81](https://github.com/0xShug0/audio.cpp/pull/81).
8893
- [@5uck1ess](https://github.com/5uck1ess) for improving Citrinet CTC decoding through the SentencePiece model in [#49](https://github.com/0xShug0/audio.cpp/pull/49).
8994
- [@dkruyt](https://github.com/dkruyt) for the first multipart transcription upload support in [#25](https://github.com/0xShug0/audio.cpp/pull/25).
90-
- [@CaptainArni](https://github.com/CaptainArni) for fixing PocketTTS empty output when switching cached voices in [#22](https://github.com/0xShug0/audio.cpp/pull/22).
95+
- [@CaptainArni](https://github.com/CaptainArni) for fixing PocketTTS empty output when switching cached voices and keeping the Windows CUDA build path healthy in [#22](https://github.com/0xShug0/audio.cpp/pull/22) and [#93](https://github.com/0xShug0/audio.cpp/pull/93).

‎README.md‎

Lines changed: 58 additions & 244 deletions
Large diffs are not rendered by default.

‎docs/community_models/models.md‎

Lines changed: 4 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -2,12 +2,15 @@
22

33
Community model ports live under `community_models` to make the ownership boundary clear while keeping them available through the normal audio.cpp CLI and server paths. Some community-contributed models graduate into the core model tree when they become part of the main release surface.
44

5-
The review bar for community models is intentionally lighter than core model integrations, so contributors can share useful ports earlier. They should still meet these practical expectations:
5+
The review bar for community models is intentionally lighter than core model integrations, so contributors can share useful ports earlier. The model does not need to be fully promoted into the core model tree on day one, but it should still be reproducible and honest about its limits.
6+
7+
Practical expectations:
68

79
- RTF should be below 1.0.
810
- VRAM usage should stay stable across multiple requests. If memory needs to be optimized, use `mem_saver` to balance performance and VRAM instead of hiding leaks.
911
- Long-form generation should work correctly. The shared long-form TTS/clone test cases live in `tools/audiocpp_cli/audiocpp_cli_longform_tts_clone_cases.json`.
1012
- Use existing framework modules and patterns as much as possible.
13+
- Include exact build/run commands, generated WAVs or output artifacts, backend coverage, parity or path-test results when available, and timing/memory notes. [PR #19](https://github.com/0xShug0/audio.cpp/pull/19) and [PR #63](https://github.com/0xShug0/audio.cpp/pull/63) are good examples of contributors providing enough detail for maintainers to reproduce and review the model.
1114

1215
## Current Community Models
1316

‎docs/gguf.md‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -102,6 +102,12 @@ Q8 packaging notes:
102102
tensors in Q8 in addition to the default converter selection. `conditioner.embed`,
103103
`cond_embed`, and Mimi conv tensors are not forced to Q8 because tested outputs
104104
drifted or the current conv path casts quantized conv weights back to F32.
105+
- `qwen3_tts` Q8 should keep speaker-sensitive components in their original
106+
16-bit type. The tested Base Q8 package quantizes the talker transformer and
107+
projections, talker code-predictor heads, and speech-tokenizer encoder/decoder
108+
projection or linear weights, while leaving the speaker encoder, lookup, and
109+
codebook-sensitive tensors unquantized. Quantizing those speaker-side tensors
110+
can produce long-form quality problems such as large silence.
105111

106112
## Build The Converter
107113

@@ -303,3 +309,6 @@ Compatibility with older binaries:
303309
Quantized GGUF support is model- and route-specific. A model may load successfully but
304310
still drift in length, waveform similarity, or recognized text, so validate the exact
305311
route you plan to ship.
312+
313+
For measured 16-bit vs Q8 speed and peak VRAM results, see
314+
[GGUF Q8 performance](reports/gguf_q8_performance.md).

‎docs/maintainers/loader_and_catalog.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -19,7 +19,7 @@ For every **installable, standalone** `ModelPackage`:
1919
| `ModelPackage.family` | The loader family string advertised by the C++ loader |
2020
| `model_specs/<family>.json` | Present when the family uses package-spec loading |
2121
| `registry.cpp` entry | Uncommented `make_<family>_loader()` (or the family's actual factory name) |
22-
| README package table | Lists the package; use **Unavailable** when not installable |
22+
| `docs/model_manager.md` package table | Lists the package; use **Unavailable** when not installable |
2323

2424
Dependency / subcomponent packages (`standalone=False`, with
2525
`parent_package_id`) do **not** need their own loader.
@@ -33,7 +33,7 @@ If a loader is not ready for this release tree:
3333
1. Keep it **commented out** in `src/framework/runtime/registry.cpp`, and
3434
2. Mark matching catalog packages as `UnsupportedSource(reason=...)`, **or**
3535
remove them from `CATALOG`, and
36-
3. Mark the README package row **Unavailable**.
36+
3. Mark the `docs/model_manager.md` package row **Unavailable**.
3737

3838
Do **not** leave a live `SnapshotSource` for a commented-out loader.
3939

@@ -97,7 +97,7 @@ builds. It:
9797

9898
- Parses active vs commented `make_*_loader()` calls in `registry.cpp`
9999
- Compares them to installable standalone packages from `model_manager.py`
100-
- Cross-checks the README recommended package table
100+
- Cross-checks the `docs/model_manager.md` recommended package table
101101
- Does **not** require a compiled binary
102102

103103
```bash

‎docs/model_manager.md‎

Lines changed: 142 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,142 @@
1+
# Model Manager
2+
3+
`tools/model_manager.py` downloads or assembles supported model packages into the
4+
framework's expected `models/` layout.
5+
6+
This tool is still useful for safetensors-based packages and a few composite model
7+
layouts, but it is gradually becoming a legacy path as audio.cpp moves toward
8+
standalone GGUF packages.
9+
10+
If a model has a ready-to-use GGUF package, prefer that route first.
11+
12+
## GGUF Downloads
13+
14+
Ready-to-use GGUF packages are published here:
15+
16+
- Core released models: [audio-cpp/audio.cpp-gguf](https://huggingface.co/audio-cpp/audio.cpp-gguf)
17+
- Community OuteTTS package: [mirek190/audio.cpp](https://huggingface.co/mirek190/audio.cpp/tree/main/Text%20to%20audio%20(TTS))
18+
19+
For support status and tested precision coverage, see the [GGUF guide](gguf.md).
20+
For measured 16-bit vs Q8 speed and peak-VRAM results, see the
21+
[Q8 performance report](reports/gguf_q8_performance.md).
22+
23+
## Dependencies
24+
25+
- Python 3
26+
- `torch`
27+
- `safetensors`
28+
- `PyYAML`
29+
- Network access to the upstream model source
30+
31+
## Commands
32+
33+
- `list` shows the available package ids
34+
- `list --json` prints a machine-readable package catalog
35+
- `info` shows the target layout, required files, and install source for one package
36+
- `info <package> --json` prints machine-readable package details
37+
- `install` downloads or converts one package into a models root
38+
39+
The runtime loader catalog is also available from:
40+
41+
```bash
42+
audiocpp_cli --list-loaders --json
43+
```
44+
45+
## Quick Start
46+
47+
List installable packages:
48+
49+
```bash
50+
python3 tools/model_manager.py list
51+
```
52+
53+
Inspect one package:
54+
55+
```bash
56+
python3 tools/model_manager.py info qwen3_tts_1_7b_base
57+
```
58+
59+
Install into the default `models/` directory:
60+
61+
```bash
62+
python3 tools/model_manager.py install qwen3_tts_1_7b_base
63+
```
64+
65+
Install into a custom models root:
66+
67+
```bash
68+
python3 tools/model_manager.py install vevo2 --models-root /path/to/models
69+
```
70+
71+
Overwrite an existing install:
72+
73+
```bash
74+
python3 tools/model_manager.py install pocket_tts --overwrite
75+
```
76+
77+
Install a converter-style package that needs a source file:
78+
79+
```bash
80+
python3 tools/model_manager.py info voxcpm2_audiovae
81+
python3 tools/model_manager.py install voxcpm2_audiovae --source-file models/VoxCPM2/audiovae.pth --models-root models --overwrite
82+
```
83+
84+
## Package Notes
85+
86+
For shared audio.cpp GGUF packages, the model manager installs the default `q8_0`
87+
GGUF. Other precision variants can be downloaded directly from
88+
[audio-cpp/audio.cpp-gguf](https://huggingface.co/audio-cpp/audio.cpp-gguf).
89+
90+
`Yes` means Hugging Face has a ready-to-use repo that the framework can download
91+
as-is. `No` means the tool must assemble, convert, or post-process files before the
92+
framework can use them.
93+
94+
Packages whose loaders are not registered in the current release tree are listed as
95+
**Unavailable**; see [loader/catalog sync notes](maintainers/loader_and_catalog.md).
96+
97+
| Package id | Model | HF ready-to-use repo |
98+
|---|---|---|
99+
| `ace_step` | ACE-Step 1.5 Turbo/Base | No |
100+
| `chatterbox` | Chatterbox | **Yes** |
101+
| `citrinet_asr` | Citrinet ASR converted layout | No |
102+
| `fish_audio_s2_pro` | Fish Audio S2 Pro GGUF Q8_0 | **Yes** |
103+
| `heartmula` | HeartMuLa | No |
104+
| `higgs_audio_stt` | Higgs Audio STT | No |
105+
| `higgs_audio_v3_tts_4b` | Higgs Audio v3 TTS 4B GGUF Q8_0 | **Yes** |
106+
| `htdemucs` | HTDemucs | No |
107+
| `hviske_asr` | Hviske ASR | **Yes** |
108+
| `irodori_tts_500m_v3` | Irodori-TTS 500M v3 | No |
109+
| `irodori_tts_600m_v3_voice_design` | Irodori-TTS 600M v3 VoiceDesign | No |
110+
| `index_tts2` | IndexTTS-2 | **Yes** |
111+
| `mel_band_roformer` | Mel-Band RoFormer MLX | **Yes** |
112+
| `miocodec_25hz_44k_v2` | MioCodec 25Hz 44.1kHz v2 | No |
113+
| `miotts_1_7b` | MioTTS 1.7B | No |
114+
| `moss_audio_tokenizer_nano` | MOSS Audio Tokenizer Nano | No |
115+
| `moss_audio_tokenizer_v2` | MOSS Audio Tokenizer v2 | No |
116+
| `moss_tts_nano_100m` | MOSS-TTS-Nano 100M | No |
117+
| `moss_tts_nano_100m_model` | MOSS-TTS-Nano 100M model subcomponent | No |
118+
| `moss_tts_local_v1_5` | MOSS-TTS-Local Transformer v1.5 | No |
119+
| `nemotron_asr` | Nemotron ASR | **Yes** |
120+
| `omnivoice` | OmniVoice | **Yes** |
121+
| `outetts_1_0_1b` | OuteTTS 1.0 1B with IBM DAC codec and Qwen3-aligned voice cloning | No |
122+
| `pocket_tts` | PocketTTS | **Yes** |
123+
| `qwen3_asr_0_6b` | Qwen3 ASR 0.6B | **Yes** |
124+
| `qwen3_asr_1_7b_hf` | Qwen3 ASR 1.7B HF | **Yes** |
125+
| `qwen3_forced_aligner_0_6b` | Qwen3 Forced Aligner 0.6B | **Yes** |
126+
| `qwen3_tts_0_6b_base` | Qwen3 TTS 12Hz 0.6B Base | **Yes** |
127+
| `qwen3_tts_1_7b_base` | Qwen3 TTS 12Hz 1.7B Base | **Yes** |
128+
| `qwen3_tts_1_7b_custom_voice` | Qwen3 TTS 12Hz 1.7B Custom Voice | **Yes** |
129+
| `qwen3_tts_1_7b_voice_design` | Qwen3 TTS 12Hz 1.7B Voice Design | **Yes** |
130+
| `seed_vc` | SeedVC-MLX | **Yes** |
131+
| `sortformer_diar_4spk_v1` | Sortformer diarization 4 speaker v1 | **Yes** |
132+
| `stable_audio_3_medium` | Stable Audio 3 Medium | **Yes** |
133+
| `stable_audio_3_small_music` | Stable Audio 3 Small Music | **Yes** |
134+
| `stable_audio_3_small_sfx` | Stable Audio 3 Small SFX | **Yes** |
135+
| `supertonic_3` | Supertonic 3 | **Yes** |
136+
| `vevo2` | VeVo2 | No |
137+
| `vietneu_tts_v3_turbo` | VieNeu-TTS v3 Turbo | **Yes** |
138+
| `vibevoice_1_5b` | VibeVoice 1.5B | **Yes** |
139+
| `vibevoice_7b` | VibeVoice 7B | **Yes** |
140+
| `vibevoice_asr` | VibeVoice ASR | **Yes** |
141+
| `voxcpm2` | VoxCPM2 | No |
142+
| `voxtral_realtime` | Voxtral Mini 4B Realtime GGUF Q8_0 | **Yes** |
Lines changed: 59 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,59 @@
1+
# GGUF Q8 Performance
2+
3+
This report compares tested 16-bit GGUF packages against Q8_0 GGUF packages on CUDA.
4+
In the measured routes, Q8_0 improves wall time by up to **1.53x** and lowers
5+
peak VRAM by up to about **37%** compared with the matching 16-bit GGUF package.
6+
7+
Notes:
8+
9+
- Offline session rows exclude the warmup request.
10+
- "Speed vs real time" means generated or processed audio duration divided by wall time.
11+
- "Q8 vs 16-bit" means 16-bit wall time divided by Q8 wall time; values above `1.0x` mean Q8 was faster.
12+
- `qwen3_tts` uses the regenerated `q8_v2` GGUF, which keeps speaker-sensitive tensors in 16-bit storage.
13+
- `higgs_audio_tts` 16-bit long-form uses `text_chunk_size=256`.
14+
15+
## Summary
16+
17+
Q8 gives the clearest end-to-end wins on the larger AR-style models:
18+
19+
- `higgs_audio_tts`: Q8 is **1.38x-1.53x faster** on warmed requests and lowers peak VRAM from **9.1 GiB to 5.9 GiB**.
20+
- `fish_audio`: Q8 is **1.26x-1.34x faster** on warmed requests and lowers peak VRAM from **11.8 GiB to 7.5 GiB**.
21+
- `voxtral_realtime`: Q8 is **1.31x-1.38x faster** in offline ASR and lowers peak VRAM from **10.7 GiB to 7.6 GiB**.
22+
23+
Smaller or already memory-light models still load and run with Q8, but the speed gain can be modest. Treat Q8 as a measured route choice, not a guaranteed win for every model.
24+
25+
## TTS Offline Long-Lived Session
26+
27+
| Model | 16-bit speed vs real time | Q8 speed vs real time | Q8 vs 16-bit | 16-bit peak VRAM | Q8 peak VRAM |
28+
|---|---:|---:|---:|---:|---:|
29+
| `pocket_tts` | 68.5x-99.0x | 78.7x-102.0x | 1.01x-1.21x | 2060 MiB | 1932 MiB |
30+
| `chatterbox` | 6.2x-8.2x | 6.2x-8.7x | 1.01x-1.11x | 4333 MiB | 4309 MiB |
31+
| `omnivoice` | 7.2x-24.4x | 8.2x-29.0x | 1.08x-1.22x | 3250 MiB | 3157 MiB |
32+
| `qwen3_tts` | 5.3x-7.0x | 5.8x-7.9x | 1.01x-1.12x | 7873 MiB | 6397 MiB |
33+
| `fish_audio` | 2.5x-2.6x | 3.1x-3.4x | 1.26x-1.34x | 12093 MiB | 7669 MiB |
34+
| `higgs_audio_tts` | 6.1x-6.7x | 8.8x-10.1x | 1.38x-1.53x | 9326 MiB | 6024 MiB |
35+
36+
## TTS Offline Long-Form
37+
38+
| Model | 16-bit speed vs real time | Q8 speed vs real time | Q8 vs 16-bit | 16-bit peak VRAM | Q8 peak VRAM |
39+
|---|---:|---:|---:|---:|---:|
40+
| `pocket_tts` | 82.6x | 85.5x | 1.04x | 2098 MiB | 2298 MiB |
41+
| `chatterbox` | 7.9x | 8.5x | 1.06x | 5033 MiB | 4554 MiB |
42+
| `omnivoice` | 38.6x | 43.7x | 1.13x | 3175 MiB | 3069 MiB |
43+
| `qwen3_tts` | 5.9x | 6.6x | 1.12x | 9412 MiB | 8138 MiB |
44+
| `fish_audio` | 2.5x | 3.3x | 1.27x | 12261 MiB | 9228 MiB |
45+
| `higgs_audio_tts` | 6.1x | 8.5x | 1.41x | 11878 MiB | 9129 MiB |
46+
47+
## ASR Offline Long-Lived Session
48+
49+
| Model | 16-bit speed vs real time | Q8 speed vs real time | Q8 vs 16-bit | 16-bit peak VRAM | Q8 peak VRAM |
50+
|---|---:|---:|---:|---:|---:|
51+
| `voxtral_realtime` | 11.1x-12.5x | 14.7x-16.7x | 1.31x-1.38x | 10909 MiB | 7754 MiB |
52+
| `nemotron_asr` | 277.8x-384.6x | 285.7x-400.0x | 1.03x-1.12x | 5125 MiB | 4028 MiB |
53+
54+
## ASR Streaming Long Audio
55+
56+
| Model | 16-bit server TTFT | Q8 server TTFT | 16-bit client TTFT | Q8 client TTFT | 16-bit speed vs real time | Q8 speed vs real time | 16-bit peak VRAM | Q8 peak VRAM |
57+
|---|---:|---:|---:|---:|---:|---:|---:|---:|
58+
| `voxtral_realtime` | 207.308 ms | 179.896 ms | 550.526 ms | 530.558 ms | 4.7x | 5.4x | 12616 MiB | 8972 MiB |
59+
| `nemotron_asr` | 205.007 ms | 214.822 ms | 488.629 ms | 499.453 ms | 31.6x | 33.4x | 2816 MiB | 2497 MiB |

0 commit comments

Comments
 (0)