Use this thread to:
Request new features
Pick up open features TODOs
Ask or find answers to common questions
PRs for these features are greatly appreciated! If you want to work on a TODO, please comment first to avoid duplicated effort.
Features
UI:
FAQ
Docker can't run
This is usually a CUDA driver compatibility issue. See https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html#cuda-toolkit-and-corresponding-driver-versions

Vulkan
Nvidia Jetson Orin
Multi-component models: YuE2, Minimax Music3, AuK
- If you’re using binaries built from source, you may encounter a
model directory contains 2 GGUF files error, causing the server/CLI to refuse to load the model. There are two ways to address this:
(1) Use --model-spec-override /absolute/path/to/model_spec.json
(2) Build with AUDIOCPP_DEPLOYMENT_BUILD=ON
Kokoro and other TTS models with built-in voices
The preset voices are usually packed in the GGUFs. /v1/audio/voices lists server-configured presets and voice-library files, rather than inspecting model-specific GGUF contents. This keeps the shared endpoint independent of each model’s voice format. The same applies to KittenTTS and MagpieTTS. E.g.
curl -i 'http://127.0.0.1:8894/v1/audio/voices?model=magpie' should return
HTTP/1.1 200 OK
{"voices":[]}
You still can select built-in voices directly by name, or add voice_presets to the server config to expose them in the list. The model-specific docs list the available voice name (not all voices are packed).
AuK/Auk Flash
Honest take: only certain tasks this model (official Python impl) supports work as expected.
We validate parity with upstream Python for the tested configuration, not whether AuK's output meets every quality expectation. If a result is disappointing, listen to the corresponding Python reference WAV linked below first.
https://huggingface.co/audio-cpp/AuK-Base-and-Flash-GGUF
A similar Python result points to the upstream model's behavior, not necessarily an audio.cpp conversion issue.
#612 also provides human evaluation of the results.
Placeholder...
Use this thread to:
Request new features
Pick up open features TODOs
Ask or find answers to common questions
PRs for these features are greatly appreciated! If you want to work on a TODO, please comment first to avoid duplicated effort.
Features
UI:
FAQ
Docker can't run
This is usually a CUDA driver compatibility issue. See https://docs.nvidia.com/cuda/cuda-toolkit-release-notes/index.html#cuda-toolkit-and-corresponding-driver-versions

Vulkan
set "GGML_VK_ALLOW_SYSMEM_FALLBACK=1"HeartMuLa 3B: failed to allocate HeartCodec scalar decoder graph -- Can't use Windows shared VRAM. Workaround: set "GGML_VK_ALLOW_SYSMEM_FALLBACK=1" #434Nvidia Jetson Orin
Multi-component models: YuE2, Minimax Music3, AuK
model directory contains 2 GGUF fileserror, causing the server/CLI to refuse to load the model. There are two ways to address this:(1) Use
--model-spec-override /absolute/path/to/model_spec.json(2) Build with
AUDIOCPP_DEPLOYMENT_BUILD=ONKokoro and other TTS models with built-in voices
The preset voices are usually packed in the GGUFs.
/v1/audio/voiceslists server-configured presets and voice-library files, rather than inspecting model-specific GGUF contents. This keeps the shared endpoint independent of each model’s voice format. The same applies to KittenTTS and MagpieTTS. E.g.curl -i 'http://127.0.0.1:8894/v1/audio/voices?model=magpie' should returnYou still can select built-in voices directly by name, or add
voice_presetsto the server config to expose them in the list. The model-specific docs list the available voice name (not all voices are packed).AuK/Auk Flash
Honest take: only certain tasks this model (official Python impl) supports work as expected.
We validate parity with upstream Python for the tested configuration, not whether AuK's output meets every quality expectation. If a result is disappointing, listen to the corresponding Python reference WAV linked below first.
https://huggingface.co/audio-cpp/AuK-Base-and-Flash-GGUF
A similar Python result points to the upstream model's behavior, not necessarily an audio.cpp conversion issue.
#612 also provides human evaluation of the results.
Placeholder...