Live work only. Completed threads move to HISTORY.md; technical deep-dives to
LEARNINGS.md.
Before you pick something up: re-read this section on origin/main, add a
## CLAIMED <date> — <what> block naming your worktree, and push that claim
to main before you start. Several agents run here at once; a claim that lands
with the work is a claim that did nothing. Delete it when the work lands, or if
it goes stale for more than a day.
MioTTS and the opt-in Echo scheduler source are integrated; see
the integration report and HISTORY.
Q4 tooling is present, but no Q4 candidate has passed runtime acceptance.
Large artifacts/logs stay under /mnt/storage/crispasr/miotts-echo-20261002/
and /mnt/storage/crispasr/miotts-echo-integration-20261003/.
- Four decoder recipes are defined in
tools/index_echo_quant_recipes.py: plain Q4 baseline; F16 sensitive tensors with Q8 attention/down and Q4 gate/up; all gate/up Q4 with down Q8; and only layers 4–27 gate/up Q4. All 177 original F32 tensors, including 24 recurrent convolution matrices, must retain F32; the acoustic tower/connector remain original. Source pair:cstr/index-echo-9b-GGUF@dffbadf0f173446fee0364a0807803d2b2fb6f49. - Kaggle q4-guards v1 passed the F16 control on two T4s, then repository creation failed with HTTP 403 before any candidate ran. All 31 outputs and terminal logs are retained. CPU preparation was removed from the GPU wrapper.
- Hosted CPU preparation 37044026371 was cancelled to correct F32 guards. Corrected 37046153440 produced a 5.05 GB plain Q4 decoder, but its HF upload commit failed with HTTP 400 (private repository storage quota). No candidate has successfully uploaded artifact pins or GPU acceptance. Private staging is a transfer choice, not a runtime requirement; do not retry the quota-limited route unchanged.
- GPU-only q4-guards v2 has not been pushed. Preparation-run/revision/checksum pins are deliberately unset; a successful producer/transfer is required.
Next: establish a feasible artifact transfer route; prepare/audit/pin all
candidates on hosted CPU; then run GPU-only canonical stage/magnitude/cache,
exact full decoded-output and TTS→ASR gates without relaxing thresholds.
Public weights and defaults remain unchanged until a candidate passes.
Original feature and integration proof source remain available as
archive/miotts-echo-q4-source-20261003 and
archive/miotts-echo-integration-proof-20261003.
- #461: reporter at RTF 1.01 (Arc B390, 8 steps,
voxcpm2-q8_0-locdit-f16.gguf) on a pre-#478 build; asked to re-bench on main (estimate ~0.85). What remains is the compute-bound LocDiT matmuls (22 columns) at 10 steps — only a faster Vulkan kernel for that shape would move the default-quality number. - #478: native AVX2/F16C CPU F16/Q8 RALM batching is now validated and default
(see
HISTORY.md/PERFORMANCE.md). Reporter Windows/249-position Khmer measurement remains external. Q4 batching stays opt-in: up to 1.83% state error and one extra-word readback; GPU and other CPU ISAs need their own proof. CRISPASR_VOXCPM2_VAE_DW_SHIFTis default-on for Vulkan only; CUDA/Metal unmeasured.
Deferred by the maintainer. Assessment posted on the issue: WavLM-large + a Kaldi-style GMM-HMM (model.npz ~2 MB, decision trees, projections) + 41k-word lexicon + beam Viterbi. Inference code MIT, but model files AND outputs (the timestamps) are under the nyra health Non-Commercial Research License with share-alike and a contractual-binding clause (3.3) -> would ship as a gated NC GGUF; espeak-ng OOV fallback is GPL and cannot be bundled. WavLM exists in-tree (MioCodec). Not started.
CLAIMED 2026-09-26 — §247 shared scheduler profiler, then hidden-F16 audit
Worktree .claude/worktrees/roadmap-445-438-337-perf, branch
feat/445-orukeet. The prerequisite model and GPU stages are complete: Orukeet
and R2T2 (#445), both Hojo checkpoints (#438), and native gfx1100 Vulkan
Qwen3-TTS (#337) are on main with model or hardware proof. The remaining
sequence is strict:
- move the FastConformer scheduler callback into
src/core/sched_prof.hand expose one opt-in profiler usable by scheduler-based runtimes — DONE:CRISPASR_SCHED_PROFILE=1is wired through Canary CTC, Cohere, FireRed-ASR, Granite Speech, Moonshine, Moonshine Streaming and Paraformer; the legacyCRISPASR_FC_PROFILEswitch still works; - run the profiler/metadata audit over quantized GGUF families, fix the largest
hidden-F16 matmul offenders one at a time, and prove each A/B before defaulting.
Cohere DONE 2026-09-28: the published Q4 artifact was requantized once;
all 96 pointwise weights changed F16 -> Q8_0, no other tensor changed, the
JFK transcript stayed identical, encoder profile improved 10242.59 ->
9447.52 ms and process wall time improved 10.93 -> 10.16 s. A remote range
read reconstructed the complete 1,289,178,240-byte GGUF and confirmed
Q8_0=96. The successful build also refreshed
${KAGGLE_ACCOUNT}/crispasr-ccache. NOW: apply the same artifact-time audit to FireRed-ASR's 32 retained F16 pointwise matrices (~300 MiB). Granite and Moonshine have no large F16 residuals; Paraformer's are mostly tiny FSMN kernels.
The port shipped (issue closed; record in HISTORY "PLAN compaction 2026-09-30"). Voice Design and Voice Direction are REFUSED rather than downgraded: the cache topology and branch index exist, per-branch prompt assembly and the logits combine do not.
Fixed and shipped in v0.8.32, issue closed (record in HISTORY). Only open:
tools/kaggle/sidon-quant-cuda/ would answer whether CUDA MMQ (sm >= 61) shares
the broadcast-RPE defect; P100 draws cannot (MMQ off below DP4A).
Shipped (record in HISTORY). Port the CTC head when the author uploads the .nemo.
- Linear-resampler gap on 44.1/48 kHz compressed input through the glint decode paths (~28 vs ~38 dB after the decoder fix).
- Speed (M1 audit 2026-08-19: encoder compute-bound, decoder ~35 % of GPU wall):
- persistent decoder step graph (
canary_decode_steprebuilds per token, ~4.2 ms/tok on Metal;core_rnnt_ggml::Decoderpattern, est. ~1.3x); - chunk-batched encode/decode for long-form (NeMo runs batch_size=8). Both need byte-identical Metal+CPU A/B and CUDA validation before a default flip.
- persistent decoder step graph (
The transcript-side damage is FIXED (3b1bc0b2, see HISTORY): [Silence] is a
Content value the MODEL emits and we no longer pass it through as transcript
text, so it cannot reach an SRT or suppress the empty-transcript warning.
Still open: why the 7B says it at all. Both members of the reporter's atempo pair
(ko-test stretched 3.26 s -> ~6 s at atempo 0.535 / 0.525) come back with no
utterance, on CPU and CUDA, in EVERY arm including one with all four #369 fixes
rolled back — so this is not something we introduced. The 1.5B BitNet checkpoint
transcribes the same two files, so it is specific to the 7B. A 2x time-stretch is
not exotic input, and a long recording containing a slow passage would lose it.
Next: dump speech_features for a stretched clip against its unstretched
original. The encoder is trustworthy now (cos 0.999926 vs upstream's own
modules), so if the conditioning matches, the divergence is the LM's. Also check
the prompt's duration string ("This is a 6.11 seconds audio") against the 46
speech tokens for an inconsistency the model could read as "mostly empty".
Fallout from 51b99d1b, recorded rather than fixed on the way past. The 1.5B now
gets its own plain-text instruction, and in that mode it returns prose rather
than the JSON array — which is what "plain text output" means upstream. So there
are no per-utterance timings and no speaker labels, while vibevoice-bitnet
still declares CAP_DIARIZE and CAP_TIMESTAMPS_CTC.
That is the same class of false claim as the CAP_TEMPERATURE removed in
23107227: a capability bit is a promise about output the framework then acts
on. Two defensible fixes and they need a decision, not a reflex:
(a) drop both caps for the 1.5B — honest, and --diarize then warns; or
(b) have the adapter switch to CRISPASR_VIBEVOICE_ASR_PROMPT=json when the
user actually asks for diarization or timestamps, trading the non-English
quality back for the structure they asked for.
(b) is friendlier but makes output quality depend on an unrelated flag, which is
the kind of thing that gets rediscovered as a bug later.
vibevoice-asr-bitnet-* (TQ2_0 LM weights) yields an EMPTY transcript on Metal:
ggml_metal_library_compile_pipeline: failed to compile pipeline:
base = 'kernel_mul_mm_tq2_0_f32'
Error: Function kernel_mul_mm_tq2_0_f32 was not found in the library
ggml/src/ggml-metal/ggml-metal.metal contains ZERO occurrences of tq2_0 — no
mul_mm, no mul_mv, no dequant — and ggml-metal-device.m has no TQ2_0 entry
either. The type is not supported at all, yet a pipeline for it is still
requested, so it fails hard instead of falling back.
Same clip, same build, only the backend differs:
ko-mic-cue-kept.wav CPU (-ng) -> 내일 오전에 회의 자료 교육 보내주세요.
Metal -> (nothing; pipeline compile error)
Impact: every Metal user of a TQ2_0 model gets silence. Ternary/BitNet GGUFs are what people reach for on laptops, so this is the wrong platform to be missing.
Two questions before fixing: (a) why is a TQ2_0 matmul scheduled onto Metal when
the device declares no support — a type absent from the support switch should
route to CPU, so something is bypassing that; (b) whether the fix is a real
kernel_mul_mm_tq2_0_f32 (check upstream ggml first — it may already exist) or
an explicit unsupported-declaration so the scheduler falls back cleanly. The
second is small and unbreaks the platform immediately.
Found while reproducing #369, and NOT that issue's cause: the reporter is on Windows CPU/Vulkan and sees wrong-language output, not silence.
These are the loose ends from 619e74b6..4b875be0. None block anything.
- fr/es have not been checked for either German defect. Both were found by
method, and the method transfers:
(a) citation stress —
espeak_fr.tsv/espeak_es.tsvwere generated the same way (one word at a time), so they bake in the same isolation stress German had.tools/gen-g2p-de-unstressed.pyis the generator; pointing it at another language is a few lines. (b) out-of-vocabulary symbols — scan the G2P output against the model's owntokenizer.ggml.tokens. Three lines of Python, and in German it found we were deletingʏout of every München. Worth running for every non-English Kokoro model we ship, not just fr/es. - Regenerate
espeak_de.tsvwith--tie. Our dictionary has no tie marks, socore_phoneme's Germants→ʦis a blanket rewrite that cannot tell an affricate from a compound seam. espeak emitst^sonly for the affricate, so a tied dictionary makes the collapse exact — and would letCRISPASR_KOKORO_DE_MISAKI_ALPHABETbe judged on its merits instead of on an approximation. Regenerate + re-upload tocstr/g2p-dicts. - The German tied-alphabet collapse needs a listening test, or a kikiri-tts
model to test against. It matches the published recipe and made the ASR
round-trip worse on the hui base we ship; the hypothesis is that that model
predates that part of the recipe.
kikiri-tts/kikiri-german-{victoria,martin}are explicitly "misaki 0.9.4 + espeak-ng" and would settle it. - Five
tools/kaggle/*/kaggle_harness.pybundles are gitignored —mimo-cuda-rvq-309,moss-tts-quants,streaming-diarize-300,whisper-ja-760M-convert,whisper-punc-308. They exist only on the machine that made them, so a fresh clone (and the CI checkout) has no bundled fallback harness for those kernels — which is the exact failurecheck-kaggle-harness-sync.pywas written to prevent, and it cannot see it because it only compares copies that are present. Eithergit add -fthem or teach the checker to fail on a kernel dir with no bundle. - English
that/read/used/object/console/useneed a POS tagger — already recorded below with the measurement that says a full spaCy port buys 0.34%. The cheap slice is a closed-class rule forthatalone (~24 of the ~32 residual tokens, no model).
The landed part: DiarizeMethod::FoxNose + options exposed in the Rust crate
and Dart, and the real bug behind the report — the hand-maintained Rust and
Dart mirrors of the APPEND-ONLY crispasr_diarize_opts_abi were never updated
when #324 appended the FoxNose fields, so every diarize_segments call from
those bindings had the C side read 24 bytes past the caller's allocation.
Both mirrors now carry the 48-byte layout; crispasr-sys has a size/offset
layout test, the flutter smoke test pins the DiarizeMethod indexes, and the
c_api struct comment now lists every hand-written mirror to update on the next
append.
Still open from the issue's asks:
DONE (second #332 landing):crispasr_session_output_sample_rate()(+ channels getters).crispasr_session_output_sample_rate/input_channels/output_channelsin the C ABI (per-backend rate table mirroring the CLI adapters'tts_sample_rate(), 0 = no audio output), wired through Rust / Ruby / JS / Java / C# (the surfaces that exposeinput_sample_rate), pinned bytests/test-session-abi-nulls.cppand documented indocs/bindings.md§"Session audio-format getters". A new TTS backend must add its ctx to the getter's table —docs/contributing.md§5b.3 records the duty. Go / Dart / Python don't exposeinput_sample_rateeither; extend all four getters together there if anyone asks.Same-benchmark DER for the pyannote+embedder path— already existed; the note here originally claimed the pyannote path had no DER on the shared benchmark, which was wrong. The cross-method table lives in the #326 NOW section below and indocs/diarization-speakers.md"#326": pyannote+embedder 7.81 % vs foxnose 7.32 % mean DER on the same 8 VoxConverse dev files (whisper-tiny segments, 0.25 s collar), with the 3.18 %-vs-7.32 % foxnose discrepancy explained there (own turns vs ASR segments as speech regions). Nothing left to run for #332; the estimator under-count remains #326's open accuracy item.
Landed: parallel 60 s chunked pyannote inference (e517273d), the SPK_MASK
powerset transposition worth 15 DER points (15aad6f8), and the over-clustering
fix (a719c89d). See docs/diarization-speakers.md "#326".
Measured end to end on the 8 VoxConverse dev files, whisper-tiny segments,
0.25 s collar, tools/der_voxconverse.py:
| path | mean DER |
|---|---|
| raw posteriors, no clustering (the chunking A/B harness only) | 33.37% |
--diarize-method pyannote --diarize-embedder auto |
7.81% |
--diarize-method foxnose (#324) |
7.32% |
Reproduce either arm in one command:
python tools/der_voxconverse.py --prepare <voxconverse>/data --audio-dir /tmp/vox
python tools/der_voxconverse.py --audio-dir /tmp/vox --model ggml-tiny.bin \
--args "--diarize --diarize-method pyannote --diarize-embedder auto"
The cap no longer decides the count, but the BIC estimator that replaced it is wrong on half the shard in the other direction:
file GT hyp DER%
esrit 5 5 3.29
fsaal 7 6 3.80 under
jyirt 4 3 9.33 under
mesob 4 4 16.90 <- worst file, count is RIGHT
nnqfq 5 5 3.59
rcxzg 4 4 9.71
tiams 5 3 9.61 under
willh 2 3 6.23 over
Two separate threads, and it matters not to conflate them:
a. Count. 4 of 8 wrong, 3 of those under. HISTORY's estimator survey ("the estimator is not FRAGILE, it is BIASED") measured the same shape with margins of 6-14%, so this is not a tie-breaking problem and a silhouette tweak will not fix it. NME-SC was already tried and LOST (archived). Next candidate is calibrating the BIC penalty against embedding-window count, since the errors concentrate on short files. b. mesob has the RIGHT count and the WORST DER (16.90%). Nothing to do with counting — it is confusion within 4 correctly-estimated speakers. Diagnose separately; a count fix cannot touch it.
GATE for any count change: mean DER must beat 7.81% AND mesob must not regress.
Unclaimed. Everything needed is built and proven; this is blocked on disk, not on work.
Do NOT stamp parler or csm. Checked: the verdict table resolves parler-tts
and csm by BACKEND NAME (crispasr_speaker_identity_models.h), which is
already independent of the filename — renaming one changes nothing, so stamping
them moves 14 GB for no behavioural gain.
kartoffel-orpheus-de-{natural,synthetic} is the opposite case: one orpheus
backend serves several checkpoints, so the verdict keys on a filename substring
and a rename drops natural from real_person to unknown — silently losing an
Art. 50(4) disclosure. That is the case the stamp exists for.
Why it is not done. stamp-speaker-identity.py rewrites rather than patches,
so it needs source AND output live: 13.2 GB peak for the f16, 7 GB for the q8_0.
This machine has 11 GB on / and 5.9 GB on /Volumes/backups, and other agents
build here — filling the disk would take them down with it. Not worth a
rename-robustness gain.
When there is room, stream it one file at a time rather than
snapshot_download-ing 24 GB:
for f in kartoffel-orpheus-de-natural-{q4_k,q8_0,f16}.gguf; do
hf download cstr/kartoffel-orpheus-3b-german-natural-GGUF "$f" --local-dir src/
python models/stamp-speaker-identity.py --input src/$f --output out/$f \
--speaker-identity real_person \
--evidence "card: fine-tuned primarily on natural human speech recordings; 19 speakers extracted from podcasts/lectures/OER"
# verify tensors with a RAW BYTE comparison, then upload, then rm both
done
The synthetic sibling takes --speaker-identity synthetic and the evidence
"card: trained on synthetic German speech with emotion and outburst control".
⚠ Verify with .tobytes(), not np.array_equal — the latter returns False
whenever NaNs are present and will call a byte-identical file corrupt.
Scope narrowed after checking where the stamp actually buys anything. The
verdict table resolves parler-tts and csm by BACKEND NAME
(crispasr_speaker_identity_models.h), which is already independent of the
filename — renaming one of those checkpoints changes nothing. Stamping them
would move 14 GB for no behavioural gain.
kartoffel-orpheus-de-{natural,synthetic} is the opposite case: one orpheus
backend serves several checkpoints, so the verdict keys on a filename substring
and a rename drops natural from real_person to unknown — silently losing an
Art. 50(4) disclosure. That is exactly what the stamp exists to fix, so those
two repos (24 GB) get it and the others do not. Everything needed is built and proven, this is bandwidth.
Stamp crispasr.voice.speaker_identity into the repos whose verdict is
established but which were left for size: cstr/parler-tts-mini-v1.1-GGUF
(real_person), cstr/csm-1b-GGUF (synthetic),
cstr/kartoffel-orpheus-3b-german-{natural,synthetic}-GGUF (real_person /
synthetic). Upload only, no runtime code:
./models/stamp-published-voices.sh <downloaded-dir>
It asks crispasr --print-speaker-identity per file and skips unknowns, so no
verdict is restated. Verify tensors are byte-identical afterwards with a RAW
BYTE comparison — np.array_equal returns False whenever NaNs are present and
will report a good file as corrupt.
CrispTTS hash-chains its consent log (SHA-256 per line + a sibling .anchor
file, since a chain cannot detect truncation of its own tail; Art. 17 erasure
handled by re-chaining survivors plus a [CHAIN-REBUILT] marker). It is well
built. Porting it is still the wrong first move here, for four reasons:
- We would be chaining the wrong layer. Our record identifies the voice by
NAME (
voice=alice.wav), never by content — there is no hash of the reference anywhere in the tree. That file can be swapped a minute later and the record still "verifies". A chain protects the SEQUENCE of records; if each record is an unbound assertion, a perfectly chained log of them proves nothing. CrispTTS's log is worth chaining because it already ties to a hash. - The threat model does not support it. This is a self-attestation BY the operator, who controls the binary, the file and the anchor. A chain does not defend against the party it is recording. It defends against a third party editing the file without the tooling — the narrow case. CrispTTS says "tamper-evidence, not tamper-proofing" and is right; the risk is readers hearing the stronger claim.
- Persisting more creates the liability. CrispTTS had to build erasure, retention pruning and rebuild records BECAUSE they persist identifying data, then prove re-chaining cannot launder later edits. Every field stored is a field that must be deletable on request.
- Library vs application. CrispTTS and Susurrus are applications. CrispASR is a library + CLI embedded by others (four bindings, an HTTP server, Wyoming). The right shape for an embedded component is to emit a clean structured record and let the host own durability; a bespoke chained log duplicates journald/CloudWatch/SIEM and is weaker than any of them.
Real tamper-resistance is a STORAGE decision — append-only permissions, object-lock/WORM, shipping off-box — and the library must not pretend to substitute for it.
ref_sha256=in every[CONSENT]line. SHA-256 of the file the backend will actually open (resolve_voice_path()output), via the header-onlycrispasr::shaalready vendored for C2PA. This is the change that turns the record from an assertion into evidence, and a hash carries far less data-protection weight than the recording or a name.- Correlate the record to the output it authorised. A per-process
run_idin the[CONSENT]line and in the post-synthesis audit line, so a disputed clip can be walked back to the attestation. On the server add a request id — today a consent line and its output are unlinkable. --consent-log <path>, JSON Lines, default off. Today the only route is redirecting stderr, which interleaves the record with model-load noise and progress output. A separable sink is what lets operators route it into infrastructure that IS append-only.- Document the division of responsibility in docs/eu-ai-act.md: the operator is the controller, this is their artefact, tamper-resistance is theirs to provide, and here is the field to key erasure on.
Hash chaining. The one real case is the Docker server run as an appliance where no host audit infrastructure exists and the operator wants to show a regulator the log was not casually edited. Revisit only then, and only after (1).
⚠ Legal note: as read here, Art. 50(4) is a DISCLOSURE duty, which CrispASR already discharges unconditionally through marking. The consent attestation goes to GDPR lawful basis, where the accountability duty sits with the operator as controller, not with the tool. Engineering judgement, not legal advice — put it to counsel before it appears in a compliance claim.
-
C# and WASM bindings are source-only-verified. No
dotnetor emsdk on the Mac, so theIntPtr/PtrToUtf8marshalling inbindings/csharp/CrispASR/NativeMethods.csand the twoemscripten.cppsites were not compiled.bindings-csharp.ymlis the real check — watch that job. -
Three ASR backends still unaudited for
CAP_PUNCTUATION_NATIVE:lfm2-audio,fastconformer-ctc,wav2vec2— no local GGUFs. The CTC pair almost certainly needs the pass (unpunctuated by construction);lfm2-audiois an LLM decoder and is the likely one to need the flag. Method:FIREREDPUNC_DEBUG=1 … | grep PUNCDBGand readin=— do NOT use--no-punctuation, it strips after the fact and inverts the answer. -
Delete the duplicated fallback copies (READY — this is "Ready to take" #1).
Fourteen
.cpp/.hfiles exist twice: once incrisp_punc/src,crisp_lid/src,crisp_truecase/src(what the shared libraries build, and what CrispEmbed consumes viaadd_subdirectory) and once insrc/(whatsrc/CMakeLists.txtbuilds when those directories are absent from a checkout). They are the same implementations twice.tests/test-copies-in-sync.cppnow makes the duplication safe — it byte- compares all 14 pairs and asserts the pair list is exhaustive — but safe is not the same as gone. Two copies had already drifted before that test existed (see HISTORY 2026-08-03), and #308's capitalisation fix was dead code for months for the same reason.The task: make the fallbacks unnecessary, then remove them.
- Find out whether a partial checkout without
crisp_punc/,crisp_lid/,crisp_truecase/is still a real distribution mode.src/CMakeLists.txtbranches on their absence — check whether anything actually ships that way, or whether the submodules are always present now. - If nobody needs it: delete the
src/copies, drop the CMake fallback branches, and deletetests/test-copies-in-sync.cppwith them. A guard for a hazard that no longer exists is dead weight. - If someone does need it: the copies must stay, and so must the test. Say who needs it, in this file, so the next person does not re-derive it.
Do not simply delete the
src/copies and see if the build goes green — the fallback branch only compiles when the shared directories are missing, so a normal build will pass either way. Test by moving the three directories aside and configuring from clean. - Find out whether a partial checkout without
Rounds 1-2 complete (HISTORY). Open, beyond the "carried out of #316" section:
- POS-dependent words (
that/read/used/…): the tag-dependent remainder is 0.34 %; the cheap slice is a closed-class rule forthat(no spaCy port). - misaki's reduced vowels
ᵊ/ᵻare not modelled (context rule needed).
Started from #305 (reporter: whisper-vad-asmr + firered-vad slow / single-core). The recurring pattern: the model compute is fine (ggml threads / Accelerate BLAS), but the audio FRONT-END (STFT/mel/fbank) is a scalar single-threaded per-frame loop — and for long audio the mel can out-cost the encoder (measured in whisper-vad). Parallelizing over the independent frame axis (per-thread FFT scratch, reductions keep order) is bit-identical and a large win.
DONE (per-backend, std::thread, gated + bit-identical):
- whisper-vad-encdec — GPU move (3.3×) + parallel mel (~30% long-audio) + requant
(b5eb7556 / 9f867909 / 8324719e).
CRISPASR_VAD_ENCDEC_*. - firered-vad — parallel DFSMN convs + fbank FFT, ~3.7× (8c1d6a10).
CRISPASR_FIRERED_VAD_SERIAL=1. - marblenet-vad — parallel mel front-end, ~1.6× (7b2f11bc).
CRISPASR_MARBLENET_VAD_SERIAL=1. - Audited silero (whisper.cpp native ggml, already threaded) + webrtc (subband GMM, no FFT) — no change needed.
STATUS: the CPU front-end parallelization vein is MINED (2026-07-26). Full sweep done. Nothing clean+validatable-locally remains — do NOT keep sprinkling std::thread:
core/istft.h(shared by outetts_wavtok/kokoro/cosyvoice3 +8 vocoders) is single-threaded BUT it is overlap-add — adjacent output frames write overlapping samples (data race, unlike the mel's disjoint writes), and n_fft is tiny for most consumers (kokoro=20, cosyvoice3=16) so the per-frame IRFFT is already cheap. Poor risk/reward — SKIP (would need per-thread out buffers + merge or a stride-coloring scheme for a marginal win). 2026-08-30 re-audit: the "miocodec has a scalar FFT" line below was misattributed —src/miocodec.cppcontains no FFT of its own; its onecore_istft::istftcall sits behindmiocodec_extract_stage, whose only caller is the diff harness (decode is a stub;miotts.cpp:1770calls core_istft directly). Measured miotts share: iSTFT ≈ 2.0% of synthesis CPU (n_fft=392/hop=98/T=1332 — 3.5 s of 176.6 s) — under the 5% gate, so still SKIP. For the record, a BIT-IDENTICAL parallel design DOES exist if a heavy consumer ever appears: parallelize onlyirfft_hermitianinto a T×n_fft frame buffer (each frame self-contained, ~2 MB scratch at miotts sizes, ~99% of the stage) and keep the overlap-add strictly serial — the naive frame-parallel OLA is both a race and FP-non-associative, and per-thread buffers + ordered merge does NOT restore bit-identity, but the irfft split does by construction. Gate asCRISPASR_ISTFT_SERIAL=1if ever done.- Own-mel backends NOT on core_mel (f5_tts, gemma4_e2b, titanet, chatterbox_s3gen, outetts_wavtok, ecapa_lid): their FFT runs ONCE on the reference clip or on TTS output where the DiT/decoder dominates — marginal fractions, not the per-long-audio bottleneck the ASR/VAD mel was. Not worth the churn.
- moonshine uses a shared mel helper (no local scalar loop of its own). The remaining lever is TIER 2 (GPU ports) only — see below + the handover prompt.
KEY FINDING that reframes the fleet rollout: core/mel.h::compute (used by
~28 ASR/TTS backends: parakeet, canary, canary_ctc, nemotron, qwen3_asr, cohere,
glm_asr, granite_, higgs_stt, ark_asr, voxtral/4b, moss_, mini_omni2,
lfm2_audio, qwen3_tts, cosyvoice3, chatterbox, indextts, mimo, piano_transcription
…) ALREADY has a §176f parallel-STFT path — but it is #ifdef _OPENMP only,
and this macOS/AppleClang build has OpenMP_CXX_FLAGS=NOTFOUND (no libomp), so
it is compiled out on macOS (and any non-libomp build). That is why ZERO
backends set allow_parallel_stft=true — the flag is a no-op on the dev box. The
mel projection matmul in compute() is serial too (not even OpenMP-gated).
TIER 1 — port core_mel STFT to std::thread (PORTABLE) — DONE (528d672f). One
change to src/core/mel.cpp: portable std::thread STFT path (OpenMP kept when
_OPENMP), DEFAULT parallel with CRISPASR_MEL_SERIAL=1 opt-out, threshold T≥256.
Bit-identical (disjoint power[] rows, same reduction order). The mel projection was
already Accelerate/BLAS-threaded, so the STFT was the sole single-threaded piece.
Validated parakeet-tdt-0.6b-ja (M1): STFT 1070-frame chunk 20.19→5.52 ms (~3.7×),
transcript BIT-IDENTICAL serial vs parallel. Lifts all ~28 core_mel backends.
Neutral for heavy-LLM-decoder ASR (mel is a tiny fraction); real win for
encoder-bound ASR (parakeet/canary/nemotron) and long audio. allow_parallel_stft
per-backend flag now redundant (kept for back-compat).
TIER 2 — CPU-hardcoded backends that could go GPU (per-model, MAJOR effort — NOT
started; needs models + Kaggle validation). ggml_backend_cpu_init() with no GPU
path, conv-heavy enough to benefit: mel_band_roformer (source sep — STRONGEST,
already promised as a GPU port under #296 below; but 1284 lines already
BLAS-optimized, model not local, and #296 was Kaggle-validated → a full
transformer+iSTFT ggml-graph port is a focused multi-hour project that must be
diff-harness + Kaggle validated, NOT doable+trustworthy locally), piano_transcription
(cblas=0 — not even BLAS yet; a cheaper first step is BLAS/threads before GPU),
pyannote_seg (already uses threads), openvoice2, TTS codecs (miocodec has a scalar
FFT — VAD-style parallelize; miotts/tada_encoder). Small classifiers (ecapa_lid,
lid_fasttext, bert_encoder, marblenet) stay CPU (launch-bound). Each needs a
ggml-graph port + Metal/Vulkan landmine handling + diff-harness validation. RECOMMEND
as a dedicated session per model (get the model, port, validate on Kaggle), starting
with mel_band_roformer (highest impact, already promised).
TIER 3 — mined. §176 runtime-opt campaign is 18/20 DONE; the 2 open are <2% (measure-first). Do not chase.
--separate (mel-band-roformer) was ~24 min for 11 s on Linux/Windows (fast only
on macOS/Accelerate). Root cause: CPU-only forward with an Apple-only BLAS gate +
naive O(N²) iSTFT + scalar attention + redundant weight de-quant. FIXED on main to
24 min → 56 s (~26×), cos=1.0 via: portable OpenBLAS linear(), FFT iSTFT,
attention→SGEMM + BLAS-thread-pin, per-layer weight hoist. Validated on Kaggle
(cos + per-stage profile each round). Replied on the issue
(#296 comments 5062263388 / 5062697259).
Explicitly PROMISED on the issue — must follow through:
- OpenBLAS in the shipped binaries — AUDITED against the published v0.8.25
artifacts, not the workflow source (
b0c548e7).- Windows: DELIVERED.
crispasr-windows-x86_64-cpu.zipshipsopenblas.dll(1.78 MB) beside the exe. The reporter's platform has the fast path; this bullet's old "Windows CLI jobs set up NO BLAS" is stale. - Linux: was BROKEN, and worse than a slow fallback. apt's OpenBLAS is
dynamic, so the binary carries
NEEDED libopenblas.so.0while the tarball shipped only crispasr, crispasr-quantize and libc2pa_c.so — on a host without OpenBLAS it does not start at all (error while loading shared libraries). Read off the real tarball with objdump. Fixed byscripts/bundle-openblas.sh(the OpenBLAS half ofbundle-c2pa.sh; RUNPATH is already$ORIGIN), wired into the x86_64 + arm64 Package steps. - Both platforms could ship a degraded binary on a GREEN run — the Windows
vcpkg step is
continue-on-errorand CMake falls back to scalar silently. Both Package steps now assert the ARTIFACT and fail with::error::. ⚠ UNVALIDATED until the next release run: the bundling copy branch cannot execute on macOS. The no-op branch, the detection branch against the real shipped Linux binary, and the YAML parse were verified locally.
- Windows: DELIVERED.
- GPU / ggml-graph port (the real long-term fix, promised as "a GPU path is
tracked"). Rewrite the forward (transformer + iSTFT) as a ggml graph → SIMD +
threads + GPU everywhere, eliminating the scalar/BLAS/OpenMP scaffolding. Same
Apple-only-BLAS pattern also slows htdemucs (the other
--separatebackend) on Linux/Windows — the ggml port is the template that fixes both. Needed for full-song latency (56 s/11 s still extrapolates to ~15 min/song). Validate with the mel-band-roformer diff harness (per-stage cos≥0.9995).
The #266 rework (closed-roster, cluster-level speaker identification — DONE,
see docs/speaker-db-clusters/PLAN.md + HISTORY) left one parked item (F9):
the diarize → merge → global-cluster → identify orchestration lives in the CLI
(crispasr_apply_global_speaker_stages(), examples/cli/crispasr_run.cpp).
Session-ABI/bindings compose the primitives themselves; the server exposes no
speaker-db (deliberate). If a binding or the server ever needs the full named
pipeline, hoist the orchestration into src/ (the parakeet_orchestrate
pattern / multi-surface lesson) rather than duplicating it per surface. The
compliance invariants (consent + --expect-speakers closed roster + post-only)
must move with it — see docs/diarization-speakers.md §2.
Pending roadmap items. Each is self-contained with files, approach, and
effort estimate. Completed items have been moved to HISTORY.md.
Numbering convention:
§Nrefers to PLAN items (sections in this file).#Nrefers to GitHub issues on CrispStrobe/CrispASR. They are independent sequences and numbers may collide. When in doubt, PLAN items are always written as§Nand GitHub issues as#N.
Latest release: v0.8.32 (tag v0.8.32, + pub.dev crispasr 0.8.32). Release
notes live on each tag. Only the CURRENT version's RELEASE_NOTES_v*.md is kept
at the repo root — older ones are removed when the next is written, since each
tag carries its own copy and the GitHub release carries the published body.
release.yml hard-fails if the file for the tag being cut is absent.
The v0.8.18 → v0.8.20 train shipped in one session (2026-07-21): each patch was
a genuine fix that only surfaced on a real tag run — see "Recent completions"
and LEARNINGS.md "a green release job is not a shipped artifact".
Recent completions (2026-07-22):
- #292 --max-new-tokens ignored — SHIPPED. moss-diarize (and 9 more ASR
backends) hardcoded the decode cap, so
--max-new-tokensdid nothing and a long single-pass truncated (reporter's 300 s file stopped at 164 s). Fixed across all 10: contextmax_new_tokensfield defaulting to the old constant (no regression) + setter + cap/KV read it; CLI forwards only whenmax_new_tokens_explicit; session C-ABI forwardss->max_new_tokens. moonshine excluded (its 194 is an architectural short-form limit). Also addedchunk_idto segments (part 2: diarize speaker labels are chunk-local). CUDA-validated on the reporter's exact backend (moss-diarize q4_k):--max-new-tokens64→4096 raised output 216→366 words CPU / 232→382 CUDA, chunk_id 4 distinct, tabcnn CPU/CUDA parity 0 mismatches (6/6, ~18 min,tools/kaggle/cuda-292-maxnewtokens/). SeeLEARNINGS.md"a hardcoded decode cap" + memory [[kaggle-full-harness-regime]].
Recent completions (2026-07-21):
- #290 canary-qwen long audio — SHIPPED (v0.8.19). The backend declared
CAP_INTERNAL_CHUNKINGwith no chunker (src/canary_qwen.cpphas zero chunking code vs 62 hits in parakeet.cpp), disabling BOTH dispatcher safety nets → one full-length encoder pass → O(T²) attention (384 MiB→10.2 GiB) and sparse output. Fix: drop the false capability. SeeLEARNINGS.md"a capability flag is a promise". - Lib-delivery bugs (CometBeat) — SHIPPED (v0.8.19 rpath, v0.8.20 flat+gpu).
6 of 7 lib bundles were unloadable as delivered: macOS baked the CI runner's
build path into LC_RPATH; all 5 Linux bundles used
$ORIGIN/../../ggml/src(one level too high). The old gate only checked deps were PRESENT, never that the loader could FIND them. Newtools/verify-lib-bundle.shrelocates + dlopens;tools/package-lib-bundle.shflattens tolib/+ rewrites rpaths. SeeLEARNINGS.md"presence is not resolvability". - iOS shipping for the first time — SHIPPED (v0.8.20). Added a
build-xcframeworkjob to release.yml (build.yml is tag-excluded by design); fixedbuild-xcframework.shto includelibglint.ain the combined archive. - #291 C# binding neglect — SHIPPED (v0.8.20). VadSegments returned
centiseconds while documented as seconds; added
CrispASR.Logging; bound the 7 task backends (tab/beats/chords/piano/pitch/separate/convert) inSessionMusic.cs; added the first-ever C# CI. C# was also missing from the contributing.md binding-parity list. See memory [[csharp-binding-neglect]].
Recent completions (2026-07-17):
- Roadmap accuracy sweep — audited the PLAN's OPEN/NOT-STARTED headers against
the code; ~11 items were already shipped but still marked open and are now
corrected: §169 (qwen3-asr ChatML prompt), #128 (Piper TTS), §66 (pub.dev
crispasr 0.8.11), python_find_lib, #60o (MTLBinaryArchive pipeline cache), §155 (CONV_TRANSPOSE_1D — all phases + Metal/Vulkan/CUDA col2im kernels), #58 (MOSS-Audio-4B), #101 (OmniVoice), §229 (GGML_LLAMAFILE ON), plus §57/§106/§224/§247 sub-items. SeeLEARNINGS.md"verify PLAN OPEN items against code". - #227 VAD boundary reuse — SHIPPED (CLI
--vad-export/--vad-import+ servervad_export/vad_import); shared serializer, unit-tested. - #91 CLI parity —
--offset-t/--durationnow honoured by every backend on both CLI + server (sharedcore/audio_window.h, unit-tested);--print-confidencefixed for non-whisper backends (was a parsed-but-dead flag). - #201 TADA on-the-fly voice cloning — in-memory make-ref (no temp GGUF) on
both the C-ABI/session and server/adapter surfaces, opt-in
(
CRISPASR_TADA_WAV_CLONE=1). Roundtrip gate pending (see below). - Docs — new
docs/benchmarking.md(§227 fair-measurement recipe);-amaligner aliases enumerated indocs/cli.md(§105).
Recent completions (2026-07-11):
- #242 moss-diarize: SHIPPED — joint ASR + diarization + timestamps, 0.9B model,
diff harness 4/4 cos=1.0, GGUFs on HF, full 12-point checklist. See
HISTORY.md. - #200 dots-tts PatchEncoder: FIXED — added missing RoPE (theta=10K) + QK-norm;
full pipeline now runs e2e (LLM → DiT → PEnc → vocoder → WAV), ASR roundtrip passes.
Wired
--tts-steps/--tts-cfg-scale. GPU already supported. SeeHISTORY.md. - #215e gallocr UAF audit: DONE — all 19
cached_*_gfsites audited across the codebase. 7 backends fixed (canary, canary_ctc, kyutai_stt, moonshine_streaming, nemotron, paraformer, sensevoice). SeeHISTORY.md. - Generation-health gate: DONE —
src/core/generation_health.hwith 5 checks- 16 unit tests. Non-breaking additive.
- qwen3-tts-perf (#245): ANALYZED — profiled on CPU: dispatch overhead (build+ reset+alloc = ~5ms) is <0.1% of per-frame cost (5000ms compute). The O15 path already caches the graph and uses a dedicated sched. Skip-realloc is broken on CUDA+Metal. The bottleneck is pure matmul compute; perf wins require GPU where the ~5ms overhead becomes significant. Handover removed.
- Untrusted-input parser hardening: SHIPPED — multi-agent security audit of the
audio demuxers + GGUF loader found 6 DoS/OOB defects (MP4 stsz/stco/co64 count +
co64 offset overflow, WebM lacing, WAV + AU size clamps, GGUF split mmap bounds),
all fixed + ASan-validated. See
HISTORY.md+LEARNINGS.md.
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: Upstream the ggml empty-key fix to ggml-org (outbound public PR, left for a human)
FIXED 2026-08-10 — the stb_vorbis SEGV linux-fuzz-smoke found.
vorbis_deinit walked comment_list_length entries of a NULL comment_list.
comment_list_length is read straight from the file at
examples/stb_vorbis.c:3660 and set BEFORE the array is allocated, so the
allocation-failure return one line later (an attacker-sized count makes
setup_malloc return NULL) left the two out of step. Reachable from the public
crispasr_audio_load on untrusted input.
⚠ An earlier CrispASR patch had already hardened the SIBLING path here — the
partially-filled array whose unassigned slots were freed as if valid — by
zeroing the allocation. It could not help when the allocation never happened.
Fixing the case you can see and leaving its neighbour is the recurring shape:
the guard now lives in vorbis_deinit (defends every path in, present and
future) rather than at the Nth caller, plus the length is reset on the error
path so the struct's invariant holds.
Reproduced deterministically rather than waiting for the fuzzer. The crash
input is 102 hand-crafted bytes — an Ogg page carrying a Vorbis ident header
and a comment header declaring comment_list_length = 0x3FFFFFFF. Kept as
tests/fuzz/regressions/ogg-huge-comment-count.ogg, and the smoke-fuzz job now
copies tests/fuzz/regressions/ into its corpus, so every fixed crash becomes a
deterministic gate instead of a coin flip. Before/after on matching builds: 2
SEGV lines → 0.
⚠ libcrispasr is a SHARED library — the fuzz harness picks it up by rpath, so
"rebuild the harness" tests the new code with an old-looking binary. Rebuild the
dylib when doing a before/after, or the control silently becomes a second copy
of the experiment. (Cost me one invalid control here.)
The job also now uploads the crashing input on failure. It previously kept only the stack trace, which for a stochastic job means the reproduction is gone.
Noted, not fixed: on macOS the same input flows past stb_vorbis into the
AudioToolbox fallback (crispasr_at_decode → ExtAudioFileOpenURL) and
libFuzzer reports an out-of-memory there. That is a resource limit inside
Apple's decoder on malformed input, not memory corruption, and that path does
not exist on the Linux CI. Worth a size sanity-check before handing a file to
AudioToolbox if anyone wants it.
Both validate the TTS port against the Python reference. Neither is claimed.
What: Dump talker LLM logits at each generation step in both Python reference and C++ runtime, compare — validates the text-decoder input so the sampler is a faithful port over verified logits.
How: (a) add talker_logits_step_N capture to the Python reference dumper (tools/reference_backends/qwen3_tts.py) via a generate-time hook; (b) add matching C++ stage to crispasr_diff_main.cpp; (c) run on the TTS diff harness.
C++ SIDE DONE 2026-08-10 (#337), and it settled the report. Three levers,
which only work together:
CRISPASR_QWEN3_TTS_GREEDY=1 (argmax), CRISPASR_QWEN3_TTS_REPLAY_CODES=<file>
(16 codec ids per frame — teacher forcing), and
CRISPASR_QWEN3_TTS_DUMP_LOGITS=<dir> (raw per-frame logits, before the
repetition penalty and suppress mask).
Result: CPU vs Metal, fully pinned, worst cos 0.999870, mean 0.999940, 0/49 argmax disagreements. Given identical history the backends agree at every step, so the free-running divergence is entirely trajectory divergence seeded by ~1e-4 arithmetic. That is the rigorous version of the #337 verdict; the earlier free-running comparison suggested it but could not separate the two.
⚠ Partial teacher forcing is a trap. Replaying codebook 0 alone leaves the
15 residual codebooks sampled (code_pred_generate_15 MUST sample — greedy
there is documented to produce silent output), so the per-frame input embedding
still diverges and the "teacher-forced" diff bottoms out at cos 0.849 — pure
artefact. Pin all 16 or measure nothing.
⚠ Every per-synthesis dump must be tagged. One --tts run generates twice
(the utterance, then the spoken AI disclaimer). talker_%04d.f32 and
generated_codes were both frame-only names, so the disclaimer OVERWROTE the
utterance's dumps and a directory held two utterances with no way to tell.
Both now carry _s%02d. This produced one entirely bogus cross-backend table
before it was spotted — the tell was argmax_gpu[k] == argmax_cpu[k-2].
Python reference side DONE 2026-08-10. _hooks.capture_per_call() written
(the module referenced an _iter_capture that never existed) — one capture per
call instead of first-call-only. tools/reference_backends/qwen3_tts.py now
emits talker_logits_step0..15 and, at last, generated_codes: that stage had
been in DEFAULT_STAGES since the backend was written and was NEVER produced,
because the ids only exist inside generate_voice_clone (the outer call returns
audio). Captured by wrapping tts.model.generate.
Verified by running it: 17 tensors, generated_codes (19, 16) int32 — exactly
the layout CRISPASR_QWEN3_TTS_REPLAY_CODES consumes — and
talker_logits_step0 (1, 147, 3072) = the PREFILL, with steps 1+ the AR steps
at (1, 1, 3072). Note that indexing when diffing: step0 is not frame 0.
⚠ Env: the shared conda base has transformers 5.x; upstream Qwen3-TTS pins
4.57.3 and importing qwen_tts against 5.x dies on check_model_inputs().
Do NOT downgrade the base. tools/reference_envs/qwen3-tts/requirements.txt
now scaffolds it; a --system-site-packages venv inherits torch and shadows
only transformers. qwen_tts comes from the clone at ~/code/Qwen3-TTS via
PYTHONPATH, not pip.
Prep landed 2026-08-10: qwen3_tts_sum_frame_embed() factors the 16-codebook
embedding sum out of the AR loop so a harness entry point can build the SAME
per-step talker input without a second copy — duplicating it is how a harness
drifts from the runtime it checks, which is what #338 was. Verified
output-neutral: PCM byte-identical before/after (the WAV bytes differ only in
the C2PA/watermark metadata, which carries a per-run id — compare PCM, not the
container, when checking a refactor here).
Still open, and here is the actual blocker. A per-step talker input is NOT just the codec-embedding sum:
next_emb[step] = sum_{cb=0..15} embed_cb(frame[cb]) + trailing_text_hidden[step]
The trailing term is prompt-derived state computed from the synth text inside
the generate path. So a harness stage cannot simply prefill from the
reference's talker_inputs_embeds and then step — it also needs trailing,
which the reference does not currently dump and the runtime does not expose.
That dependency is why this stage does not exist yet; it is not just wiring.
Two ways forward, pick one before writing code:
- Dump
trailing_text_hiddenas a reference stage too, and pass it into aqwen3_tts_talker_logits_replay(embeds, n_tokens, codes16, n_frames, trailing, n_trail)entry point. Most faithful, and it makes the dependency explicit in the archive. - Have the runtime construct the whole prompt itself from the same text +
voice wav the reference used, and use the reference's
talker_inputs_embedsONLY as a structural gate (cos ≈ 1 before trusting any logits — the guide's input-alignment rule). Less plumbing, but it assumes the two prompt builders agree, which is the thing being tested.
Option 2's gate is worth having either way.
The superseded note: CRISPASR_QWEN3_TTS_DUMP_LOGITS=<dir>
writes raw per-frame talker logits (f32), dumped BEFORE the repetition penalty
and the suppress mask so a diff isolates the forward from the sampling policy.
It already paid for itself — it is what proved the reported "GPU miscompute"
was cos 0.99992 at frame 0, i.e. ordinary backend arithmetic amplified by the
AR loop. What remains is (a): the Python reference hook, which turns a
cross-BACKEND comparison into a cross-IMPLEMENTATION one against the blueprint.
Test: needs a TTS model (qwen3-tts or tada). 0.6B Q8_0 (941 MB) fits on VPS; TTS gen slow on CPU (~105x RTF) — use short "Hi." input, 2-3 frames.
What: Dump Python's sampled token IDs and replay them in C++ (instead of re-sampling) so sampling-enabled downstream stages diff deterministically despite torch-vs-mt19937 RNG mismatch.
How: (a) Python dumper captures sampled_token_ids as a 1D int32 tensor in the reference GGUF; (b) C++ diff harness reads them and feeds the backend's step function instead of sampling; (c) compare downstream stages (codec, vocoder) vs the Python reference that used those tokens.
Files: tools/reference_backends/<tts_backend>.py + crispasr_diff_main.cpp.
Requested by the CometBeat opus (voice-svc) agent via its docs/PLAN.md
coordination note (read 2026-07-20). CometBeat is splitting a singing-voice-
conversion stack into pure-Dart (lightweight/offline) and native-via-CrispASR
halves. We own the real-time-critical heavy vocoders; they keep the
HuBERT/ContentVec encoder, Harvest F0, and a lightweight DDSP-SVC synth as the
web/offline fallback.
Converter -> numpy spec -> ggml graphs -> convert() -> session C ABI, every
step at cos 1.00000000 against torch. crispasr-diff rvc <model> <ref> <wav>
reports 48 stages, including convert_e2e which runs the real
rvc_svc_convert() and reproduces the reference audio from ContentVec + F0 +
speaker id + replayed noise alone. Live test: 4 cases / 12825 assertions.
Deliberately NO CLI verb — the input is ContentVec features, which we do not
produce. docs/bindings.md documents the session surface.
ARTIFACTS: stored in the PRIVATE repo cstr/rvc-svc-GGUF (verified
private: True server-side before upload) — rvc-40k-f32.gguf (105 MB) plus
rvc-ref.gguf (23 MB), the 60-stage parity reference. That reference is the
valuable half: crispasr-diff rvc <model> <ref> <wav> re-runs all 48
comparisons with no torch, no checkpoint and no RVC repo.
REMAINING (packaging, not correctness):
- Registry entries. BLOCKED, and deliberately so on two counts: the checkpoint's
licence is unscoped (converted
--license other), AND the repo is private, so a registry URL would fail auto-download for anyone without auth. Scoping the base checkpoint's terms and re-stamping the tag comes first. - F32 ONLY. An f16 GGUF converts but the graph is F32-only today (ggml_scale is
F32-only; an F16 operand reaching ggml_add trips
GGML_ASSERT(src1->type == GGML_TYPE_F32)). It REFUSES rather than producing subtly wrong audio, and is noted as such in the header. - Dart/Flutter wrapper for
crispasr_session_convert*(CometBeat is the consumer; the C ABI they need is done). - Two agreed parameters were CORRECTED rather than implemented —
protectis provably inert without the FAISS index, andrms_mix_rateneeds the source waveform our seam never receives. See SVC_RECORD_SHAPES §9b.
Contract CONFIRMED (see docs/music-transcription/SVC_RECORD_SHAPES.md).
Source traced; see docs/music-transcription/RVC_BLUEPRINT.md. Two findings
change the job before any code:
- The ask is ~3x bigger than "the vocoder". The seam CometBeat wants
(features + F0 + speaker -> audio) is
SynthesizerTrnMs768NSFsid.infer()(models.py:664), which runs a transformer encoder (enc_p, with ann.Embedding(256, ...)on the coarse pitch), a normalizing flow (4x coupling + flip, reversed), AND the NSF vocoder. Only the third is HiFi-GAN. Ask CometBeat whether they want all three or onlydec— if onlydec, they must sendz, not ContentVec features, which changes the contract we just froze. - Inference is STOCHASTIC, at two independent sites: the latent sample
z_p = m_p + exp(logs_p) * randn * 0.66666(models.py:684), and SineGen's random initial phaserand_ini(:325) plus additive noise (:358). So output is not reproducible run-to-run, waveform correlation is an invalid acceptance test, and the diff harness MUST replay the reference's noise (same pattern asinput_featin btc /input_audioin mel-band-roformer). Settle this before the graph, not after.
Numerical hazards already spotted (phase cumsum drift in f32, F.interpolate
mode, upp = prod(upsample_rates) derived not configured, noise_convs
stride schedule) are itemised in the blueprint.
What: port the RVC NSF-HiFi-GAN generator to ggml/native-FFI. CometBeat
feeds us ContentVec features (from their hubert.dart), F0 (from RMVPE —
already done on their side) and a speaker id; we return converted audio.
Seam: they expect a CrispasrSession.convert(...)-style entry point. That
is a THIRD task-shaped surface after --separate/--pitch/--chords, so it
follows the same pattern: its own session entry points rather than riding on
transcribe(), plus the CLI dispatcher and the wasm/Go arms. Note the
crispasr_detect_backend_from_gguf trap — register the arch there too, not just
in the CLI (see the BTC entry in docs/music-transcription/PLAN.md).
Blocking coordination — DRAFTED, awaiting their reply: see
docs/music-transcription/SVC_RECORD_SHAPES.md, a concrete proposal for every
record shape (layout, dims, frame rate, F0 units, unvoiced encoding, speaker id,
return PCM) with reasons, so it can be accepted or amended rather than discussed
in the abstract. Items we have not yet verified against the RVC reference are
marked [UNVERIFIED] and must not be built against. The sharpest open question is
§3: who resamples F0 onto the feature timebase. The feature/F0 record shapes
must be agreed with the opus agent BEFORE their API freeze. Pin down, in writing: ContentVec feature
rate + dimensionality + dtype, F0 units (Hz vs cents vs MIDI) and hop, whether
F0 is voiced-masked, speaker-id encoding, and the sample rate of the returned
audio. Do this first — it is cheap now and expensive after the freeze.
Licence: RVC's own code is MIT, but the WEIGHTS in circulation are a mess (many community models are of unclear provenance, and some RVC forks carry non-commercial terms). Scope licences per-checkpoint before shipping any registry entry, exactly as the music-transcription scoping pass did.
What: port Beatrice v2; its low-latency design suits the native path.
Licence: MIT — this entry previously said "custom/NON-COMMERCIAL" and that was
wrong. fierce-cats/beatrice-trainer ships LICENSE = MIT (Copyright (c) 2024
Project Beatrice), and its README states in terms that "このリポジトリ内の
ソースコードおよび学習済みモデルは MIT License のもとで公開されています" —
the source and the trained models. So no acceptance gate is needed; a
cc-by-nc-sa-4.0-style tag here would have withheld permission the licence
actually grants.
Two real non-MIT signals exist but neither applies to this port: beatrice.lib
(the closed inference engine used by the VST "under permission") is irrelevant
because we port from the MIT source, not that binary; and the "Beatrice JVS
Corpus Edition" carve-out is a different distribution, not this repo.
Feasibility: CONFIRMED — architecture and weights are both published. An
earlier WebFetch summary of this same repo claimed it held "training scripts
only, not inference code or the model architecture". That summary was false;
beatrice_trainer/__main__.py is 4519 lines and defines the entire path
(PhoneExtractor, PitchEstimator, VectorQuantizer, ConverterNetwork,
Vocoder). Another instance of [[blueprint-summary-is-not-the-source]] — the
file listing settled in one call what the summary got backwards.
Weights (assets/pretrained/), all MIT, are split per component — the
obvious "load the checkpoint" assumption fails:
| file | contents |
|---|---|
122_checkpoint_03000000.pt (14.7 MB) |
phone_extractor only |
104_3_checkpoint_00300000.pt (7.1 MB) |
pitch_estimator only |
151_checkpoint_libritts_r_200_02750000.pt.gz (153 MB) |
net_g (177 tensors, the multi-speaker LibriTTS-R base) + net_d (discriminator, training-only → skip) |
Scope is larger than §CB1, and the wire contract is different.
ConverterNetwork.forward takes raw waveform (x: [batch, 1, wav_length],
plus target_speaker_id, formant_shift_semitone, optional
pitch_shift_semitone) and returns 24 kHz audio. Beatrice therefore owns phone
extraction and pitch estimation itself — unlike RVC, it needs no ContentVec
from the caller, which simplifies CometBeat's side but means porting three
networks rather than one.
Non-obvious details to verify while reading (each a silent bug if assumed):
CausalConv1d/WSConv1d— weight-standardised causal convs, not standardnn.Conv1d. The standardisation is part of the forward pass.VectorQuantizeris injected intophone_extractor.headvia a forward hook (enable_hook). This session already lost time twice to hooks that silently never fire; the numpy spec must confirm the hook is active by asserting the quantised path changes the output, not by trusting it ran.embed_quantized_pitchis a fixed sinusoidal table (built in__init__,requires_grad_(False)) — it may or may not be present in the checkpoint, so the converter must rebuild it rather than assume it was saved.key_value_speaker_embeddingis initialised with every speaker row copied from row 0, so speakers look identical until trained — an untrained-looking A/B is not necessarily a port bug.self.melspectrogramsis loss-only; it is not on the inference path.- Output rate is hardcoded 24000 with
hop_length = 24000/100— the same 100 Hz frame rate §CB1 uses, so the record shapes carry over.
Progress: two of three components ported and validated.
| component | state |
|---|---|
PitchEstimator |
DONE — crispasr-diff beatrice, 30 stages + e2e, 0 failed |
PhoneExtractor |
DONE — crispasr-diff beatrice-phone, 69 stages + e2e, 0 failed |
ConverterNetwork + Vocoder |
blueprint read done; converter and graph NOT started |
Next step: the ConverterNetwork/Vocoder port. It is the largest of the three and differs in kind from the two frozen extractors:
- The vocoder is a source-filter / impulse-response synthesiser, not a
HiFi-GAN — a 512-tap IR per frame plus aperiodicity and post-filter, driven by
overlap_add. That is why Beatrice is real-time-cheap. - It uses
WSConv1d/WSLinear, which neither ported component does, so the unbiased-variance trap (torch.var_meanis ddof=1, numpy defaults to 0 → uniform ~20 % weight mis-scale) becomes live for the first time. - Its attention is
CrossAttentionover the speaker embedding, notnn.MultiheadAttention; the PhoneExtractor'smha_subsequence()does not transfer. overlap_adddraws a random initial phase, so this component needs §CB1's injectable-noise discipline — and the injector must patchtorch.rand, not justrandn_like.- The lookahead alignment (energy shifted 1 frame, quantised pitch and pitch features 2, all reflect-padded) is the highest-risk silent bug: getting it wrong misaligns content against pitch by 10–20 ms and no input-aligned per-stage check can see it.
Details in docs/music-transcription/BEATRICE_BLUEPRINT.md.
Effort: unestimated until the blueprint read is done, but larger than §CB1 (three networks, custom conv variants). Do NOT start before §CB1's record shapes are agreed.
The remaining open item for full 12B support (a new converter map + backend audio path for the 640-dim unified encoder) is a larger port, scoped but not started.
Status: config/adapter-parity guards DONE (tests/test-tada-params.cpp defaults-audit
now covers library + CLI + c_api; full ~40-adapter sweep clean, cosyvoice3/f5-tts session-config
bugs fixed). Generation-health header + unit tests DONE. Three extensions + one generalisation remain.
TODO (partial — status verified 2026-07-17):
- Per-step talker logits in the diff — DONE for the qwen3_tts exemplar.
tools/reference_backends/qwen3_tts.pycapturestalker_logits+cp_step{0..14}_logits;_iter_capture.pydocuments per-step talker_logits. Generalise to other TTS backends if/when a second consumer needs it. - Wire generation-health checks into backends' live tests — still OPEN. Shared header
src/core/generation_health.h(check_not_empty / duration_plausibility / no_ngram_loop / not_truncated / tts_duration / trailing_silence) +tests/test-generation-health.cppare done; still need per-backend live-test integration (needs models). - Replay-token dual-mode reference — PARTIAL. Noise-replay infra exists
(
_iter_capture.pywritesnoise.binfor C++ to replay); the sampled-token dual-mode (replay Python's argmax picks instead of re-sampling) is the remaining piece. - Generalise the defaults-audit pattern across backends. Per-backend table of (param → upstream default) checked against the params struct, so "knob declared but dead / default diverges from upstream" fails CI everywhere.
#201 follow-up — generate a TADA voice ref from audio+transcript at query time (C-ABI + server DONE gated; roundtrip pending)
Switch-voice, offline --make-ref, --align, and CLI query-time inline cloning
(--tts "…" --voice sample.wav --ref-text "…") all shipped.
C-ABI / session half — DONE (opt-in, feat/tada-201-server-abi). In-memory
make-ref, no temp GGUF:
tada_set_prompt_values()— in-memory counterpart oftada_load_prompt(the latter now reuses it, so file + in-memory paths are identical).tada_make_ref_from_pcm()insrc/tada_tts.{h,cpp}—tada_encoder_encode(validated) →tada_set_prompt_values. Provably equivalent towrite_ref_gguf()+load_prompt()(same tensors), so no new graph math.crispasr_session_set_voice(s, "ref.wav", "<transcript>")decodes to 24 kHz, resolves encoder + language-matched aligner (explicit → next-to-model → cache), and bakes the prompt.crispasr_session_tada_set_makeref_models()sets the GGUF paths; PythonSession.set_voice(path, ref_text)already routes here.- Gated default-OFF:
CRISPASR_TADA_WAV_CLONE=1— without it a.wavvoice keeps the historical-2reject, so default behaviour is byte-identical.
Remaining before flipping the gate on by default:
- Decoded-output roundtrip (HARD RULE #3) — synth reference → set_voice(wav,
text) → synth clone → ASR +
speaker-cosine(clone,ref) > cosine(baseline,ref)via the PythonSessionAPI ontada-1b(+tada-encoder-f16.gguf+tada-aligner-en.gguf). Not run on the dev box (memory-pressured).
Server / adapter half — DONE (opt-in, same gate). The HTTP server uses the
backend adapter (crispasr_backend_tada.cpp), a distinct surface from the
session C-ABI:
TadaBackend::clone_from_wav()— gated (CRISPASR_TADA_WAV_CLONE=1) helper called from bothinit()(first voice) andapply_request_voice()(the server's per-request switch). Resolves the WAV against--voice-dir(bare name →<dir>/<name>.wav+ companion<dir>/<name>.txtfor ref-text, the qwen3-tts convention), resolves encoder + aligner (explicit → next-to-model → cache → auto-download), decodes to 24 kHz, and applies viatada_make_ref_from_pcm— no temp GGUF. Off/failed → the historical reject./v1/audio/speechgained aref_textbody field (→tts_ref_text). The existingconsent_attestationgate already fires for a.wavvoice.- Docs:
docs/server.md(ref_text field + updated tada voice note).
Still OPEN (optional): cache baked ref keyed by (audio hash, transcript) to skip re-running the aligner on repeat requests.
Files: examples/cli/crispasr_backend_tada.cpp, examples/cli/crispasr_server.cpp,
src/tada_tts.{h,cpp}, src/tada_encoder.*. Aligner is language-specific
(tada-aligner-<lang>.gguf) — must match audio language. CLI helper
tada_run_aligner_pipeline is the reference implementation.
Effort: Medium-large. Pipeline is CLI-proven; work is server lifecycle + per-request wiring + caching + the ~1.3 GB memory gate. Lower priority than the shipped switch-voice half. Tracked on #201.
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: GPU-native codec on RADV / chunked-decode design (deferred, conditional on a RADV user needing it)
x86 CPU-only Linux validation of moss-transcribe / higgs-stt / ark-asr passed at
dcc7e47b (all verbatim, beam == greedy, Go LDFLAGS drift clean). Audit tool
tools/check-backend-wiring.py (ccc04a02) ships; 49 canonical PASS required.
Follow-ups (LOW, not blocking):
- Fix handover to
cmake --build build(all targets) beforectest -L unit— VPS run only builtcrispasr/crispasr-diff, ran 2 unit tests. Or have VPS build all. - Install Go toolchain on VPS — DONE 2026-09-02: go1.23.4 at
/mnt/volume1/go-toolchain/go, symlinked as~/.local/bin/go(on PATH). - Optional: promote to a standing post-push Linux smoke (Routine/cron).
- Orphaned ctest-label audit (2026-09-02, follow-up to the dead-
clifind): cross-checked everyLABELSin tests/CMakeLists.txt against every-Lfilter in .github/workflows + ci/. ONE real orphan found and fixed:ci/run.shranctest -L main(upstream whisper.cpp's label — nothing here carries it), so both build.yml CPU jobs tested NOTHING while green; now-L unit, which also gives the unit suite its only Debug-build run. Labels that exist but are deliberately not CI-run (model-gated, local/live only):live,base,small,medium,large,tiny;en,benchmark,integration,espeak;piper;tts— do not "fix" these.
Multilingual + beam spot-checks (LOW, either machine):
- #421 arm64 HiFT SIMDCONV failures — the
-L unitfix's first full-ci dispatch ran 1782 tests on arm64 (previously ZERO) and 4 fail: chatterbox + cosyvoice3 SIMDCONV scalar/SIMD identity (tests/test-*-hift-simdconv.cpp:66/68, instant assert). Possible silent vocoder divergence on arm64 Linux/Android. Unclaimed; details + log pointers on the issue. - moss-transcribe multilingual check — RESOLVED 2026-09-02, premise was
WRONG: the upstream card says "intended for English automatic speech
recognition" (this line had conflated it with the Diarize sibling). Ran the
clips anyway (
samples/paraformer_zh.wav+ a bananamind-de synthesized German sentence —audio_samples/never existed): both come back as rough ENGLISH translations at rc 0, faithful to the model (prompt mirrors the upstream processor, instruction-less). Actioned: moss-transcribe now declaressole_language()="en"on the adapter + the session guard table, so-l zhis an explicit pre-dispatch rejection instead of silently wrong output, and-l autoshort-circuits without an LID download (#227). - higgs/ark
-bs 4noisy-clip beam test — DONE 2026-09-02 on Kaggle T4 (tools/kaggle/beam-noisy-ab/, ${KAGGLE_ACCOUNT}/crispasr-beam-noisy-ab-higgs-ark v2): jfk + additive white noise at 10 dB and 5 dB SNR (seed 42), both backends, greedy vs-bs 4. Result: WER 0.000 in every cell — even 5 dB noise doesn't dent either model on this clip, and beam-4 wins nothing while costing ~3x wall (higgs 2s->7s, ark 4s->12s per clip). Greedy stays the right default. Residual (only if someone cares later): white noise is not accent — a FLEURS accented-speech rerun would need new fixtures; the harness takes any wav list.
Backend-wiring coverage gaps (LOW cleanup; re-list via python tools/check-backend-wiring.py):
- missing reference dumper — RESOLVED by 2026-09-02: all five now exist
in
tools/reference_backends/(fastconformer_ctc.py,wav2vec2.py,m2m100.py,kyutai_stt.py,gemma4_e2b.py). Original caveat kept for the record (m2m100 text-only MT, gemma4-e2b shares gemma path, encoder components diff via host backends) — confirm per-backend before adding, not a blanket gap.
moss-transcribe backend (OpenMOSS-Team/MOSS-Transcribe-preview-2B) shipped
9f3c5ede — q4_k verbatim on jfk.wav, validated vs PyTorch ref via crispasr-diff.
#218 (greedy n-gram loop collapse + 30 s-seam dup) is fixed. Remaining optional work:
- Publish f16 + q8_0 to
cstr/MOSS-Transcribe-preview-2B-GGUF(q4_k + card live; f16/q8_0 held back for WLAN bandwidth). Re-stage from/Volumes/backups/ai/moss-transcribe-preview-2b-{f16,q8_0}.gguf,hf upload-large-folderinto the existing repo. Both already produce the verbatim transcript locally. - GPU validation beyond Metal. Metal (default) + CPU both verbatim; CUDA/Vulkan
untested. LM reuses
core_attn::kv_self_attn(covered by §192/#200 Vulkan F16-GQA guard); check the encoder's windowedflash_attn_ext+ conv front-end on CUDA. - Multilingual eval. Authors report 4.87 % avg WER; only English (jfk) validated here. Model is zh/en — spot-check a Chinese clip.
Note: encoder/adapter run F16 (cos ~0.98 vs f32 ref, byte-exact at layer 0 — pure F16
weight precision, not a bug). f32 encoder path not worth it since decode is verbatim.
Loop-fix opt-out: CRISPASR_MOSS_TRANSCRIBE_NO_LOOPFIX=1.
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: Broader eval to flip windowed-attn default / declare CAP_UNBOUNDED_INPUT (mechanism ready)
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: Document that single-pass blueprint skips quiet leading audio (chunked covers more) in README
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: Investigate mega-asr long-form at 4-bit vs bf16 blueprint (candidate follow-up)
| Priority | Item | Effort | Status |
|---|---|---|---|
| HIGH | #221 Issue #89 hardening + v0.8.8 | Medium | 5 steps: CI regression guard (a), server-path mirror (b), Vulkan sanity (c), q4_k registry/UX (d), release (e). |
| DONE / LOW | §176 Runtime optimization pass | Phased | 18/20 done. 2 OPEN (low-value, measure-first): §176c device-resident KV (Dia measured ~1.2% of decode → DEFERRED; compute-bound decoders make this <2%), §176l Kyutai RVQ (genuinely scalar, but no local model to validate). Do not treat as HIGH. |
| MEDIUM | #52 Qwen3-TTS — perf pass | Medium | talker + code_predictor + codec + ECAPA + codec_encoder done; step-4 perf pass open (~137 ms/frame → real-time). O15 broken on CUDA and default-OFF (61c42bfb) — main perf lever disabled; root cause is ggml_set_rows KV scatter or fixed-Lk causal mask on CUDA (crash on first code_pred call). Baseline O15=OFF: 27.4 ms/frame, WAV OK. |
| MOSTLY DONE | #57 Commercial-friendly TTS expansion | Phased | Phases 1–3 + Turbo + native voice cloning shipped; #83 S3Gen fix landed. VoxCPM2, kugelaudio, gwen-tts, kartoffelbox-turbo, CosyVoice3 all shipped + registry-wired (verified 2026-07-17). Remaining: only the Darwin-TTS-1.7B-Cross / AMAImedia Qwen3-Darwin family unported. → HISTORY §82, upstream-prs/09–11. |
| MEDIUM | #51c MiMo-V2.5-ASR F16 step decode | Small | F16 step-decode validation blocked behind ≥32 GB box. Base runtime + Q4_K shipped → HISTORY §56. |
| MOSTLY DONE | #58 MOSS-Audio-4B-Instruct | Large | Runtime + GGUFs shipped, diff cos≥0.999, Kaggle P100 CUDA PASS. Remaining: flash-attn encoder, sweep transcript-extraction fix. → see HISTORY. |
| MOSTLY DONE | §221 TADA encoder --make-ref |
Medium | C++ encoder runtime + GGUF converters + diff harness shipped; GGUFs at cstr/tada-encoder-GGUF. WIP: cos_mean=0.94 parity (F16 precision), C++ BPE tokenizer for end-to-end --make-ref. → HISTORY §221. |
| LOW | #95 IndexTTS Chinese TN binary alternative | survey only | Python INDEXTTS_TEXT_NORMALIZER hook shipped. Hand-roll (#95a) is next when a user reports a digit/date prompt that breaks; OpenFST vendoring (#95b) only after #95a grows past ~5 cases. |
| IN PROGRESS | #97 More Parakeet variants | Small per-variant | TDT/TDT+CTC + rnnt 0.6b/1.1b DONE. parakeet-unified-en-0.6b surveyed: 24L 600M Unified-FastConformer-RNNT, NOT converter-only — 8× subsampling (vs 4×) + Dynamic Chunked Convolutions are new; ~80% overlap with #81. Offline mode may work through existing converter; realtime-EOU blocked on #81 cache-aware streaming. |
| LOW | #106 TEN-VAD | Small | Feasible VAD backend (C-compatible, 16 kHz / 10-16 ms frames, prebuilt libs + ONNX). License is the gate: Apache 2.0 plus extra no-compete / own-app-only conditions from Agora. |
| MOSTLY DONE | #114 Long-form transcribe chunking-default ladder | Medium | Chunking/streamed-default shipped for all ASR backends + per-chunk AED re-injection + LCS dedup + word-snap. Remaining: EN-FLEURS retokenization artifacts (out of scope). → see HISTORY. |
| MOSTLY DONE | #125 multi-backend bug sweep from montvid | Medium | 12 findings; P1–P6b all DONE. Remaining: P0 Blackwell retest; mimo-asr -np empty-transcript retest. → see HISTORY. |
| LOW | #127 Coverage gaps from 2026-05-26 sweep | Small | (a) omniasr-llm DONE; (c) cohere-asr-ja DONE (repo CKHO/cohere-asr-ja-GGUF) — still needs JA fixture sweep for PERFORMANCE.md table. (b) OPEN: mimo-asr local test doesn't run in CI (4.2 GB Q4_K doesn't fit runner disk). |
Open follow-ups from §79:
- #73 cohere long-form rerun. flash_attn_ext shipped on canary + cohere (
193a736). On JFK (~11 s) canary q8_0/q4_0 −17% under flash (win) but cohere q8_0/q4_0 is +11% (regress); F16 ties both. Before promoting flash as cohere's recommended path, validate on a multi-minute clip — if the crossover is workload-dependent, docs must recommend cast-on-read for short audio, flash for long. Until then PERFORMANCE.md notes flash as available-but-regresses-on-JFK for cohere. - encoder-decoder #69a (canary, cohere, kyutai-stt). Cross-attention layout has no
<prefix><N>.*block-tagged tensors; needs bespoke per-backend predicates. Own design problem. - #06 FA per-head mask (A1000 perf step). Removes 72 CPU splits/chunk (per-head additive mask in
fattn.cu:423+ four kernel variants). ~300-500 LOC acrossfattn.cu/fattn-common.cuh/fattn-mma-f16.cuh. Expected ~10-15% wallclock on top of postsiglu. Follow the WDDM-warm bench protocol (LEARNINGS) for the new baseline before starting.
Batched-CFG B=2 T3 decode shipped behind CRISPASR_CHATTERBOX_T3_CFG_B2=1
(default OFF); greedy-token bit-identical to legacy on CPU (all quants), GPU+F16,
GPU+quant. Path works — what's left is proving/delivering the speedup.
Files: src/chatterbox.cpp (build_graph_t3_kv_b2, run_t3_kv_b2,
ensure_t3_b2_f16_weights, decode loop ~§214). See HISTORY + PERFORMANCE §214.
TODO:
- Quiet-machine A/B + default-flip decision (HIGH, ~1 h). CPU floor ~34 %
was measured on a contended M1 — unreliable. Re-measure on a quiet host
(alternating order, min-of-N, token-parity gate
CRISPASR_CHATTERBOX_TEMP=0). Flip default ON for the CFG path only if the win holds (keep env + legacy path forever for bisection). - Generalize B=2 to the other CFG backends — see §215.
Closed/won't-pursue (do not reopen): cached/bucketed B=2 step graph (build+alloc
<1 % on both backends, and risks the §186 Metal buffer is nil crash); GPU B=2
as a speed win (GPU ~4× slower/step than CPU — its value is enabling T3-on-GPU +
GPU+quant F16-dequant, not beating the CPU default which stays production).
Full-tree audit (§214, PERFORMANCE.md): only s3gen CFM + chatterbox T3 batch cond+uncond into one B=2 forward; every other CFG backend runs two sequential B=1 passes. Where the per-step forward is GPU-run, dispatch-bound, AND the dominant cost, apply the chatterbox-T3 B=2 pattern: batch over ne[2]=2 for heavy GEMMs, split per-batch attention/KV-cache, tag GGML_PREC_F32, gate behind an env (default OFF), parity-gate vs the sequential path. On GPU + quantized weights, dequant the batched weights q*→F16 GPU-resident (s3gen dequant_cfm_f16 / T3 ensure_t3_b2_f16_weights trick — Metal's batched ne[2]=2 quant mat-vec misses the PREC_F32 exact-dot kernel and degenerates).
General caveat (voxtral §93 lesson): batched-CFG is a MODEST (~1.2–1.3×), GPU-dispatch-bound win. Before each port confirm the stage is GPU-run AND dispatch-bound AND dominant — measure, don't assume the T3 −42% transfers.
Already resolved (see HISTORY): §215a dia-encoder B=2 done; §215b tada — MEASURED NON-GOAL, do NOT port (talker only ~24% of loop, asymmetric graph paths); §215b bucket-floor follow-up SHIPPED backend-conditional default (tada_default_bucket_min(): Metal/CPU→64, CUDA/ROCm/Vulkan/WebGPU→512; CRISPASR_TADA_BUCKET_MIN overrides).
Open candidates, prioritized (high step count × dispatch-bound first):
- §215a dia decoder (HIGH).
src/dia_tts.cpp run_dia_synth— encoder already B=2, but the decoder AR loop still runs cond/uncond as two passes / two KV caches (run_dia_decode_step). Mirror chatterbox T3: B=2 decode-step graph, split per-batch KV write/read, F16-dequant on GPU+quant. Largest payoff — long AR loop, CFG every step. - §215c zonos (MED).
src/zonos_tts.cppkeeps two KV caches (kv_k/kv_k_uncond), decodes sequentially (~L1740–1842). Batch AR backbone B=2, split dual-KV attention; keep dual-KV CFG + optional random speaker embed independent per batch. - §215d voxcpm2 (MED).
src/voxcpm2_tts.cppLocDiT runslocdit_calltwice per ODE step (condmu+ zero-mu, ~L2752–2778). Diffusion (fewer steps) so lower payoff, but each LocDiT forward is heavy. Keep cfg-zero-star blend per batch. - §215e f5 (MED, risky).
src/f5_tts.cpprunsdit_forwardtwice per CFG step (v_cond+v_uncond, ~L1563–1566). Same B=2-DiT shape as voxcpm2. §176h: a standalone B=2 DiT graph is correct but a runtime B=2 corrupted batch-1 (F5-runtime-specific, [[project_ggml_batched_fused_graph_alloc_bug]]) — run the parity gate especially carefully; may not be worth it. - §215f cosyvoice3 (LOW — likely WON'T).
src/cosyvoice3_tts.cppexplicitly declined batching (~L3027–3030): 22-block diffusion twice/step, per-call overhead small vs forward. Only revisit if a profile shows dispatch overhead matters; else document as deliberate non-goal. - §215g kugelaudio (BLOCKED).
src/kugelaudio.cppCFG is a TODO (cfg_scale read but unused, negative path unimplemented). Implement sequential CFG first (correctness), then consider B=2.
Shared infra: factor the F16-dequant-of-matmul-weights helper (duplicated in ensure_t3_b2_f16_weights + s3gen dequant_cfm_f16) into a core_* helper only when a third backend needs it.
§210 follow-up — shape-stable bucketed decode for remaining LLM/AR backends (CUDA-graph capture) (OPEN, CONDITIONAL)
Status: template landed in granite-speech (PR #207) — fixed-Lk KV bucket written via
ggml_set_rows, in-graph ggml_argmax, fused F16 embed. Unlocks CUDA-graph capture
(~9–13× decode on Ampere+, engages automatically in ggml-cuda for shape+pointer-stable graphs)
and Metal gallocr allocate-once. Survey done (LEARNINGS §210). Done, no work: granite_speech,
mimo_asr, dots_tts. irodori-tts DiT persistent graph implemented + byte-identical parity (default
OFF, gated CRISPASR_IRODORI_PERSIST_GRAPH).
- CUDA-graph win is gated to Ampere+ (sm_80+):
ggml-cuda.cu:4329. Project's usual test GPUs T4 (sm_75) / P100 (sm_60) are gated OUT — only RTX 5090 / A100-class benefit. - On Metal there is no throughput win (granite decode host-encode 1.8% / GPU 98%, GPU-bound Q4_K GEMVs); shape-stable rewrite buys only memory-pressure robustness, not speed.
- Each port is a MANUAL graph rewrite + byte-identical diff-harness validation — NOT delegable to agents (runtime graph code). Budget one focused session per backend.
Port a backend ONLY when actually deployed on Ampere+ CUDA, and measure first
(CRISPASR_METAL_PROFILE for host/GPU split; confirm no "disabling CUDA graphs" GGML_LOG_DEBUG line
on an Ampere+ GPU). Smaller decoders may have a larger host-encode fraction than granite's 1.8% — measure, don't assume.
src/granite_speech.cpp:granite_build_argmax_decode(~L2474),granite_dec_use_gallocr(~L2362; gallocr on Metal/CPU, sched on CUDA so capture fires).src/mimo_asr.cpp: cachedstep_t1_gf+step_t1_fixed_kv_len(L211), set_rows scatter at runtimekv_indices(L958, L1026), skip-plan reuse (L1507–1522) — good set_rows reference.
ASR-LLM (prioritized by likely server deployment):
- voxtral —
src/voxtral.cpp,LkL1019; rebuild + reset/alloc L1159/L1192/L1244. - qwen3_asr —
src/qwen3_asr.cpp,LkL1210. - voxtral4b —
src/voxtral4b.cpp,LkL1411 (decode viacore_greedy_decode). - gemma4_e2b —
src/gemma4_e2b.cpp,LkL995 (core_greedy_decode::run_with_probs_cb~L1631). - glm_asr —
src/glm_asr.cpp,LkL1322. - higgs_stt —
src/higgs_stt.cpp,LkL1061 (Qwen3-1.7B decoder). - ark_asr —
src/ark_asr.cpp,ark_build_decoder_graph(L643) /ark_run_decoder(L729),LkL674. - lfm2_audio —
src/lfm2_audio.cpp,LkL933; already on gallocr (L875) but growing-shape, not yet capturable.
TTS AR decoders (LOWER priority — heavier per-step compute = more GPU-bound, payoff diluted):
csm_tts, indextts, bark_tts, moss_audio, mini_omni2 naive growing-shape. qwen3_tts,
chatterbox (T3), tada_tts, parler_tts, vibevoice already have per-backend perf work — audit individually.
- Allocate KV at
kv_max_ctxonce; pick fixedbucket_len(≤ cache cap). - Rewrite step graph to fixed
[0, bucket_len)KV view; write new token K/V viaggml_set_rowsat runtime indexn_past. Mask(bucket_len, 1), set host-side each step. Topology byte-identical across steps. - Move argmax in-graph (
ggml_argmax); keeplogitsas graph output for callers. - Make embed capturable: if
token_embdk-quant, in-graphGET_ROWShost-syncs and disables capture — fuse a F16 embed or pass a pre-computed F32 embed input (granitefused_embed). - Cache the cgraph; gate gallocr-vs-sched like
granite_dec_use_gallocr(sched on CUDA/HIP so capture engages; force sched if a CPU layer split exists). Env opt-out. - Bound
n_past < bucket_len(granite OOB guardc5035969).
- Byte-identical transcript vs legacy path on jfk + fleurs_60s — no merge without it.
- Measure before/after:
CRISPASR_METAL_PROFILE+ per-step compute-µs accumulator likeCRISPASR_GRANITE_DEC_PROFILE. On real Ampere+ confirm capture engages, A/B decode RTFx. M1 wall time is noise — gate on the instrumented per-step quantity. - irodori DiT: flip its default ON (or gate
cc>=800) only once a real Ampere A/B shows a win.
Effort: ~1 focused session per backend. Do highest-deployment ASR backend first; stop if its measured Ampere+ CUDA A/B doesn't justify the next.
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: Speex / WavPack decoder support (LOW, only if a corpus needs them)
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: Runtime asrTranscribe-equivalent smoke test in node/browser (LOW nicety)
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: language-instruction prompt template helper (revisit if a 3rd consumer appears)
All other legs fixed (server/CLI chunking §176 prefix guard, KV stride leak at
71f0639, quant recipe b36248c1+5c8add40 — three HF repos regenerated). Reporter
retested regenerated q8_0 (sha256-confirmed current HF file): still broken. Signature:
q8_0 breaks, f16 clean, same RX 9070 XT (RADV GFX1201, int dot: 1); Metal + MoltenVK
clean on the exact q8_0. Suspicion: RADV GFX1201 shader miscompile on the quantized-
matmul paths (int-dot MMQ/MMVQ, KHR_coopmat dequant) — shader logic proven fine
(forced int-dot + MMVQ clean on M1). Vendored ggml base 2026-05-05, no matching
upstream fix found.
TO DO:
- Reporter runs the knob matrix one-at-a-time on a broken sample (env-only, no
rebuild):
GGML_VK_DISABLE_INTEGER_DOT_PRODUCT=1,GGML_VK_DISABLE_MMVQ=1,GGML_VK_DISABLE_COOPMAT=1,GGML_VK_DISABLE_F16=1, anchorCRISPASR_N_GPU_LAYERS=0. Also Mesa upgrade / AMDVLK cross-check. - Once one knob is confirmed → device-targeted safe default (RADV GFX12xx) in
vendored
ggml-vulkan+ upstream report to ggml-org/llama.cpp with minimal repro. - Note for thread: reporter's old
vibevoice-1.5b-bf16.ggufpredates--include-decoder; current f16 oncstr/vibevoice-1.5b-GGUFhas the decoder tensors.
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: moonshine-streaming-tiny/small/medium need new streaming runtime
Runtime (src/parakeet.cpp) dispatches TDT/CTC/RNNT via GGUF flags; converter
(models/convert-parakeet-to-gguf.py) reads hparams from model_config.yaml +
cross-checks tensor shapes. Most FastConformer-encoder + TDT/CTC/RNNT checkpoints
are converter-only, no new C++. Shipped: tdt-0.6b-v2, tdt-1.1b, tdt_ctc-110m,
tdt_ctc-1.1b, rnnt-0.6b, rnnt-1.1b (all on HF, registry-wired). See HISTORY.
Still open:
nvidia/parakeet-unified-en-0.6b— direct-zip extraction + GGUF conversion already WORK (v4/v5 Kaggle kernels; 1181 MB F16, same arch as standard parakeet: d_model=1024, n_layers=24, vocab=1025, pred=640, joint=640). CrispASR SIGABRTs because runtime assumes 4× subsampling but this model uses 8× (3 strided convs, not 2). Fix: makeparakeet_build_pre_encodeinsrc/parakeet.cpp(orcore/fastconformer.h) handlesubsampling_factor=8— 3 Conv2d layers with strides [1,2] instead of 2. Tensor shapes already in GGUF; graph builder just needs the extra conv layer.nvidia/parakeet_realtime_eou_120m-v1— streaming + end-of-utterance head. Needs cache-aware FastConformer streaming (cf. #81 Nemotron) + an EOU head. Not converter-only.
Won't do (unless a user asks): parakeet-ctc-0.6b-Vietnamese — already
runtime-supported (CTC), known gap not active work.
See also #98 Hotwords (orthogonal; lights up biasing on every Parakeet variant once CTC-WS trie lands).
User-supplied vocabulary the ASR prefers when in doubt (names, jargon, product/place names). Helps only the biased subset, but lift there is large. Phased so each phase covers a family of backends. ~1 week total covering 9 of 14 backends. (Full upstream-support survey → HISTORY.)
Phase A — generic CTC-WS phrase-boost trie (2–3 days). Covers parakeet-ctc / -tdt / fastconformer-ctc / omniasr in one shot (model-agnostic on the logit stream).
- New shared helper
src/core/asr_context_bias.{h,cpp}— Aho-Corasick trie over piece-id sequences, configurable per-phrase boost; emits a per-frame log-prob bias vector the CTC/TDT decoder shallow-fuses into argmax/beam scoring. Pure CPU, no ggml graph. - Wire-in:
parakeet_ctc_decode+parakeet_tdt_decode(src/parakeet.cpp:999+,parakeet.cpp:1670dispatch). Phrase tokenisation via the backend's existing SentencePiece model so users pass human-readable strings. - CLI:
--hotwords "Acme Corp,Sandra Berenz,GPU-PB"and/or--hotwords-file <path>(one phrase/line, optional^Nboost suffix); envCRISPASR_HOTWORDS=...for the OpenAI-server path. ~250–400 LOC incl. beam-rescoring. - Reference to mirror: NeMo CTC-WS notebook + TurboBias
BoostingTreeC++ (arxiv.org/html/2508.07014v1).
Phase B — --hotwords → LLM prompt-prefix helper (1 day). Covers funasr, granite-plus,
voxtral, qwen3-asr (each already has a system-prompt path).
- Tiny
src/core/helper renders a hotword list into each backend's prompt template (graniteKeywords: …; funasr hotword token block; qwen3 free-text context). One template registry, one call site per backend. - Wire-in:
funasr_transcribe_ex,granite_nle_transcribe, voxtral, qwen3-asr. ~150 LOC + per-backend template strings.
Phase C — parakeet TDT joint-net boost (Transducer-native) (1–2 days). Per-step bias on joint-net output when the partial hyp matches a trie prefix (mirrors NeMo MBS hotwords).
- DECISION-GATE: defer until Phase A is shipped + benchmarked; only do it if Phase A on TDT undershoots NeMo's reference numbers.
Out of scope: whisper initial_prompt (already upstream via --initial-prompt), MiMo
PromptASR (no upstream flag — park), cohere/moonshine/kyutai-stt/glm-asr (no upstream hook).
Validation: tests/test_hotwords.py — synthetic clip with a rare name (e.g. "Berenz")
through each Phase-A backend with/without --hotwords Berenz; assert unbiased misspells,
biased nails it. Phase B: assert prompt-prefix matches upstream Python byte-for-byte.
Port ships with ggml_flash_attn_ext on encoder+adaptor (FUNASR_NO_FA=1 to opt
out), fused QKV (DONE), and single-token embed fast path (CRISPASR_FUNASR_EMBED_FAST,
default ON, DONE §180). None affect correctness — pure throughput. Remaining:
- Per-step LLM decode graph cache — DONE 2026-09-02, but NOT as recorded
here. The literal design (
fixed_kv_len = kv_max_ctx, one graph) was already scaffolded default-OFF in the file with an M1 3× regression note, and re-measured on x86 as +69% decode CPU — one wasted KV key costs ~0.28 ms/tok through the GQA repeat+cont, so a fixed-Lk graph loses far more to KV bandwidth than it saves in graph prep (~2.1 ms/tok). Shipped instead: bucketed Lk, width 16 (CRISPASR_FUNASR_STEP_BUCKET, measured optimum; break-even w≈15), FIFO of 4 live graphs, padded slots masked to -inf (flash-attn skips fully masked positions → bit-identical by construction), default ON withCRISPASR_FUNASR_STEP_CACHE=0opt-out. Byte-identical 18/18 (3 clips × f16/q4_k × pre-change/off/on, re-verified independently), quant-bcast audit clean (0 hits, q4_k+q8_0, both arms). Net on x86: neutral within ±8% noise (decode is weight-bandwidth-bound; step prep was only 2–3.8 ms/tok here, not the ~30 ms the M1 note implied) while removing 93% of per-step graph prep — the win lives on platforms where graph construction is expensive (M1). qwen3_asr adoption: only with the bucketed variant, never fixed-Lk. - Encoder graph cache — exact-T_lfr half DONE 2026-09-02 (default ON,
CRISPASR_FUNASR_ENC_CACHE=0opt-out): repeat-length calls skip the 7.5–23.9 ms graph build (cache hit = 0 ms; sched_alloc 2.4–6.2 ms still paid). The padded-bucket half is dropped as designed: the SANM FSMN branch is a width-11 depthwise conv over TIME (core/sanm.h:128), so a padded frame leaks into the last 5 real frames of all 70 blocks and zero-padding doesn't save it (LayerNorm of a zero row = bias ≠ 0) — a bit-identical version needs a time mask on V insidecore_sanm— and the economics are upside down anyway (build is ~10–24 ms against a 2.3–6.1 s encoder; padding T_lfr 183→256 would spend seconds to save milliseconds). - Two-pass: CTC fast pass → Fun-ASR-Nano LLM rescore. Upstream checkpoint has
0 CTC tensors (LLM-style by choice); the only public trained CTC head is
csukuangfj/funasr-nano-with-ctc(Apache-2.0, encoder+adaptor+CTC, no LLM, frozen encoder = upstream). Two patterns: (a) csukuangfj CTC head + upstream encoder/adaptor/LLM — single shared encoder forward, no vocab remap (cleaner, single-author trust); (b) fallback SenseVoice-Small fast pass (gold trust, extra encoder + vocab translate).- Phase A (measure, no code): tensor-list csukuangfj
model.ptto confirm his encoder == upstream byte-identical; run his CTC head + upstream encoder+adaptor on a zh+en ground-truth set, measure CER/WER vs pure-LLM path. Need within ~3–5% rel for rescore net win. Write to LEARNINGS.md; proceed to B if in bounds, else use fallback (b). - Phase B (impl): grow
models/convert-funasr-to-gguf.pyto optionally pick upctc_decoder.*from a with-ctc checkpoint →funasr-nano-with-ctc-q4_k.gguf, auto-download fromcstr/funasr-nano-with-ctc-GGUF(mirror csukuangfj + attribution). Opt-inCRISPASR_FUNASR_TWOPASS=1(requires companion GGUF, else single-pass) +--asr-rescoreCLI flag. One encoder forward forks into CTC + LLM heads: greedy CTC → per-frame probs; skip LLM if avg per-frame conf >0.95, else CTC top-K as LLM decode-prefix candidates. Expected 2–4× on high-conf clips, neutral on hard audio.
- Phase A (measure, no code): tensor-list csukuangfj
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: granite_speech embed fast path (VPS bench blocked) + deepen bench stages / encoder graph cache
Status: the 2026-06-20 full Kaggle GPU sweep (tools/kaggle-benchmark-all-backends.py, streamed to
cstr/crispasr-kaggle-progress/full-backend-sweep/) ran 59 backends; 10 apparent failures reduced to
a handful of genuine CUDA-path bugs. Most resolved (vibevoice/lfm2-audio §206/kugelaudio §209/
fastpitch+speecht5 §204/chatterbox §205 — all in HISTORY). Remaining open:
TODO (open):
- f5-tts — runs once given a reference voice but TIMEOUT at 120 s in re-test.
Timeout half DONE (registry already carries
timeout_s=600for f5-tts,tools/test-all-backends.py— ≥240 satisfied); the Kaggle re-run to settle pass-vs-stuck is still pending; passes on M1 Metal locally. - orpheus (TTS) — fixed §215 (Metal + CPU bucket both ASR-roundtrip verbatim on M1), stays
opt-in
CRISPASR_ORPHEUS_BUCKET=1(~30% slower on M1 unified memory, may win on CUDA). CUDA cross-check still pending (Kaggle${KAGGLE_ACCOUNT}/crispasr-orpheus-talker-cudaend-to-endorpheus_synthesize). - chatterbox (TTS) — 0-byte (~14 s) with
--voice <wav> --i-have-rights; the #83 S3Gen GPU fix was Metal-validated — re-check the CUDA S3Gen path. - cosyvoice3 (TTS) — dies in 0.1 s even with a reference voice; passes on M1 Metal.
Flow-matching + HiFT — cheapest to bisect. §205's mixed-radix FFT fix covers CosyVoice3's
n_fft=400mel (same heap overflow as chatterbox) — likely resolved, needs Kaggle CUDA re-test.
Method: per-backend JSONs in the dataset have timing context; reproduce on a CUDA worker
(Kaggle T4/P100 or A1000) with CRISPASR_VERBOSE=1 + CRISPASR_<BACKEND>_DEBUG=1. Several pass on
M1 Metal → CUDA-path-specific; cross-check Metal first. Small models also fit the 8 GB CPU-only VPS
where the diff harness can drive the fix if the bug reproduces on CPU (how §204/§206 were fixed).
Base runtime + Q4_K + fused-QKV shipped (HISTORY §56/§64); 51a mmap loader (§62) and 51b step-decode KV reuse (§60) DONE. Only F16 step decode remains — no code change needed (runtime is dtype-agnostic; mmap loader already wired). Blocked behind ≥32 GB RAM: F16 working set (~16 GB) thrashes on this 16 GB box.
- On a 32+ GB box, validate JFK transcript byte-equality + decode speedup:
CRISPASR_GGUF_MMAP=1 ./build-ninja-compile/bin/crispasr --backend mimo-asr \ -m /path/to/mimo-asr-f16.gguf --codec-model /path/to/mimo-tokenizer-q4_k.gguf \ -f samples/jfk.wav - Gate: if F16 prefill hits ≥1× realtime as predicted, ship F16 as the
recommended quant on
cstr/mimo-asr-GGUFand demote Q4_K to memory-tight fallback. Until then both shipped, Q4_K default.
Effort: 0 LOC (validation only).
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: perf pass — clean quiet-machine FUSED_QKV bench (F16/Q4_K), fusing 15 cp steps into one graph
Variants 4.1 / 4.1-plus / 4.1-nar shipped bit-exact on JFK → HISTORY
§61. Remaining: speaker labels + word-level timestamps for the plus
variant via chat_template (~50 LOC, template-only).
In-process libespeak-ng phonemization (behind CMake CRISPASR_WITH_ESPEAK_NG
AUTO/ON/OFF, CRISPASR_HAVE_ESPEAK_NG=1) with popen fallback + LRU cache SHIPPED.
German voice cascade (Option 1/2a/2b), phonemizer diff harness, and cache-clear
ABI all DONE (see HISTORY). Remaining open: Mandarin tone numbers, Japanese kanji,
optional German stage-2 fine-tune.
-
German stage-2 fine-tune (optional, out of scope of this item). Native German path already ships (dida-80b backbone + kikiri voicepacks, auto-routes when both
kokoro-82m-f16.ggufandkokoro-de-hui-base-f16.ggufare in the same dir). For deployable single-speaker production quality, run Stage-2 fine-tune on one HUI speaker (~half-day A40). Track separately if needed. -
Mandarin tone numbers. espeak-ng emits digit-suffixed tones (
ni2χˈɑu2) not in the 178-symbol kokoro IPA vocab → dropped at tokenization, losing tone. Investigate--ipa=2(no tone numbers) + separate tone embedding, or switch Mandarin G2P (e.g.pypinyin). Symptom is auto-captured bytools/check_kokoro_phonemizer_parity.py. Effort: ~an afternoon. -
Japanese kanji. espeak-ng falls back to English for kanji (日本語 → "Chinese letter") inserting non-IPA
(en)…(ja)markers. Pre-process kanji→kana with a Japanese frontend before espeak. MIT-clean approach: MeCab (BSD-3) + unidic-lite (MIT) morphological analysis → reading extraction → feed kana to espeakja. NO kakasi/pykakasi (GPL-3, viral). Impl: libmecab C API via dlopen (like espeak) OR a shipped kanji→kana flat-file dict (simpler, less accurate on rare/compound; generate offline with fugashi/cutlet, both MIT). Effort: ~an afternoon.
Files: src/kokoro.{h,cpp}, examples/cli/crispasr_backend_kokoro.cpp.
May 2026 sweep of high-traffic HF TTS models. Filter: permissive license + reusable architecture + reasonable effort. Sequenced so each phase unlocks a family of finetunes (e.g. Phase 3 Chatterbox stack also unlocks Phase 5's CFM solver).
License triage that drives ordering (candidates for later phases):
| ✅ Permissive (commercial OK) | ❌ Non-commercial — defer | |
|---|---|---|
| Qwen3-TTS-{Base,CustomVoice} (Apache 2.0) | Orpheus-3B family + Kartoffel_Orpheus (llama3.2) | SebastianBodza/Kartoffelbox-v0.1 (CC-BY-NC-ND) |
| ResembleAI/chatterbox base (MIT) | HumeAI/tada-3b-ml (llama3.2) | marduk-ra/F5-TTS-German (CC-BY-NC) |
| SebastianBodza/Kartoffelbox_Turbo (CC-BY-4.0, gated) | mlx-community/fish-audio-s2-pro (Fish-Audio Research) | |
| oddadmix/lahgtna-chatterbox-v0/v1 (MIT) | amphion/Vevo1.5 (CC-BY-NC-ND) | |
| openbmb/VoxCPM2 (Apache 2.0) | mlx-community/Voxtral-4B-TTS-2603 (CC-BY-NC; upstream Mistral Apache OK) | |
| FINAL-Bench/Darwin-TTS-1.7B-Cross (Apache 2.0) | ||
| AMAImedia Qwen3-1.7B-TTS-Cross-Darwin AWQ (Apache 2.0) | ||
| g-group-ai-lab/gwen-tts-0.6B (MIT) | ||
| kugelaudio/kugelaudio-0-open (MIT) |
TO DO:
Resolve license gap before depending on CosyVoice 3— MOOT:src/cosyvoice3_tts.cpp- registry entry already shipped. Likewise VoxCPM2, kugelaudio, gwen-tts, and kartoffelbox-turbo (German) are all ported + registry-wired (verified 2026-07-17).
- Remaining unported from the permissive column: Darwin-TTS-1.7B-Cross and the
AMAImedia Qwen3-1.7B-TTS-Cross-Darwin family (no
src/*darwin*, no registry entry). - Phase 2+ ports otherwise not detailed here — scope from the permissive column when picking the next family.
Phase 1 — DONE (see HISTORY.md + git log).
| Model | License | Approach | Priority |
|---|---|---|---|
| Wav2Vec2 Conformer | Apache-2.0 | Conformer attention variant | Medium |
| Qwen2-Audio 7B | Apache-2.0 | Whisper encoder + Qwen2 LLM | Medium |
| OmniASR larger (1B/3B/7B) | Apache-2.0 | Same converter, bigger models | Medium |
| NeMo Canary-Qwen-2.5b | Apache-2.0 | FastConformer + Qwen2.5 decoder | Medium |
| Paza / Phi-4 | MIT | 14B multimodal, defer to llama.cpp | Low |
| XiaomiMiMo/MiMo-V2.5-ASR | TBD (check) | LLM-style multimodal speech (Qwen3-ASR pattern) | Medium — user-requested #35 |
| google/gemma-4-E2B | Gemma terms | Conformer + Gemma 4 decoder | Medium — user-requested #35 |
From llama.cpp (MIT): Ultravox (Whisper enc + Llama 3.2), Gemma 4 Audio (Conformer, chunked attn, streaming), LFM2-Audio (Conformer variant, position embeddings).
- CT-Transformer (FunASR) Apache-2.0 — Medium. SANM 3-layer (vocab 272727),
zh+en, FunASR/RapidPunc production default.
modelscope/punc_ct-transformer_zh-cn-common-vadrealtime-vocab272727-pytorch. SANM primitives already in CrispASR (src/core/sanm.h). New aliasct-punc; VAD-realtime variant emits per-segment punc for streaming. - bert-restore-punctuation (MIT, en) — Low.
- xashru/punctuation (Apache-2.0, XLM-R+BiLSTM-CRF, 40+ langs) — Low. (FireRedPunc, fullstop, punctuate-all, PCS, all truecasers — DONE, see HISTORY.)
| # | Optimization | Applies to | Expected gain | Status |
|---|---|---|---|---|
| O2 | Fused QKV pre-merge | LLM decoders | ~10-15% attn (GPU) | API ready in core/attention.h; CPU gain <1%, defer to GPU |
| O5 | Pipelined mel+encode | LLM backends, CPU | ~15-20% | TODO |
| O6 | Batched encoder (GPU) | All + GPU | 3-5x | TODO |
| O7 | Speculative decoding | LLM backends | 2-4x decode | TODO |
| O4 | Beam search for LLMs | Audio-LLM backends | Quality | DONE except mimo-asr, blocked on #115 |
Guidance: only move LARGE, REUSED matmuls onto ggml/GPU; persistent subgraphs per decode step > one-off graphs; never dequant (native Q4_K matmul 9.3× faster than F32 OpenMP). Candidate from CrispEmbed: SentencePiece Viterbi DP optimal tokenizer.
-
.m4a,.mp4,.webmcrash with upstream ffmpeg integration — needs fix or robust fallback. -
.aiff,.wma, raw PCM not supported without pre-conversion. Consider bundling a lightweight M4A/AAC decoder or improving the ffmpeg path.
- OmniASR-LLM beam search — beam=2+ with N hypothesis KV caches.
Status: Dart wrapper is published on pub.dev (crispasr 0.8.11, manual publish — see §66). Rust + Python wrappers have publishable metadata + passing dry-runs but are not yet on crates.io / PyPI (both 404, verified 2026-07-17) — blocked on the one-time registry creds bootstrap in §66. All are thin FFI/ctypes shims over the C ABI in src/crispasr_c_api.cpp — they do NOT bundle the native lib (user must have libcrispasr.{so,dylib,dll} installed).
- PyPI — at https://pypi.org/manage/account/publishing/ add a pending publisher: owner
CrispStrobe, repoCrispASR, workflowrelease-wrappers.yml, environmentpypi. Then push anyv*tag. (OIDC trusted-publishing, no token.) - crates.io — generate a token at https://crates.io/me, add as
CARGO_REGISTRY_TOKENrepo secret. Publish order:crispasr-systhencrispasr. - pub.dev — at https://pub.dev/packages/crispasr/admin (after first manual publish/claim) enable automated publishing, tag pattern
v{{version}}. Or first-publish locally viadart pub publishwith owner creds.
Update _find_lib() in python/crispasr/_binding.py to probe in order: (1) $CRISPASR_LIB_PATH; (2) sys.prefix/lib/; (3) Homebrew/Linux paths (/opt/homebrew/lib, /usr/local/lib, /usr/lib); (4) existing repo-relative fallbacks. If none found, raise RuntimeError linking to install docs.
.github/workflows/release-wrappers.yml, tag-triggered (v* only, not every commit), runs in parallel: python -m build && twine upload (PyPI OIDC); cargo publish -p crispasr-sys && cargo publish -p crispasr (crates.io); dart pub publish --force (pub.dev OIDC). Version bumps stay manual — bump pyproject.toml/Cargo.toml/pubspec.yaml together in the tag commit.
Effort: Low per wrapper.
After the pure-Python release is out and stable, add a cibuildwheel pipeline (manylinux2014 + macOS arm64/x64 + Windows) bundling libcrispasr.* via auditwheel/delocate/delvewheel. Same optional path for Rust (crispasr-sys vendoring native build like tch-rs/onnxruntime-sys). Defer.
Status: Session TTS API (incl. qwen3-tts variant routing) is wrapped across all 7 bindings (commit 65e0a61 + Dart follow-up). The non-Session ABI (~80 of the 136+ crispasr_* exports in src/crispasr_c_api.cpp) is still C-ABI-only or partially wrapped on most bindings. Rust + Python are canonical full-coverage; the diarize surface (segment-level + #107 P6 embedder/clustering/cache) also landed in Dart/Flutter + Go.
Not now. Open a capability×binding only when a concrete consumer asks (e.g. "Java VAD", "Go streaming"). Reference commits for the pattern: 4f476c3 (TTS surface sweep), 65e0a61 (variant detect).
Each = ~3-12 exports + an idiomatic result type per binding:
- Forced alignment —
crispasr_align_words,align_words_abi,align_result_*. - Diarization (segment) —
crispasr_diarize_segments[_abi]. Missing in Java, Ruby, JS. - Diarization (embedder+clustering) —
crispasr_speaker_embedder_*_abi,crispasr_speaker_cluster_abi,crispasr_pyannote_cache_*_abi(#107 P6). Missing in Java, Ruby, JS. - Language ID —
crispasr_detect_language[_pcm],crispasr_lid_free_cache. - VAD —
crispasr_vad_segments,crispasr_compute_vad_slices,crispasr_stitch_vad_slices,crispasr_vad_remap_timestamp,crispasr_vad_free. - Streaming —
crispasr_stream_open/feed/get_text/flush/close,crispasr_stream_run_decode. - Punctuation —
crispasr_punc_init/process/free/free_text. - Model registry —
crispasr_registry_lookup[_abi],registry_lookup_by_filename[_abi],crispasr_detect_backend_from_gguf. - Cache —
crispasr_cache_dir_abi,crispasr_cache_ensure_file_abi.
- Java (
bindings/java/) — JNI exposes onlycrispasr_session_*+*speaker_name*. Add JNI wrappers forcrispasr_diarize_segments_abi+ the 9crispasr_speaker_*_abi/crispasr_pyannote_cache_*_abiexports + idiomatic helper class. ~250 LOC. - Ruby (
bindings/ruby/) — onlySession.transcribe. Needs Ruby FFI diarize bindings. ~200 LOC. - JS/WASM (
bindings/javascript/) — no speaker surface; depends on the WASM build linking pyannote-seg/titanet/indextts_voc. Start withcrispasr_diarize_segments_abi(no model deps beyond existing wasm whisper); defer embedder primitives.
Effort: ~150-300 LOC per binding. Suggested ordering once a consumer asks: (1) Streaming (Go/Java), (2) VAD+alignment (Dart/mobile), (3) Diarization+LID+punc, (4) Registry+cache.
Optionally relocate crispasr/ + crispasr-sys/ (repo root) under bindings/rust/ to match C-family bindings. Do NOT rename the crates (names are correct/idiomatic). Move both dirs together (relative path = "../crispasr-sys" dep, no workspace). Consumer-safe (crates.io resolves by name+version). Before moving, audit: (a) downstream repos using git+path dep on the subdir (CrispEmbed/CrisperWeaver), (b) internal CI/scripts//build_go refs, (c) docs path refs. One deliberate commit. Not worth churn unless the root-dir ambiguity bothers.
Parent #65 (session-API word-confidence parity) shipped → HISTORY §65
(main batch + vibevoice / moonshine-streaming + gemma4-e2b token-prob
API + Go/Java/Ruby parity in 5534588 + d963e3a). Only residual:
JS/emscripten word accessors — leaving until a JS consumer asks (the
current JS binding is TTS-focused).
Status: PARTIAL (verified live 2026-07-17). Auto-trigger silenced —
tags: ['v*'] push trigger on release-wrappers.yml is COMMENTED OUT (failed on
every release since v0.5.0, confirmed v0.5.4 gh run view 25248028443). Workflow
stays on workflow_dispatch only.
Live registry state:
- pub.dev
crispasr— DONE, latest 0.8.11 (manually publisheddart pub publish --force, see~/code/pupdev.mdhandover 2026-07-15). No bootstrap needed for the package to exist; only the pub.dev-admin "automated publishing" toggle remains if tag-triggered republish is wanted. - crates.io
crispasr+crispasr-sys— DONE, both 0.8.23 (manually published 2026-07-27 withCRISPASR_LIB_DIRset so the verification build takes build.rs path 1, no cmake). The account needed a one-time verified email first. crates.io consumers must link a pre-built lib (CRISPASR_LIB_DIR); the package does not vendor the C/C++ sources, so a from-source build needs the git dependency (build.rs hardened to say so). - PyPI
crispasr— DONE, latest 0.8.23 (also 0.8.22; both manually uploaded 2026-07-24, verified 2026-07-27 — the published 0.8.23 wheel is byte-identical topython/crispasr/_binding.py). Pure-Pythonpy3-none-anywheel; the nativelibcrispasris still installed separately by the user.
So all three registries (crates.io, PyPI, pub.dev) are now bootstrapped; only the optional CI/OIDC auto-trigger remains (below).
-
crates.io — ALREADY DONE (both crates first-published 2026-07-27). The recipe used (needs a verified email on the account + a prebuilt lib so the verification build skips cmake):
export CARGO_REGISTRY_TOKEN=... # token from https://crates.io/me CRISPASR_LIB_DIR=/usr/local/lib cargo publish --manifest-path crispasr-sys/Cargo.toml --allow-dirty sleep 30 CRISPASR_LIB_DIR=/usr/local/lib cargo publish --manifest-path crispasr/Cargo.toml --allow-dirty
For tag-triggered CI, add
CARGO_REGISTRY_TOKENrepo secret (Settings → Secrets → Actions) — CI build.rs will cmake from the checkout, so noCRISPASR_LIB_DIRneeded there. -
PyPI — ALREADY DONE (package first-published manually,
crispasr 0.8.220.8.23, 2026-07-24). Only optional remaining step for tag-triggered republish: at https://pypi.org/manage/account/publishing/ create a pending publisher — OwnerCrispStrobe, RepositoryCrispASR, Workflowrelease-wrappers.yml, Environmentpypi. Manualtwine uploadof a bumped version also works without it.
-
pub.dev (Dart) — ALREADY DONE (first-published manually to
crispasr 0.8.11).. Only remaining optional step: https://pub.dev/packages/crispasr/admin → enable Automated publishing, Repositorycd flutter/crispasr && dart pub get && dart pub publishCrispStrobe/CrispASR, Tag patternv{{version}}— do this only if you want tag-triggered republish (manualdart pub publishworks without it). -
Auto-trigger — DONE.
release-wrappers.ymlnow runs onpush: tags: ['v*']and publishes Rust (crates.io) + Dart (pub.dev). Python is handled separately byrelease-python-wheels.yml(see below), so the PyPI job was removed fromrelease-wrappers.ymlto avoid a double upload.
Ships the model I recommended: bundled CPU wheels → PyPI, GPU wheels → a
PEP 503 index on GitHub Pages (--extra-index-url .../whl/{cuda,vulkan}/),
plus a pure-Python sdist fallback. It runs on workflow_run after
release.yml finishes and REUSES the libcrispasr-<platform>[-cuda|-vulkan]
bundles that release.yml attaches to the GitHub Release — no native rebuild.
tools/stage_libs.py copies the libs into the crispasr package,
_binding.py:_find_lib() probes the package dir first, wheels are retagged per
platform with wheel tags. CPU matrix: linux x86_64 + arm64, macOS arm64
(Metal), windows x86_64. GPU: CUDA (linux + windows) + Vulkan (windows),
carrying +cuda/+vulkan local versions. Verified end-to-end locally on
macOS-arm64 (install → _find_lib picks the bundled dylib → CDLL loads).
REGISTRY SECRETS (all set 2026-07-27 via gh secret set):
PYPI_API_TOKEN(pre-existing),CARGO_REGISTRY_TOKEN(added — crates.io CI publish was otherwise skipped;gh secret listconfirms both).
GPU INDEX HOSTING: Pages already serves main/root (legacy Jekyll, the README
landing). The GPU index is committed to main under whl/ by the workflow
(Jekyll serves it at .../whl/{cuda,vulkan}/), leaving the README landing
untouched — no Pages reconfiguration. Index links point at Release-hosted
wheels, so main only carries tiny index.html files, regenerated from all
releases each run ([skip ci] commit, push-with-rebase-retry).
Assumption to verify on the first tagged run: linux wheels are labelled
manylinux_2_28_*; if a bundle needs newer glibc, bump the tag (auditwheel show).
- 60d F16 mimo-asr re-upload (HF). HF F16 (
cstr/mimo-asr-GGUF) is still legacy unfused layout; runtime fallback works but misses the 1.7× fused-QKV per-step decode. Needs a fresh BF16→F16 run (killed at 22 min on the 16 GB / 99%-full box, disk-thrash). Run on a 32+ GB box with non-99%-full external, thentools/patch_mimo_asr_fuse_qkv.pypatches to fused layout (~5 min). - 60e per-backend Q8_0 KV cosine validation.
CRISPASR_KV_QUANT={f16,q8_0,q4_0}wiring landed across 9 backends (defaults F16, bit-identical until opted in). Only mimo-asr diff-harness-validated at q8_0. Remaining 8 (qwen3_asr, voxtral, voxtral4b, granite_speech, gemma4_e2b, glm_asr, omniasr, orpheus, qwen3_tts) each need aCRISPASR_KV_QUANT=q8_0 crispasr-diff <backend>pass (≥0.98 gate) before any default-flip. ~5 min each, warm cache, zero code. - Vibevoice CUDA cache reuse re-test.
backend_needs_fresh_pred_graph()bypasses the pred-head graph cache on Metal+Vulkan+CUDA (CUDA on presumption). On a CUDA box runCRISPASR_VIBEVOICE_REUSE_PRED_GRAPH=1, confirm noGGML_ASSERT(src_backend_id != -1). If clean → drop CUDA from bypass list, recover ~30% per-synthesis caching. If assert fires → keep gated off; proper fix = recompute view→backend mapping fromview_src->bufferinggml_backend_sched_split_graph. - SYCL/HIP/ROCm cache-bypass extension. Same shape as CUDA; extend the
backend_needs_fresh_pred_graph()prefix list when a report comes in or a maintainer audits the upstream sched reset path there. - Per-backend
MADV_RANDOMpost-prefill.core_gguf::mmap_advise_random()exposed but unused; add one call between prefill and decode loop inmimo_asr_transcribe/qwen3_asr_transcribe/voxtral_transcribeetc. when a 32+ GB-box benchmark shows benefit (marginal on Q4_K; F16 is where it matters, unmeasurable on 16 GB). - Disk5 cleanup.
/Volumes/backupsat 99%. Safe to delete local unfusedmimo-asr-q4_k.gguf(superseded bymimo-asr-q4_k.fused.gguf+ HF fused) once future A/B not needed. - CI legacy
build.yml. Legacy whisper.cpp matrix (triggers on nonexistentbranches: [master]+tags: v*), failing every tag push since v0.4.x. Doesn't block (ci.yml/release.yml are the gates). Delete or repair after auditing whether any build-matrix combo isn't covered by new ci.yml.
Effort: Medium-large. Status: not started; design settled.
Goal: a --stream TTS mode that emits a 24 kHz PCM chunk every K AR steps
(stdout / HTTP) while the AR loop continues, dropping time-to-first-byte from full
wall-clock to "K AR steps + one chunked-VAE pass". NOT the Intel-Arc Vulkan
workgroup bug (already fixed via CPU fallback 31795a7 / VIBEVOICE_VAE_BACKEND=cpu).
geneing's chunked_vibevoice.patch (#52) nailed the chunk decomposition but
regressed on per-call sched_reset+sched_alloc_graph overhead — start there.
- Persistent VAE compute-graph reused across chunks (mostly mechanical; the
overhead that killed geneing's prototype). Mirror qwen3-tts
O15graph reuse (src/qwen3_tts.cpp:1037): build once atLk = max_chunk_latents, pin topology, reuse cached gallocr plan; cost = oneset_rows-style write op/call, not a rebuild. Benchmark to confirm the regression is gone before proceeding. - Causal padding on the σ-VAE conv stack — left-pad each chunk with previous chunk's tail, drop first L output samples, so chunk decode == full decode at boundaries (avoids phase artefacts). Ref: kokoro/voxtral4b causal-conv1d paths.
- Chunked transfer in HTTP TTS endpoint — depends on #58 (
POST /v1/audio/speech) landing first; wire cpp-httplib chunked-transfer forAccept: audio/wav; chunkedorstream=true.
- vibevoice (σ-VAE) — primary, largest win (positioned as realtime backend).
- qwen3-tts codec decode — 12 Hz vocoder; already has
O15graph reuse, extend to chunked output. - kokoro iSTFTNet — straight-line generator, cleaner chunking but iSTFT inverse window has the same boundary-artefact problem.
- Skip orpheus (SNAC already emits 24 kHz PCM single-pass — no win).
src/vibevoice.cpp / vibevoice_tts.cpp (chunked decode + graph reuse + causal
pad); examples/cli/crispasr_backend_vibevoice.cpp (--stream stdout PCM);
examples/cli/cli.cpp (--tts-stream flag); examples/server/server.cpp
(chunked-transfer, after #58); docs/tts.md; LEARNINGS.md (per-call ggml graph
overhead trap + graph-reuse cure).
Multi-chunk look-ahead beyond one chunk; kokoro/orpheus chunking as separate items; any change to AR decoding (only post-AR codec/VAE side is chunked).
Parent #75 round 1 shipped → HISTORY §81 (PR #63 merged + corrective batch + 75a/75b + 75d chunking + 75c-opt-1 server-side speed resampler). Remaining follow-ups: 75c-opt-2 (native-backend duration knobs) and 75e (streaming / mp3 / upload).
Remaining gaps documented in follow-up items: 75c-opt-2 (native-backend duration knobs), 75e (streaming response, mp3/opus encoding, voice upload/delete).
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: WER benchmarking on standard sets, GPU perf testing, decide whether streaming becomes default path
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: voxcpm2_tts sched migration; re-publish lfm2-audio GGUFs at Q5_K minimum quant
CLI knobs CrispASR exposes that CrisperWeaver doesn't. None blocking; listed so
the next parity-pass audit doesn't re-discover them. (--alt N shipped — see
HISTORY.)
TO DO (open parity gaps):
-
— DONE on both dispatch surfaces (multi-surface trap — the server has its own slice loop, see [[multi-surface-dispatch-trap]]):--offset-t MS/--duration MS- CLI (
crispasr_run.cpp process_one_input, feat/offset-duration): windows the decoded PCM to[offset, offset+duration)before VAD/chunking and shifts reported segment/word/token timestamps back into original-audio time. Was whisper-internal only (cli.cpp→wparams.offset_ms). Offset-past-end exits cleanly. Verified on parakeet-ctc + jfk: timestamps land at 5.0–11.0 s; word-level JSON offsets shifted too. - HTTP server (
crispasr_server.cpp do_transcribe, feat/server-offset-duration):offset_t_ms/duration_msform fields were parsed but never applied. Same window + shift, applied after the per-slice diarize re-walk (which matches segments by the unshifted slicet0_cs). Verified live viacrispasr --server --backend moonshine+ curl:offset_t_ms=5000→ seg 5.00–11.00;offset_t_ms=4000&duration_ms=3000→ seg 4.00–7.00; offset past end →{"text": ""}.
No session-C-ABI change needed:
crispasr_session_transcribeis a low-level PCM-buffer primitive — windowing is the caller's concern there. Docs indocs/cli.md+docs/server.md. The window arithmetic (initially copy-pasted into both surfaces) was factored intosrc/core/audio_window.h(core_audio_window::compute+trim) and unit-tested (tests/test-audio-window.cpp, 9 cases, model-free) so the CLI/server can't drift and the logic is CI-guarded. Still open: the Dart binding + CrisperWeaver UI rows (out of this repo's scope). - CLI (
-
Whisper decoder fallback knobs (
--word-thold,--entropy-thold,--logprob-thold,--no-speech-thold,--no-fallback,--temperature-inc) — already in Dart binding's TranscribeOptions; just add UI rows + l10n in CrisperWeaver Advanced Options. ~half a day. -
Subtitle line formatting:
--max-len+--split-on-punctalready work for all backends (applied post-hoc viacrispasr_make_disp_segments(all_segs, max_len, split_on_punct)incrispasr_run.cpp, verified 2026-07-17). Still whisper-only:--split-on-word(referenced only incli.cpp, nocrispasr_run.cpphookup). -
--carry-initial-prompt— sticky vs reset initial prompt across segments. Edge case, ~1 hour. -
— DONE (feat/print-confidence-nonwhisper). The flag was advertised in--print-confidence--helpand parsed but did nothing for non-whisper backends (only the whisper path incli.cpphonoured it) — a silently-broken flag, not a missing feature.crispasr_run.cppnow prints each segment's tokens with an inlineword[NN%]annotation after the transcript, via a newcrispasr_print_confidenceincrispasr_output.cpp(gatedelse ifafter the--altprinter so the two don't double up). Verified on parakeet-ctc (ans[85%]) and moonshine (my[67%] ask[61%] ,[42%]— low-confidence tokens now visible); no flag → single transcript line unchanged. Docs indocs/cli.md. (JSON/WTS exports already surfaced per-tokenconfidence; this closes the stdout half.) -
Token suppression (
--suppress-nst,--suppress-regex) — niche, whisper-specific; lowest priority.
Deferred --alt N follow-ups (low priority, v1 covers common case):
beam-search alt capture (siblings ≠ greedy alts, different capture + UX);
full word-level alt enumeration (BPE sub-word: only first content token gets
alts today); alt-picker popover widget test (CrisperWeaver Riverpod + l10n).
Status: 32 ASR + 21 TTS regression entries, 0 PLACEHOLDERs. CI at
.github/workflows/regression.yml runs nightly (04:00 UTC cron) + PR smoke-only.
Matrix: 22 ASR + 7 TTS = 29 backends. First nightly (2026-06-16) green after
fixing 6 failures. Architecture shipped in tests/regression/: manifest.json
(per-backend GGUF revision SHA + ref path + expected transcript + cosine
thresholds), run_one.py driver, regression.yml. Fixtures pinned in
cstr/crispasr-regression-fixtures (fixtures.revision SHA pins the whole set).
Next steps (each ~1 h/backend):
- Flip
skip_diffon backends where ref archives exist. - Add parakeet-tdt-0.6b-v3 (English) — cache the
.nemosource locally first (nvidia/parakeet-tdt-0.6b-v3not on the dev box yet). - Add canary + cohere + kyutai-stt + moonshine — all have reference modules in
tools/reference_backends/. - Add the TTS family (kokoro, indextts, qwen3-tts, chatterbox, vibevoice) — need WAV-output checksums or ASR-roundtrip rather than transcript equality.
- Promote to release gating — hook into
release.ymlpre-publish job so a regression aborts the tag.
Today CrispASR ships INDEXTTS_TEXT_NORMALIZER=<shell cmd> + tools/wetext-normalize.py
(commit 1bfe7c5a) — covers users who already have Python + wetext. No-Python deployments
(single-binary, Windows w/o Python, embedded) need a native path. Do NOT start speculatively.
Gap: default in-process preprocess_indextts_text() handles CJK char split, a subset of
char_rep_map punctuation, ASCII upper-case. Missing vs full wetext.Normalizer(lang='zh', operator='tn'): Arabic-numeral→hanzi, pinyin tone-digit restoration, dates, times, currency,
phone numbers, math/measurements/fractions/percent, EN contractions in Chinese. This isn't
cosmetic — the model often fails to emit stop_mel_token on un-pronounceable digit inputs and
burns max_mel_tokens=600, so digit-containing prompts break without TN.
Options (in preference order — implement per trigger below, not all):
- 95a. Hand-roll high-leverage rules in C++ (recommended first). digit-string→hanzi for
the 1-billion range (
零一二三四五六七八九+十百千万亿),年/月/日,点/分time, pinyin tone-digit lookup. ~300–600 LOC, no deps, covers ~90 % of prompts. Extendsrc/indextts.cpp:preprocess_indextts_textwith anormalize_chinese_numbers()pass + golden string-in/string-out tests. Effort: ~1 day. - 95d. Tiny FST reader in own C++ — consume upstream
pengzhendong/wetext.fstfiles (fsts/zh/tn/tagger.fst+verbalizer.fst, ~1 MB) without linking OpenFST. Parser + Tropical-semiring traverser + symbol-table loader + port of wetexttoken_parser.py(~200 LOC). Total ~500–800 LOC, zero deps. Byte-stable vs upstream except FST features not implemented. Newsrc/indextts_zh_tn.{h,cpp}; invoke viaINDEXTTS_TEXT_NORMALIZER=native. Effort: 2–4 days. - 95b. Vendor
kaldifst+ OpenFST + ship compiled.fst— byte-identical to upstream.third_party/openfst(~30–50 K LOC) +third_party/kaldifst(~5 K LOC), both Apache-2.0; real build-profile cost. Invoke viaINDEXTTS_TEXT_NORMALIZER=wetext. Effort: 3–5 days + ongoing submodule maintenance. Only as fallback if 95d's FST gaps keep biting. - 95c. PyInstaller-bundle the Python sidecar — ~50–80 MB bloat, per-platform pipeline. Almost certainly not worth it; listed so nobody re-discovers it.
Not alternatives (already surveyed as insufficient): ICU Transliterator, cn2an/pypinyin
(no C++ port), HF tokenizers normalizers.
DECISION-GATE — start only when one of:
- A user files a digit/date/pinyin-tone-digit prompt that breaks audio → do 95a, smallest rule set that fixes the reported case.
- Hand-rolled list reaches 2–3 entries (use case confirmed alive) → 95d becomes the right next step rather than letting 95a grow into one-off rules.
- Only if 95d's unimplemented FST features keep biting → fall back to 95b. Never vendor OpenFST speculatively.
Status: survey-only. RapidAI ships RapidTP-Aligns as a standalone timestamp-only model that
predicts timestamps from audio alone (no ASR text). Would be a third path alongside our native
decoder timestamps (whisper/parakeet-TDT/canary/cohere/kyutai-stt) and CTC forced aligners
(canary-ctc-aligner.gguf / qwen3-forced-aligner.gguf, which need the text as input).
Value when: cross-backend timestamp consistency independent of which ASR ran; trustworthy segment boundaries even when ASR text is wrong (diarization/VAD post-proc); single-forward-pass end-of-utterance/silence for streaming.
TODO (open questions to resolve before starting):
- What architecture upstream ships (README is one Chinese line; likely small Conformer-CTC over raw audio emitting frame-level boundary labels).
- License + upstream weights (ModelScope-hosted? FunASR derivative?).
- Quality vs our existing CTC aligners (
canary-ctc-aligner-q4_k.gguf~80 MB is fast+accurate; without a clear margin this is incremental).
Trigger (stays survey-only until one fires): a user reports the CTC aligner failing on a specific audio class (heavy code-switch, multi-speaker overlap, music behind speech); OR we add a streaming ASR endpoint needing sub-100-ms end-of-utterance prediction (currently a VAD silence heuristic).
Priority: HIGH — auto path (no --vad, no --chunk-seconds) tops out at
~82 % coverage on 60 s Japanese audio; users expect >95 %. Root cause is the TDT
decoder cold-starting each independent chunk (kLongAudioFallbackChunkSeconds=30
reinitializes the LSTM per chunk), losing 5–20 s of content per chunk interior —
not boundary stitching. Fix: keep TDT LSTM predictor state across frames like
NeMo's BatchedFrameASRTDT (stateful decoder, 4 s rolling buffer, running z-norm).
Phase 1 — stateful TDT decode (core change, ~200–300 LOC; hot loop <50 lines,
LSTM state threading is the work). Split API already exists (parakeet.h:
parakeet_encode, parakeet_decode_frames).
- Add
parakeet_decode_frames_stateful— likeparakeet_decode_framesbut accepts/returns LSTM hidden state (parakeet_lstm_state). The TDT loop atparakeet.cpp:1003already usesstate; make it an in/out param instead of initializing to SOS. - Add streaming mel with running z-norm — maintain running mean/variance across
frames (NeMo
get_norm_consts_per_frame) via a newparakeet_mel_streaming_context. - Wire into
crispasr_run.cpp: when long-audio fallback triggers and backend is parakeet/canary, use streaming decode (per 4 s frame: streaming-mel → encode → decode_frames_stateful(&lstm_state) → merge with LCS) instead of independent chunk transcription.
Phase 2 — tuning (benchmarking).
4. Frame-size sweep (1.6s, 2s, 4s, 8s) on benchmark corpus.
5. Running z-norm warmup: first frame per-frame z-norm, later frames EMA.
6. LCS delay tuning: lcs_delay = (buffer - frame) / model_stride.
src/parakeet.h— addparakeet_decode_frames_stateful,parakeet_mel_streaming_contextsrc/parakeet.cpp— LSTM state in/out, streaming-mel helpersrc/core/mel.{h,cpp}— running z-norm modeexamples/cli/crispasr_run.cpp— streaming decode path inprocess_one_inputexamples/cli/crispasr_backend_parakeet.cpp— streaming transcribe methodtests/test-issue-89-long-audio-fallback.cpp— pin streaming path activation
Success: python tests/benchmark_asr.py --audio yt_60s.wav --backend parakeet-ja --settings auto reports coverage ≥ 95 % on the #89 JA audio, without --vad.
Trigger: immediate — #89 open, current fix is partial mitigation (prevents 0-output but doesn't match NeMo quality).
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: Benchmark new aligners vs canary-ctc. (-am aliases now documented
in docs/cli.md — canary-ctc, wav2vec2 [12 langs], fastconformer [18 langs],
qwen3-forced.)
Status: feasible-but-BLOCKED-on-license; no code. Fourth VAD backend candidate alongside Silero, FireRedVAD, MarbleNet, Whisper-VAD-EncDec. Upstream ships cross-platform C bindings, prebuilt libs (Linux/macOS/Windows/Android/iOS/Web), and an ONNX path; runtime target 16 kHz, 10/16 ms hop — fits our VAD surface. Aimed at low-latency streaming turn detection, lighter than Silero.
Decision-gate: upstream license is Apache 2.0 plus additional no-compete/own-app-only conditions — treat distribution as BLOCKED until legal review or an explicit internal-only use case. Do NOT wire or ship until that's resolved.
Implementation plan (once unblocked):
- Decide prebuilt-native-lib path vs ONNX path (or both).
- Add
ten-vadalias in VAD registry + CLI so--vad -vm ten-vadworks. - Add auto-download metadata for the chosen artifact(s); keep Silero default.
- Run the boundary benchmark vs Silero + FireRedVAD on the short-gap / sentence-end test set.
- Document 16 kHz-in sampling-rate handling (resample other inputs first).
Trigger: a user wants lower-latency/lower-footprint VAD than Silero, OR we want a 4th native cross-platform backend — AND the license review clears (or an internal-only path is confirmed).
125. Issue #125 — multi-backend bug sweep from montvid (mostly DONE — validation + longer-term threads open)
12 findings from user montvid (RTX PRO 6000 Blackwell sm_120, CUDA 12.6) on
CrispASR v0.6.10 eaee2319. All fix commits landed and M1-Metal-smoked; remaining
work is GPU validation + a few longer-term root-causes. Reports cached at
/Volumes/backups/code/issue125-attachments/.
Open TODO items:
- P0 mimo-asr Blackwell segfault — hardening shipped (
a5a518c8), needs GPU confirmation. Real cause was0f0f0793(ggml-backend src-mutation log/restore), NOT the FA-mask commit. Ask montvid to rebuild from95d74455+ and rerun--backend mimo-asr -m auto --auto-download -f samples/jfk.wav -l en -np -nt; expect the v0.6.9 reference transcript. If segfault persists:gdbbacktrace; next suspects are Metal-debug commits leaking to CUDA (unlikely) or a Blackwell-specific ggml-cuda bug. - P1 funasr (longer-term): root-cause the audio adaptor/encoder collapse on
Blackwell CUDA (log
frames_splicedinfunasr_init_from_file; if 0 on 11 s JFK the adaptor is the failure; ideally diff adaptor output vs upstream FunASR Python ref). Loop guard +-lwiring already shipped. - P2 firered-asr (if it resurfaces): the "JFK silence-only
<Sil>!without --vad" subtask (suspected vocab/blank-id mismatch in auto-downloaded GGUF) was not reproduced locally; reopen only on a new report. - P4 gemma4-e2b (longer-term): add init sanity logs
audio_soft_token_id,proj_dimvsd_model, "audio projection weights found". Chunking + prefers_vad already shipped. - P6c kyutai-stt — DEFERRED: streaming model on batch dispatcher is a
footgun; a
--force-long-audiocap is now UX-nice only (P6b bounded the wallclock). Defer until a user reports.
Cross-finding open sweeps:
- Wider GPU validation matrix before any future
ggml-cuda/fattn*change: add mimo-asr, glm-asr, gemma4-e2b, voxtral, granite (multi-head audio-LLM backends). - Do NOT submit
tools/upstream-prs/06-cuda-fa-perhead-mask.mdupstream until that wider matrix validation lands. - Audit all 18 registered backends in
tools/test-all-backends.pyfor silent staleness (4 were missing entries before this sweep). - Defensive sweep of every
CAP_UNBOUNDED_INPUTdeclaration — confirm the encoder is genuinely unbounded, not just the dispatcher's input shape.
Completion triggers: mimo-asr next release no longer segfaults on Blackwell (montvid or same-GPU-class validation); 5-min EN clip clean on firered-asr / omniasr-llm / gemma4-e2b.
Reporter contact: montvid, GitHub #125.
Three small holes the sweep + #115 bisect surfaced. None are urgent; recording so the next contributor doesn't rediscover the same gaps.
The original 5 min sweep and the 90 s rerun both came back BOTH_EMPTY for omniasr-llm-300m-v2-q4_k.gguf: default and --chunk-overlap 0 both hit the 20 min per-pass wallclock on M1. Probably slow, possibly also has the truncation bug like its sibling backends. Can't tell without a faster box.
Fix shape. Re-run ./tools/check-overlap-save-bug.sh omniasr-llm with PER_RUN_TIMEOUT=2400 on a Linux x86 host (the VPS) or a Kaggle CPU kernel. If default produces materially less output than no-overlap, add to the opt-out list in examples/cli/crispasr_chunk_context_gate.h. If both produce the same content, mark VERIFIED-OK in the harness comment.
tools/test-all-backends.py has had a mimo-asr registry entry since 2026-05-02 (commit 2aeaf4c4); the test_transcribe function explicitly handles EMPTY output (line 693-695). Locally the test would have caught PLAN #115's silent-empty regression at runtime — but CI doesn't run test-all-backends.py against large-model backends (mimo Q4_K is 4.2 GB, doesn't fit in the standard runner disk budget per pre-release), so the regression shipped in v0.6.10 anyway.
Fix shape. Either (a) Kaggle scheduled-CI workflow that runs the full test-all-backends.py against the 4 LLM-class backends (mimo, voxtral, gemma4-e2b, granite-4.1-2b) on each main push — patterns in tools/kaggle/crispasr-regression.py already handle the model-download + heartbeat parts; (b) cheaper, a documented make smoke-llm-backends target that release scripts run before tagging. (a) is more reliable; (b) is one afternoon of work.
Issue #123 added the JA variant to the registry + README, but no row in any of PERFORMANCE.md's cohere tables (the long-form coverage at line 1374, the cross-backend matrix at line 1535, the per-length wall-time table at line 1555). The English cohere-transcribe is benchmarked across multiple Japanese / English / multilingual clips; the JA fine-tune isn't.
Fix shape. Run the JA variant on the same fixture set the English one used (TedX / JSUT clips per the model card; samples/jfk.wav is English so won't exercise the JA tuning). Drop one extra row into each cohere table with the JA numbers. ~30 min of inference + table updates once the fixtures are downloaded.
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: File upstream ggml issue with minimal repro; revert weight split if fixed
Status: Dia CPU full speech is verified (2026-10-02); SpeechT5 still has a decoder content mismatch; Parler + FastPitch not started.
- Encoder verified (cos > 0.999 all 12 layers). Runtime
src/speecht5_tts.cpp= encoder + decoder w/ KV cache + postnet + HiFi-GAN. GGUF/mnt/storage/speecht5/speecht5-tts-f16.gguf(300 MB). Convertermodels/convert-speecht5-to-gguf.py. - TO DO: decoder content mismatch — validate decoder per-layer against the Python reference.
- DONE 2026-10-02: fixed cross-attention Q/K RoPE, nucleus threshold inclusion and delayed-BOS timing. Independent pinned F32 source versus native F16 passes all 127 steps (cosine and magnitude); all 126 real-feedback inputs match, with the old BOS timing rejected by a negative control.
- F16 default and Q8 CPU full-speech roundtrips both pass WER 0. Q8 at 1/4/8 threads produces identical audio for each of two prompts. Model capacity replaces the 200-step cap; C ABI and CLI explicit limits/reset are verified. Shared CLI/server planning preserves the whole Dia dialogue.
- Proof: full-generation receipt, and HISTORY. Q8 does not pass the strict F16 numerical threshold; its separate speech gate does pass. Metal compiles, but hosted Mac VMs supplied no usable physical GPU runtime proof.
- FastPitch: ~1000 LOC stub, converter exists, needs NeMo model.
- Parler: ~857 LOC stub w/ 3 TODOs, T5 encoder + DAC decoder.
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: mimo-asr beam (blocked on PLAN #115) and lfm2-audio beam (needs KV+conv save/restore) still API stubs
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: AudioSeal diff-harness validation per status header
Replaces GPLv3 espeak-ng static link with modular permissively-licensed
phonemization. Phase 1+2 DONE (2026-06-07): header-only core/g2p_{en,de,fr,es}.h
(LTS rules + IPA-dict loaders + neural GRU G2P for EN), espeak_dlopen.h,
phonemizer.h/cpp cascade with auto-download from HF cstr/g2p-dicts, wired into
piper_tts.cpp, 202 unit assertions + 4 live TTS→ASR roundtrips. Cascade:
pre-gen IPA dict → builtin CMUdict+neural+LTS → OLaPh MIT dicts → espeak dlopen →
espeak popen. See HISTORY for coverage table + phoneme-inventory fixes.
a. More languages (need LTS rules; OLaPh MIT dicts already exist): Portuguese,
Italian, Dutch, Swedish, Czech, Danish, Finnish, Polish. Japanese/Chinese/Korean
need rules from piper-plus. Optionally port piper-plus (ayutaz/piper-plus, MIT,
branch dev) FR (1197 lines) + ES (620 lines) — more thorough than ours (NFD norm,
PUA mapping, syllabification, stress). No German in piper-plus — our g2p_de.h
fills that.
b. GGUF-embedded dicts (TTS.cpp pattern): embed phonemizer rules /
CMUdict / neural weights per-model via phonemizer.rules.keys /
phonemizer.rules.phonemes arrays for zero runtime external deps.
c. Gruut CRF (MIT, rhasspy/gruut): dict + CRF G2P (18 MB SQLite + CRFsuite/BSD)
for higher-quality OOV (compounds/loanwords), langs de/en/fr/es/it/nl/pt/ru/sv/cs/
ar/fa/sw. Port: extract SQLite lexicon + CRFsuite model, write ~100-line C++
feature extractor, link libcrfsuite. Lower priority now OLaPh covers most words.
d. Neural G2P weight distribution: publish standalone g2p_en.json (MeloTTS v3
melotts.g2p_en_json, base64 ~4 KB); loader already in g2p_en.h.
src/core/g2p_{en,de,fr,es}.h, src/espeak_dlopen.h, src/phonemizer.{h,cpp},
tests/test-g2p-{en,de,fr,es}.cpp, tests/test-espeak-phonemize.cpp,
tests/test-piper-roundtrip.sh.
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: §167e-h (granite-nle/mimo/moss/lfm2) beam need A/B validation on real models
16 items DONE at the 2026-06-20 audit (§176a,b,d–j,m,o–t); §176k + §176n later resolved. Full audit → PERFORMANCE.md "Runtime optimization audit — 2026-06-20". Open:
Do NOT implement without first measuring the KV round-trip fraction on the specific
backend (DIA_BENCH-style instrument). Dia measured at ~1.2% of decode → DEFERRED,
not worth the non-bit-identical, high-risk device-KV rewrite (ggml KV write/read
ordering on Metal is the codebase's most bug-prone pattern). Compute-bound transformer
decoders are NOT KV-bandwidth-bound — the "dominant bottleneck" premise is wrong.
Genuinely-still-host-side backends (verified 2026-07-12 by code read):
- SpeechT5 self-attn KV — host
std::vector<float>that grows per step (speecht5_tts.cpp:246-265), re-uploaded whole every step (:1035-1041,1084). Cross-attn KV already device-resident (§202). Measure fraction first (likely <5%). - Pocket-TTS self-attn KV — host, pre-sized to
max_seq(doesn't grow,pocket_tts.cpp:311-319), but reordered past window re-uploaded per step (:1207-1228). Measure first. - (Dia — measured ~1.2%, deferred. VoxCPM2/Parler/LFM2/KugelAudio already device-resident — do NOT re-chase.)
Approach if ever justified: IndexTTS/CSM 4D on-device
[head_dim, max_ctx, n_heads, n_layers]+ggml_view_4d/ggml_cpywrites; keep both paths gated; expect low-single-digit-% ceiling. Effort: Medium/backend, expected <2% payoff — deprioritize.
Shipped the real lever (env-gated matvec graph cache CRISPASR_FIRERED_MATVEC_CACHE,
default ON, bit-identical, pure Pareto). Residual: self-attn KV is a growing
std::vector<float> with scalar O(T²) scoring loop — for very long single-pass decodes
(hundreds of tokens) a pre-alloc 4D device KV + BLAS/ggml scoring could help there.
Measure on a long clip before investing — NOT the highest-impact lever for typical
clips. File: src/firered_asr.cpp.
Fast path shipped gated: rvq_encode_group (kyutai_stt.cpp:768) calls
core_rvq::encode_euclidean_per_stage when CRISPASR_KYUTAI_RVQ_FAST=1 (default OFF);
non-uniform dim / failure falls back to scalar. Helper + extraction/transpose unit-tested
(test-core-rvq, LABELS unit) — codes identical to scalar reference. Only end-to-end on a
real model is unverified.
- To flip default: run kyutai on a Kyutai STT GGUF (clip via
tools/kaggle/kyutai-stt-2.6b-convert) with flag on vs off, assert emitted RVQ codes byte-identical, then measure speedup. Until then scalar stays default + reference. Effort: Small (one-clip validation). Files:src/kyutai_stt.cpp,src/core/rvq.{h,cpp}.
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: CUDA validation (only Metal validated); nice-to-have
Status: comparative study (2026-07-03) of what CrispASR shares with the ggml-org ecosystem
(whisper.cpp + llama.cpp libmtmd). Framing: we're siblings on ggml — kernel-level wins
(flash-attn, CUDA MMQ/tensor cores, CPU tinyBLAS/sgemm, Metal mul_mm, Vulkan coopmat, k-quants)
are inherited free by keeping ggml sync current; llama.cpp's real advantages are at the
orchestration layer (imatrix, mixed-precision quant CLI, speculative decode, continuous batching).
Full comparison tables archived to HISTORY. Actionable adoption items:
- imatrix + per-tensor quant overrides — DONE, one loose end.
crispasr-quantizetakes--imatrix <file>+ iq4_nl/iq4_xs + requant-from-q8_0, and--tensor-type <regex>=<type>(llama.cpp-parity). Producersrc/crispasr_imatrix.{h,cpp}(setCRISPASR_IMATRIX_OUT) installed on decode scheduler of the major ASR-LLM backends; A/B harnesstools/imatrix_ab.pygates on transcript CER (cosine only tiebreaks). CC0 calib set atcstr/crispasr-imatrix-calib(+tools/imatrix-calib/); Kaggle kerneltools/kaggle/imatrix-quant/. Seedocs/quantize.md. OPEN: wire the collector into the remaining (non-ASR-decoder) backends if we ever want imatrix there. Corpus rule: diversity + in-distribution + language coverage is decisive — a narrow/mismatched corpus makes imatrix worse; use CC0 Common Voice (clean license). - FA default-on audit. Upstream flipped flash-attn to baseline. Re-check our two known FA bugs —
batched-FA corruption (§176h) and the Vulkan GQA-REPEAT-f16 path — against the current
FLASH_ATTN_EXTin vendored ggml; some may already be fixed upstream. - Keep ggml sync current — kernel wins are free there. Verify pinned ggml is recent enough for
Metal
mul_mm(M1 encoder prefill), Metal FA with head_size_k≠head_size_v (#12612 — matters for partial-RoPE ASR encoders: ark rot32/hd64, higgs, parakeet), CPU RMSNorm+MUL fusion, MMQ, tinyBLAS.
- Symmetric quantized-KV (q8_0) for Qwen-family audio-LLM decoders (ark, higgs-stt, moss-transcribe). Near-lossless, shrinks decoder KV. Must be symmetric K==V type or it silently drops off the fused FA path. Validate against per-head KV-stride bugs (#171) first.
- Continuous batching if the server ever needs concurrent transcription. Bespoke single-stream decode can't multi-slot today; #171 shows per-slot KV isolation is a real hazard.
- Prompt-lookup / n-gram speculative decoding for AR ASR heads (transcripts echo their own context → high acceptance). Heed §161: the draft must reproduce the target's exact realization or generation derails.
Don't bother: i-quants (low-bit breaks audio backbones; slow codebook decode on compute-bound TTS/ASR), KV defrag (deprecated upstream; single-pass audio doesn't fragment), paged attention (not merged upstream).
Don't converge to them on: model breadth (~60 vs ~8 archs), per-model long-audio chunking (overlap-merge beats their fixed-30 s), diffusion/flow TTS with batched CFG, neural diarization, voice cloning, single-file arch-autodetect UX.
ggml_rope_extarg order /GGML_ROPE_TYPE_*enum values vs pinnedggml/include/ggml.h(revised upstream).- Current speculative-decode flag namespace (
--spec-*vs older--model-draft/--draft-max). - mtmd mel params (128 HTK bins) + LFM2-Audio/MiniCPM-o audio status are DeepWiki-sourced — confirm from
clip.cppif load-bearing.
The #89 fix (VAD slice cap + per-slice single-pass + gap-fill) is shipped and manually verified (see HISTORY + LEARNINGS). This section makes it durable and gets it released. All subitems OPEN.
Coverage is protected only by manual runs; next parakeet refactor could silently
re-break #89. Drive a live test from the reazon baseball fixture
(cstr/crispasr-regression-fixtures → parakeet-tdt-0.6b-ja/reazon_baseball_14s/ audio.wav): concatenate ×3 (42.2 s — crosses 30 s auto-chunk + 12 s slice cap),
transcribe via the session ABI (carries cap + gap-fill), assert (a) 岡本 ≥3×,
(b) last timestamp reaches the third repetition, (c) a byte floor. Wire into
tests/ + tests/env-live-tests.sh (CRISPASR_MODEL_PARAKEET_JA,
CRISPASR_FIXTURE_PARAKEET_JA).
crispasr_server.cpp has its own slice loop (crispasr_compute_audio_slices
~line 410) separate from the CLI dispatcher and session ABI. Audit whether a JA
parakeet request via /v1/audio/transcriptions gets the slice cap + gap-fill; if
not (expected), mirror the policy — consult the backend's
vad_slice_cap_seconds() like crispasr_run.cpp, or route server parakeet
handling through the session-ABI path.
#89 reporter is on AMD/Vulkan (encoder instability worse there). Fix is
policy-level (slicing) so should transfer, but confirm once via MoltenVK on M1
(GGML_VULKAN=ON build, VK_ICD_FILENAMES gotcha — see LEARNINGS): yt_60s
default run should land ≈97 % recall like Metal.
parakeet-ja q4_k TDT decode is degenerate (repetition loop; pre-existing) while
CTC over the same file is clean. Two guards: (1) check what
-m auto --model-name parakeet-ja resolves to in
src/crispasr_model_registry.cpp — if q4_k, point at q8_0 (TDT byte-identical to
F16); (2) stderr hint when a JA parakeet GGUF with ≤q4 weights loads in TDT mode:
suggest --parakeet-decoder ctc or the q8_0 file.
Everything since v0.8.7: #89 series (slice cap, gap-fill, session mirror, tools,
CTC-head GGUFs), #218 moss fixes, #217 --align-only docs. Per dev-guide release
process: notes from git log v0.8.7..HEAD --oneline --no-merges into
RELEASE_NOTES_v0.8.8.md; scripts/bump-version.sh 0.8.8; push main; wait green
CI; push tag; gh release create with notes; then remove the repo-root notes file.
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: parakeet-ja CTC aligner upload, first-word t0 clamp, OWSM-CTC eval
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: Vulkan RADV verify, TitaNet batch/F16, chatterbox_campplus scalar cure.
(openvoice2 + firered_vad scalar cures already DONE — both use Accelerate cblas_sgemm
gated by CRISPASR_OV2_FORCE_SCALAR / CRISPASR_FIRERED_VAD_FORCE_SCALAR, verified 2026-07-17.)
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: baseline JA generation quality WIP, DiT graph caching perf follow-up
sims1253/starling (CUDA-graph ASR inference, RTX 5090, bf16) benchmarks CrispASR and overlaps our fleet (granite-speech, parakeet-tdt-v3, MOSS-Transcribe, qwen3-asr, ARK-ASR-3B, higgs-audio-v3-stt, cohere-transcribe — several via our cstr/* GGUFs). Their claim: stock transformers decode is launch-bound (GPU ~10% busy) → capturing decode steps + multi-step token loops into CUDA graphs gives 27–1180× vs 3–66× stock, byte-identical, WER-verified. Repo cloned to ~/code/starling (read 2026-07-06).
Benchmark-fairness (cheap, reputational): their CrispASR adapter
(benchmarks/engines.py:722, scripts/bench_qwen3_crispasr.py) times ONE full CLI subprocess
per clip — incl. process start + multi-GB F16 GGUF disk load + CUDA weight upload every rep —
while starling/stock numbers exclude model load and use warm reps. Cold-start seconds get
labeled engine speed. No crispasr numbers published yet, but the columns could appear anytime.
- Contact author: benchmark
crispasr-server(resident model, matches their server mode) or parse our stderr phase timings; offer setup help / a resident-mode adapter PR. (The doc to point them at now exists — see below.) - DONE —
docs/benchmarking.md: the fair-measurement contract (measure transcribe time, not cold start), three methods (server resident-model, in-process Session/ctypes = the apples-to-apples path, CLI-with-parsed-stderr-line), proof-of-work rules (non-zero exit/empty = FAIL; scale check; warmup + median + absolute ms), identical-load discipline, required reporting fields, and the phase-timing env vars (CRISPASR_VERBOSE,CRISPASR_<BACKEND>_BENCH,CRISPASR_METAL_PROFILE,CRISPASR_FC_PROFILE). Links the existingtools/benchmark_asr_engines.README.mdrather than duplicating it, and is linked from README's doc index so third-party benchers actually find it. Key fact it documents: the CLI/server stderr linetranscribed Xs audio in Ys (Zx realtime)already excludes model load (timer starts after init/decode/VAD) — so benchers should parse it instead of wrapping the process intime.
DECISION-GATE — option 0 first, gate options 1–4 on it: one Kaggle T4/P100 run measuring
GPU-busy% during granite/qwen3-asr AR decode (nvidia-smi dmon or nsys). If ggml decode is
already 60%+ busy, options 2–4 are duds (cf. §210 Metal ICB 1.8%, vibevoice CPU cache A/B 0%).
Only if ~10–20% busy (starling's stock baseline) → proceed.
Optimization options (ordered by evidence):
- Keep per-step graphs capture-friendly (mostly done, free). ggml-cuda already captures stable per-step graphs. Audit other AR backends where a decode loop shares a sched with helper graphs (EOS classifiers, connectors) — alternating graphs on one sched defeats gallocr + capture. (#171 per-purpose dedicated scheds already fixed pred head / per-KV-path LM steps.)
- Multi-step token-loop unroll (the real starling edge, CUDA-only). Unroll K greedy steps
into one graph: in-graph
GGML_OP_ARGMAX→ get_rows embedding feedback → static per-step KV positions; host sync per-K instead of per-token; EOS checked per block; byte-exact for greedy; stable topology per (n_past bucket, K). Substantial engineering, EXACTLY the cached-graph minefield of #171/#184/#220 (invariant: one graph per sched, or last-allocated-only). Do not start before option 0 numbers exist. - Self-speculative draft from CTC head (granite only). Draft from granite's encoder CTC head, verify with the LLM — no extra model; we already have granite CTC infra. Caveat: their batched spec at B≥16 loses (0.76×); B=1 spec is the win.
- Fused RMSNorm/SwiGLU steps — ggml-cuda already fuses some; only audit if option 0 shows launch-bound decode with capture ON.
Anti-options (skip — negative results transfer): INT8 weight-only quant on CUDA decode (SLOWER, launch-bound not bandwidth-bound; quant is a CPU/Metal win only); shape-bucketed graph caches without eviction (per-clip capture cost depresses RTFx at high shape diversity); torch.compile encoder fusion (not byte-exact — fp32 upcast + BatchNorm amplification).
Static code read (not benchmarks) of the direct ASR peer handy-computer/transcribe.cpp
(ggml C/C++ STT, heavy roster overlap). They have the more performant core ASR
engine (CPU, flash, AR decoder, on-target encoder, features, streaming latency)
via uniform shared-infra; we win quantization QUALITY (imatrix, unmatched),
VAD-gated long-form robustness, and breadth (ASR+TTS). Full scoreboard + file:line
findings in HISTORY. Every port gated on decoded-output A/B (§176/§210/§227 rule).
- GATE on measurement (do first). One M1 + one T4/CPU encoder-RTF A/B for wins 1-2. Port each only if it moves the encoder ≥5%.
- Force
GGML_LLAMAFILE ON— DONE (verified 2026-07-17).CMakeLists.txt:199set(GGML_LLAMAFILE_DEFAULT ON)overrides ggml's default-OFF (ggml/CMakeLists.txt:113), landed with the §232 work. Note §232's closed-notes call the net effect "neutral on Q4_K/x86" but it stays ON as a low-risk default. - Metal (+Vulkan)
GGML_OP_NORM_AFFINEkernel — our fused norm op (ggml.h:500,ggml.c:3162, used 7×/block incore/fastconformer.h+core/sanm.h) has ZERO Metal/Vulkan impl → falls to CPU with GPU↔CPU copies in the hot loop (~336/canary encode). Add the kernel +supports_opcase (ggml-metal-device.m:1392-1393), OR gatecore_conformer/core_sanmto plainggml_norm+mul+add on Metal/Vulkan. Gate on M1 canary encode A/B. - In-graph argmax + used-prefix-only KV snapshot for cohere/AR decode.
Ours (
src/cohere.cpp) does full-vocab device→host + host max/softmax every token (:2606,2614). Beam search is the slow form: N sequential single-token forwards (core/beam_decode.h:412-441) each preceded by a full-KV deep-copy regardless ofn_past(core/attention.h:230-284) — the #161 driver. Fix: in-graphggml_argmax+ 1-int readback (theirarch/cohere/model.cpp:1211-1213); snapshot only then_pastprefix. - Flash head-dim padding helper + real per-backend flash-capability gate.
Ours = ~40 scattered per-model
p.flash_attn+ reactivegetenvguards, no head-dim padding, Turing sm_75 −9-10% fusion regression unhandled at model layer (fold in the opt-out from memory [[project_flash_attn_turing_regression]]). Ref theirpad_head_dim(arch/moonshine_streaming/decoder.cpp:37). - Consolidate per-backend pow2 FFT into one shared non-pow2 frontend. Ours
is per-backend pow2-only (
core/mel.h:33), backends ship duplicate hand-rolled copies (voxtral.cpp:446"cargo-cult",core/fft.h:6-8). (Optional: switch miniaudio linear resamplecrispasr_audio.cpp:406→ windowed-sinc if fidelity matters.) - (stretch) Utterance-batched encoder + batched multi-utterance decode for offline throughput — large, and the cached-graph minefield of #171/#184/§227.2. Do NOT start before action 0.
- (cleanup) Bucket-table quant tensor classifier (their
tools/transcribe-quantize/policy.cpp:50-278) to replace our per-arch if-ladder (examples/crispasr-quantize/main.cpp:559-684) — refactor, not a perf change.
Note: our single-stream conformer core is actually leaner (fused norm_affine,
fused GLU, load-time BN fold, zero-copy strided rel_shift fastconformer.h:43-47) —
port that rel_shift trick INTO them / keep it; §176 encoder-graph cache confirmed a
dud (keep CRISPASR_PARAKEET_ENC_CACHE OFF).
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: optional CPU direct-path rescue (avoid lm_head slot blit)
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: CV3 flow-steps default 10->6 flip pending human listen; transducer GPU decode pending parallel CUDA verdict
Completed work archived to HISTORY.md (PLAN compaction 2026-07-17).
Still open: voice-clone roundtrip validation, output-gain/clipping check, denoise token, rerun parity kernel
Status: Engine is competitive (CA wins/ties most of 11 shared models vs transcribe.cpp on Kaggle P100). Big losses (RNNT/TDT CPU-cblas decode, moonshine CPU-only-from-CLI bug) are FIXED and default-flipped. GPU-forwarding audit done across all CLI adapters. Remaining gaps are either architectural (moonshine-streaming) or measurement/validation, all GPU-hardware-gated (Kaggle P100). Full history + A/B tables → HISTORY.
-
[HIGH, LARGE] Moonshine-streaming 18× loss — banded/blocked windowed attention. Bottleneck is O(T²) sliding-window masked flash-attn (dense T_enc×T_enc F16 mask/layer ×6, wl=16/wr=4), NOT frame-by-frame encoding (encoder runs once,
moonshine_streaming.cpp:808; masks can't be dropped — LEARNING 17 degenerate output). Need a banded flash kernel (no native banded flash in ggml). Own campaign. Keep--backend moonshinefor offline. -
[MED] Extend persistent ggml decode to beam/RNNT/maes paths. Greedy TDT/RNNT decode is ported+flipped, but
parakeet_rnnt_decode,parakeet_tdt_beam_decode,*_maes_decode, nemotron beam still call cblaspredictor_step/joint_stepbeam_size× per step. Reusecore_rnnt_ggml::Decoder(one Decoder serves all hypotheses, state passed per call). Low risk, proven pattern. Needs a parakeet-rnnt model to validate the RNNT path. -
[MED, GPU] dia TTS — widen default beyond Metal. Default GPU on Metal only; CUDA/Vulkan opt-in (
DIA_TTS_GPU=1). get_rows contiguity crash fixed (65a5d30c) but CUDA greedy tokens DIVERGE from CPU at step 0 (M1 Metal matched, argmax 568). TODO: ASR-roundtrip the CUDA-generated audio (the real HARD-RULE-#3 test) — if intelligible it's benign FP → widen default; if garbled, bisect the decode graph for another non-contiguous/precision-sensitive op.load_weights_split(enc+DAC→GPU, decoder→CPU) is the fallback. -
[MED] Re-run P100 competitive scoreboard (v16). After decode flips + moonshine GPU-forwarding fix (
d46839ca), re-measure parakeet/nemotron/moonshine TOTAL RTF vs transcribe.cpp; updatedocs/performance.md. Run v16 kernel to confirm moonshine P100 CA ~0.012-0.015 (would make CA lead 6–5). Kernels:tools/kaggle/parakeet-ggml-decode-ab/(P100, exercises persistent path),tools/kaggle/gpu-pin-ab. Measurement, not a code change. -
[LOW] In-graph argmax for transducer greedy path. 2 int32 readbacks vs 8198-logit readback, greedy no-hotword no-sampling path only (hotword-bias/temperature still need full readback). Minor now that persistent-graph decode won.
parakeet.cpppredictor/joint step. -
[LOW] Parakeet same-version A/B (Fix 5). CA uses parakeet-tdt-0.6b-v3 (25 EU langs) vs TC's v2 (EN-only) — the residual gap may be the model. Download v2 GGUF, benchmark both, note in docs. Model local:
parakeet-tdt-0.6b-v3-q4_k.gguf(467 MB). Quick, VPS.
- No more optimisations on VPS — all remaining wins need GPU hardware (Kaggle) to validate+measure.
- Flip a GPU default only if transcript/WER parity holds AND GPU decode < cblas decode ON THE TARGET
platform. Metal is excluded from the RNNT/TDT ggml-decode flip (Accelerate cblas beats it there;
P100 win is a slow-OpenBLAS effect — LEARNING 34). Overrides:
PARAKEET/NEMOTRON_GGML_DECODE=1/0,RNNT_GGML_PERSTEP,CRISPASR_{PARAFORMER,M2M100,T5}_GPU=1/0,MOONSHINE_ALL_GPU=1,MOONSHINE_ENC_ATTN=manual,DIA_TTS_GPU=1/0. - Closed (no action): moonshine decode hybrid placement, moonshine encoder manual-attn (slower, opt-in), f5_tts GPU (7.8× base fix; cheap levers tapped out — needs Metal-capture profiling), paraformer/m2m100/t5 GPU (flipped GPU on CUDA/Vulkan), titanet/diarize (Accelerate-BLAS, GPU port poor EV), GGML_LLAMAFILE ON (neutral on Q4_K/x86, kept as default).
Status: full ASR+TTS+codec+pipeline re-verification 2026-07-11 (see PERFORMANCE.md "Runtime
Optimization Audit — Re-verification (2026-07-11)"). Gate: every item needs a target GGUF (q8_0
preferred, to isolate from q4_k quant noise) + before/after parity + latency — do NOT land a perf
change on a compile-only check. Note: most per-model TTS "flash not wired" claims are false — flash
reaches Orpheus/OuteTTS/Zonos/TADA/Chatterbox/CSM via shared core_attn::kv_self_attn
(src/core/attention.h:665,903). Real flash gaps are only manual-soft_max backends (dia, speecht5,
parler) and structurally-can't-flash relpos models (melotts, piper).
- Decode-step graph cache for remaining LLM/AR backends → §210 follow-up (shape-stable bucketed
decode / CUDA-graph capture). Templates: qwen3-tts Lk-bucket, granite §210 gallocr, mimo
step_t1_gf. - Batched-CFG (B=2) for remaining TTS → §215. Un-migrated diffusion/DiT targets: f5, dots, kugelaudio, pocket (+ dia/speecht5/parler once they get device KV). Respect Metal quant-B=2 gotcha (dequant batched-against weights q*→F16 once) and item-24 (don't CPU-batch a GPU pipeline).
- gallocr cross-call UAF audit → #215e / encoder-graph-cache removal (#235). Encoder-graph caching stays OFF (measured dud + GPU UAF).
- Decode-step graph cache — same design as CrispEmbed Tier-1 #1; CrispASR is further along (§210 CUDA-graph-capture template).
- ggml-metal ICB replay — Apple-side equivalent of §210's CUDA-graph capture; ggml-metal has no ICB path. Depends on a stable per-step graph. Shared ggml submodule — do once, both repos benefit.
| P | Area | Gap | File |
|---|---|---|---|
| P0 | melotts / piper | Scalar O(H·T²·D) relpos attention (can't flash — additive bias); HiFi-GAN 17.9s of 26.3s VPS total. Needs manual-attn ggml graph or BLAS | melotts.cpp, piper_tts.cpp |
| P0 | voxcpm2_tts | CPU-only (Metal SIGSEGV); manual per-step host KV re-upload | voxcpm2_tts.cpp:106-111 |
| P0 | openvoice2 | 16-layer WaveNet + ref-encoder Conv2d/GRU scalar CPU | openvoice2.cpp |
| P1 | voxtral/voxtral4b enc, mimo LLM dec | Attention not on flash_attn_ext (manual soft_max) | voxtral4b.cpp, mimo_asr.cpp |
| P1 | firered/glm/funasr/qwen3/omniasr/mimo | Beam = replay; add KV snapshot pool (canary/moonshine/kyutai template) | — |
| P2 | scalar CPU hotpaths | RNN-T LSTM pred+joint; granite cpu_linear+depthwise; rvq encode; istft IRFFT; titanet mel; diarize apply_xcorr |
core/rvq.cpp, core/istft.h, titanet.cpp:740 |
| P2 | parakeet/nemotron | Batched sgemm decode opt-in default-OFF — validate + flip on | CRISPASR_TDT_BATCH |
| P3 | threading | Hardcoded default-4 threads in ~90 sites; adopt whisper-core's min(4, hw) |
crispasr_c_api.cpp |
| P3 | pyannote | Runs per-slice not once-over-audio (#107); RNNoise recreates state+resamplers/call | crispasr_diarize.cpp |
(Struck P0 firered_asr — self-attn KV already cached, remaining O(T²) is scalar attention scoring that often loses on M1/CPU — and P2 align_wav2vec2_ctc — now on the §176e resident cache — dropped as non-gaps; detail in HISTORY.)
Techniques shipped for the FastConformer family (fastconformer.h consumers: parakeet, canary, canary_ctc, canary_qwen, lfm2_audio, nemotron); roll each out where it applies.
- (1) F16-weights-in-quantized-GGUF audit (the 35%-of-encoder trap): converters
storing matmul-consumed weights 3D/1D get them skipped by the quantizer's 2D rule → run
on the ~6×-slower CPU F16 mul_mat path. Method: list tensors with
GGUFReader, flag F16 tensors feedingmul_mat; or run per-node profiler and look forMUL_MAT f16high-% rows. Suspects: cohere-transcribe, firered-asr, granite/conformer_ibm, sanm/paraformer conv, moonshine, TTS vocoders (hifigan/seanet/dac k=1 — qwen3-tts FASTCONV fixed the cast side, not storage). Fix: quantizer carve-out (+Q8_0 floor + idempotency) + load-time repack viacore_conformer::repack_conv_pw_q8+ fleet requant kernel (tools/kaggle/fc-pw-requant). - (2) Generalize the per-node profiler: move
cc_prof_cb(sched eval callback, aggregates by op+src-type+shape,CRISPASR_FC_PROFILE=1) tosrc/core/sched_prof.hso the audit targets share it.CRISPASR_SCHED_PROFILE=1now covers Canary CTC, Cohere, FireRed-ASR, Granite Speech, Moonshine, Moonshine Streaming and Paraformer; the callback reports relative shares because forcing one split per node adds dispatch overhead. - (3) Fused QKV (
core_conformer::fuse_qkvis tensor-generic): bit-identical ~free win wherever Q/K/V share an input. Already deployed on ~10 backends (parakeet/canary/canary_qwen/canary_ctc/lfm2_audio viacore_conformer::fuse_qkvunderCRISPASR_FC_FUSED_QKV; voxtral/voxtral4b/qwen3_asr/higgs_stt/qwen3_tts have their own env-gated fused impls; nemotron deliberately opts out). Still open for the remaining named targets: whisper encoders, cohere, firered_asr, granite_speech, sanm/paraformer, AED decoder self-attn. - (4) Strided flash inputs: grep
ggml_contfeedingflash_attn_extacross src/ — the kernel reads strided views (llama.cpp does); each cont is a full tensor copy/layer/pass. - (5) Manual-attn-on-CUDA gate: any backend whose flash mask has ne[2]>1 (per-head
bias) silently falls back to CPU on CUDA —
GGML_SCHED_DEBUG=2on a CUDA box shows flash nodes on CPU splits. Check cohere-transcribe (Shaw rel-pos) first. Reusefc_gpu_manual_attn- BlockParams.manual_attn.
- (6) -inf pad masking + bucketed persistent graphs (
CRISPASR_FC_BUCKET): reusable base for batched inference + CUDA-graph capture; correct padding for future streaming/batching (finite mask constants get overrun by pad garbage — LEARNINGS 2026-07-12). - (7) Q8_0 floor for decode-critical tensors in sub-8-bit quants (quantizer): conv pw done; consider same floor for other high-sensitivity small tensors flagged by future A/Bs.
Suggested order: (2) profiler → (1) audit sweep with it (one Kaggle CPU kernel over the
registry, collect MUL_MAT f16 % per backend) → fix top offenders → (4)/(3) mechanical
wins alongside → (5) after the §246 CUDA per-stage data.
Context: #274 flagged 0xShug0/audio.cpp as a comparable C++ inference engine. Audit performed 2026-07-19. They share ~6 ASR and ~12 TTS models with us; the rest is non-overlapping. CrispASR is broader on ASR (43 vs ~6) and TTS (48 vs ~12) but audio.cpp covers categories we don't touch yet.
Must-compare on Kaggle (GPU + CPU kernels):
| Model | audio.cpp claim | Our backend | Metric |
|---|---|---|---|
| Voxtral Realtime | RTF 0.089 (11.2×), 15.7× Q8_0 | voxtral4b |
RTF, WER on librispeech-test-clean |
| VibeVoice TTS 1.5B | 5.15× realtime (93.9 min podcast in 18.2 min) | vibevoice-tts / vibevoice-1.5b |
RTF, MOS-proxy (ASR roundtrip) |
| Qwen3-ASR | (no claim) | qwen3 |
WER, RTF |
| Higgs Audio STT | (no claim) | higgs-stt |
WER, RTF |
| Nemotron 3.5 | (no claim) | nemotron |
WER, RTF |
| Chatterbox TTS | (no claim) | chatterbox |
RTF, ASR roundtrip |
| IndexTTS2 | (no claim) | indextts |
RTF, ASR roundtrip |
| OmniVoice | (no claim) | omnivoice |
RTF, ASR roundtrip |
Approach: build audio.cpp on Kaggle (their CMake), run their CLI on the same WAV files we use, collect wall-clock + output text, compare WER/RTF side-by-side.
All licenses verified 2026-07-19. Only open-licensed models listed (Vevo2 is CC-BY-NC-4.0 = non-commercial, skip; Stable Audio is Stability Community License = commercial-restricted, skip).
Source separation (new category):
- HTDemucs — hybrid transformer demucs, music/voice separation. MIT (Meta).
DONE (2026-07-19, VPS+Kaggle). ~2720 lines C++. Full parity with Python.
12-point checklist verified. All ops implemented:
STFT→CaC→norm→encoder(Conv2d+DConv+GLU+freq_emb)→channel_up→
2D/1D sin pos emb→LayerNorm→CrossTransformer(5 layers, self+cross attn)→
channel_down→decoder(skip+GLU+ConvTranspose2d+time_decoder)→
CaC unmask→iSTFT+time_denorm→per-source stereo PCM.
All wired: CLI adapter + factory + GGUF auto-detect + model registry
(Q4_K default) +
--separatedispatch + C API session + Python binding + Go LDFLAGS. GGUFs on HF (cstr/htdemucs-GGUF): F16 (81 MB), Q8_0 (53 MB), Q4_K (38 MB). Kaggle: 21 reference stages validated. VPS (8 GB) can't run inference (swap pressure); validated on Kaggle (16 GB). - Mel-Band RoFormer — frequency-band source separation. MIT.
The #274 reporter specifically mentioned this. Do both — they're complementary
(HTDemucs = 4-stem, RoFormer = vocal/instrumental).
TAKEN (M1/Metal session, 2026-07-19). Picked for this box because it is
non-autoregressive — no sampling, so the diff harness gives deterministic
per-stage cos verdicts with no torch-vs-mt19937 RNG mismatch (unlike an AR TTS
port), and the model is small enough to run the Python reference and the C++
port in 16 GB. No NeMo dependency (NeMo import is broken on this Mac).
Regime: read the Python blueprint line-by-line →
tools/reference_backends/ mel_band_roformer.pydumper → converter → C++ backend → per-stage diff → acceptance = decoded-output roundtrip (separated stems judged by SDR/ASR on the vocal stem, not by cos alone — HARD RULE #3).Coordination with the in-flight HTDemucs work (same category): HTDemucs is unchecked above but IS being worked (converter
a6a447587+ 21-stage reference dumper60ada0a06). We share the new--task separatesurface (CLI flag, stem output/WAV writing, backend-capability bit). Whoever lands that scaffolding first owns it; the second builds on it rather than adding a parallel one. Per-backend files (src/mel_band_roformer.*, converter, reference dumper, registry/CMake entries) are additive and conflict-free.
Voice conversion (new category):
- Seed-VC — zero-shot voice conversion. GPL-3.0 (Plachta/Seed-VC on HF). Encoder + flow-matching + vocoder. Medium effort. GPL is fine for our Apache-2.0 project as long as the model weights are loaded at runtime (not statically linked). Scoped 2026-07-19: Tiny XLSR variant = 142 MB (25M params), fits 8GB VPS (~3 GB RAM with PyTorch). Architecture: XLSR→CAMPPlus speaker emb→length regulator→DiT UViT flow-matching (30 ODE steps)→HiFi-GAN vocoder. Risk: flow-matching requires noise pinning for deterministic diffs; FunASR+ModelScope deps are heavy. Repo archived.
TTS:
- Supertonic 3 — claims 200×+ realtime on CUDA. OpenRAIL-M (Supertone/ supertonic-3 on HF). Permissive (attribution + responsible use). Scoped 2026-07-19: ~400 MB total (4 ONNX components: text_encoder 36 MB, duration_predictor 4 MB, vector_estimator 257 MB, vocoder 101 MB). Non-autoregressive flow-matching (5-12 steps). 44.1 kHz output, 31 languages, 10 voices. Runs on Raspberry Pi (~800 MB RAM). ONNX-only distribution = custom ref approach needed (dump ONNX intermediate tensors, no PyTorch hooks). Deterministic diffs. VPS-ready.
- MioTTS — voice cloning TTS. Apache-2.0 (Aratako/MioTTS-0.6B, Qwen3-based).
CODEC PARITY ACHIEVED (2026-07-19). Full pipeline: Qwen3 LLM (28L, 1024d, GQA
16/8, head_dim=128, vocab=164480) + MioCodec-25Hz-24kHz (FSQ → wave_prenet → conv →
ResNet → AdaLN-Zero decoder → ResNet → iSTFT → 24kHz waveform).
Parity results (diff harness):
- FSQ dequant: cos=1.0, wave_prenet: cos=1.0, audio: cos=0.999 (F32)
- Q8_0 (723 MB): audio cos=0.954 — acceptable for TTS
- Q4_K (397 MB): audio cos=0.069 — codec needs F16 for quality
- LLM forward: argmax matches Python (tokens correct), tail-cos lower Remaining (non-blocking for audio generation):
- Tokenizer integration (Qwen3 BPE + ChatML) for
miotts_synthesize() - CLI adapter (
crispasr_backend_miotts.cpp) + registry + 12-point checklist - Mixed quantization (LLM=Q4_K + codec=F16) for optimal quality/size
- HF upload: F16 + Q8_0 GGUFs with README
- Voice cloning: global encoder (reference audio → 128-d embedding)
Audio codec:
- MioCodec v2 — standalone audio codec. MIT (confirmed on HF + GitHub). TAKEN (VPS session, 2026-07-19). 133M params, ~530 MB F32. Architecture: WavLM-base+ encoder → FSQ quantizer (1 codebook, 12800 vocab, 25 Hz) → iSTFT decoder. Encode+decode is the full pipeline (no external vocoder for v2). Also does standalone voice conversion (swap global embedding). ~1.2 GB RAM — fits easily. Unblocks MioTTS port (shared codec). Deterministic diffs (feed-forward, no sampling). Regime: read Python → converter → backend → per-stage diff → roundtrip (encode→decode parity with original audio).
Music generation:
- ACE-Step 1.5 — music generation. Apache-2.0. 3.5B params, 8.3 GB. Scoped 2026-07-19: mT5 text encoder + 3.5B linear transformer flow-matching + DCAE latent codec + vocoder. Kaggle-only (weights alone exceed 8GB RAM). 4 min max.
- HeartMuLa — music generation. Apache-2.0 (HeartMuLa/HeartMuLa-oss-3B). Scoped 2026-07-19: 4B params, ~16 GB F32. AR LM + HeartCodec (12.5 Hz). Kaggle-only.
Diarization:
- Sortformer — NVIDIA dedicated diarization. NOT Apache-2.0 — streaming v2 is CC-BY-4.0, v2.1 is NVIDIA Open Model License. ~470 MB, 4-speaker hard cap. Scoped 2026-07-19: End-to-end (no VAD/clustering pipeline). Fast-Conformer 18L + Transformer 18L → per-frame 4-speaker labels. Community ONNX exports exist. ~750 MB RAM via ONNX, ~125 MB Q4_K GGUF. VPS-feasible but: (a) license isn't Apache-2.0, (b) 4-speaker cap vs pyannote/MOSS unlimited, (c) need quality eval first.
Skipped (license incompatible):
Vevo2— CC-BY-NC-4.0 (non-commercial only)Stable Audio 3— Stability Community License (commercial requires separate agreement)
- Benchmark first — head-to-head on shared models (Kaggle GPU kernel). If we're slower on Voxtral/VibeVoice, fix perf before adding scope.
- Source separation — HTDemucs + Mel-Band RoFormer (both MIT, #274's ask).
- TTS — Supertonic 3 (speed benchmark), MioTTS (voice cloning, Apache-2.0).
- Voice conversion — Seed-VC (GPL-3.0, runtime-loaded weights = OK).
- Music generation — ACE-Step + HeartMuLa (both Apache-2.0, new category).
- Diarization — Sortformer (evaluate vs moss-diarize first).
-
tools/sync_go_cgo_ldflags.py— add-lwebrtc-vad(CI will flag) - Dedicated live test comparing segment boundaries vs Silero/FireRed
Context: CometBeat TRANSCRIPTION_SOTA_HANDOFF.md lists these as CrispASR ggml port targets. New category: audio → note events (MIDI).
Models:
- piano_transcription_inference (ByteDance/Kong, Apache-2.0). CRNN: 4× (4-layer Conv2d + 2-layer BiGRU) sub-networks (frame/onset/offset/velocity), 88-key output at 100fps. ~172 MB checkpoint, ~86 MB F16 GGUF. TAKEN (VPS session, 2026-07-19).
- Basic Pitch (Spotify, Apache-2.0). Polyphonic, instrument-agnostic
audio → note events. DONE (2026-08-30, VPS). The network is tiny (~40k
conv weights, 112 KB F16 GGUF) — the model is really its front end.
Weights come from the
nmp.onnxthat ships inside the package (basic_pitch/saved_models/icassp_2022/, sha2562c3c1d14…59a0ec): the ONNX has the BatchNorms already folded AND carries the nnAudio CQT kernels as initializers, somodels/convert-basic-pitch-to-gguf.pycopies them bit-for-bit instead of reimplementingscipy.signal.firwin2.src/core/cqt.hcould NOT be reused — it is the direct-kernel librosa CQT (zero pad,1+n/hopframes); Basic Pitch trained on nnAudio CQT2010v2 (one top-octave kernel bank, 9 octaves of recursive x2 decimation, reflect padding,sqrt(lengths)rescale). Newsrc/core/cqt2010v2.h;cqt.his untouched so BTC parity is unaffected. Parity vstools/reference_backends/basic_pitch.py(onnxruntime), synthetic polyphonic clip: every stage cos = 1.000000 — audio window, CQT magnitude, NormalizedLog, harmonic stack, all three heads, and the stitched full-file posteriorgrams. End-to-end note events 27/27 exact (start, end, MIDI, velocity) at F32; F16 shifts 2 of 27 note ENDS by one frame at a threshold boundary. jfk.wav (16 kHz → resampled, so the resampler differs from librosa's): all stages ≥ 0.9991, note events 11/11 exact at both F16 and F32. Wired: converter +src/basic_pitch.{h,cpp}+ CLI adapter +--pianodispatch (routes on GGUF arch) + arch map + registry row +crispasr-diffbranch. GGUF not yet uploaded — registry row points atcstr/basic-pitch-GGUF, build locally until then. Not ported: pitch bends (get_pitch_bends) and MIDI file writing — the CLI emits note events, same shape aspiano-transcription. - MT3 (Google, Apache-2.0). Seq2seq multi-instrument. Feasibility check
DONE 2026-08-30 → GO,
docs/music-transcription/mt3-feasibility.md. Corrections to the old line: the checkpoint is 171.6 MB (60 M params, ~120 MB F16 — verified against the live GCS listing), not "1 GB+", and the T5X/zarr checkpoint decodes with stdlib+numpy (no JAX/t5x/TensorStore). ~80% ofsrc/t5_translate.cppreuses directly; the gotcha is MT3 uses sinusoidal ABSOLUTE positions (FixedEmbed, network.py:180/225 — zero relative_attention_bias params in the checkpoint), so the T5 runtime needs a second positional branch. Gate on note-level F1 vsopenmirlab/mt3-infer(numpy ref dumper required — mt3-infer excludes Magenta MT3 deps), NOT cosine: the risk is the tie-section cross-segment note stitching. ~8 working days.
New CLI surface: --task transcribe-music / --backend piano-transcription
→ MIDI output file.
Context: §250 covers note-event models. This covers the shared DSP those
models need, plus the roster items still absent. Status reply written back to
CometBeat: mus-textbook/docs/TRANSCRIPTION_CRISPASR_STATUS.md.
Already shipped, for the record (so nobody re-ports them): CREPE F0
(src/crepe.cpp, --pitch, Dart + WASM + pub.dev crispasr 0.8.16), HTDemucs
and Mel-Band RoFormer separation, piano_transcription (§250).
-
core/cqt.h— constant-Q transform. DONE (src/core/cqt.h,tests/test-core-cqt.cpp726 assertions,tools/cqt_librosa_parity.py). Direct time-domain Brown kernels, not librosa's recursive downsampling. Measured vs librosa 0.11.0 at BTC params: per-frame shape correlation median 0.9999 (mean 0.9721, min 0.1136), peak-bin exact match 97.6%. The mean is dragged down by exactly 3 transition frames + the tail frame, where librosa's per-octave group delay differs from our uniform centring — steady state agrees to 0.9999, and for 10 s BTC segments that latency is immaterial. Frame count differs by one at the tail; align on the leading edge. CometBeat can reusetools/cqt_librosa_parity.pyas their Dart CQT oracle — it is the exact parity harness their handover asks them to budget for.⚠️ Cost is O(n_bins · N_k)/frame; N_k ≈ 24k samples at the bottom bin. Fine offline at hop 2048, NOT a real-time path — add a sparse spectral-kernel variant behind a flag if one is ever needed. -
core/gru.h— uni/bidirectional GRU. DONE, validated against a realtorch.nn.GRUfixture: max_abs < 1e-5 (tools/gen_gru_reference.py→CRISPASR_GRU_REF→tests/test-core-gru.cpp; the parity case skips without the fixture so CI needs no torch/network). Two traps documented inline because both yield plausible-but-wrong output: (1) the reset gate multiplies the RECURRENT TERMr*(W_hn h + b_hn), notW_hn(r*h)as the textbook GRU does; (2)b_ih/b_hhCANNOT be pre-summed the way core/lstm.h folds them, sinceb_hnsits inside therproduct. (piano_transcription landed with its own BiGRU; fold it in when convenient, do NOT refactor it blind.) -
core/stft.h— forward STFT.core/istft.hcovers only the inverse. htdemucs and mel-band-roformer each carry a private copy already, so a third consumer makes it the third duplicate.⚠️ This refactors two SHIPPED backends — requires byte-identical stem output on both before it lands, or it ships gated. Lower priority than cqt/gru precisely because of that risk.
- RMVPE (MIT, lj1995/VoiceConversionWebUI). Robust vocal F0; same
360-bin salience as CREPE, so it reuses
crepe_decode_local_averageunchanged and drops into the same--pitchsurface. CometBeat's vet correction is important and correct: the shipped ONNX is a pure conv U-Net with NO recurrent layer (the paper's GRU is absent), so it ports cleanly. Input is a 128-bin mel — nearly free for us viacore/mel.h. 361 MB f32 → ~180 MB q8_0 / ~90 MB q4_k. Wire as a second capacity behind--pitch. - BTC chord recognition — SHIPPED 2026-07-20.
--chords+ session C ABI- wasm;
crispasr-diff btc13/13 at cos 1.000000 (f16 and f32); 98.6-99.2 %mir_evalagreement with the torch reference on real music. Weights are CC-BY-NC-SA and ship behind--accept-license; published ascstr/btc-chords-GGUF.core/cqt.hlanded and was NOT the last blocker -- the real bugs were a missingscale=Trueand a chunked-vs-continuous front-end mismatch. Seedocs/music-transcription/PLAN.md.
- wasm;
- Basic Pitch — see §250. DONE 2026-08-30. Note: it needed a SECOND CQT
(
src/core/cqt2010v2.h, nnAudio CQT2010v2) —core/cqt.hstays the librosa direct-kernel one for BTC/TabCNN. - MT3 — feasibility memo on T5X/JAX checkpoint conversion BEFORE any C++.
- Bind
crispasr_session_separate*in the Dart package — SHIPPED (flutter/crispasr/lib/src/crispasr.dart:4031,separate()returningList<Stem>). NOTE: until 4eccc60cb this reached htdemucs ONLY -- the session ABI had no mel-band-roformer arm, so the MIT vocals/instrumental model (the CLI default) was unreachable from Dart. Now both. - CREPE perf: per-node profile (never done — the RTF numbers came from
FLOP arithmetic, not measurement) and sweep
kBatch(64 was a guess). - CUDA validation for the CREPE path and both ggml conv fixes
(
ggml_conv_1dbatch reshape,ggml_conv_1d_dwbatch support). Everything so far is M1/Metal only, and LEARNING 35 says Metal-correct is not CUDA-correct. - File the upstream ggml PR drafted at
tools/upstream-prs/24-conv-1d-batch-reshape.md.
Background (see LEARNINGS "Backend miscomputes my pipeline ≠ op X is broken" +
project_304_cosyvoice3_vulkan_native_lm_vs_flow): CV3 is CPU-routed under Vulkan
because native synthesis is noise. This was FULLY root-caused, on a real Tesla
P100: it is not a broken ggml op (test-backend-ops passes every flow/HiFT op
on Vulkan — IM2COL, NORM, MUL_MAT, ROPE, …) and not my gallocr dispatch
(FORCE_GALLOCR on CPU is bit-identical to the scheduler). It is aggregate
precision sensitivity: Vulkan's in-tolerance per-op accumulation deltas (fp16 vs
CPU fp32) compound across the 22-layer DiT × 6 CFM Euler steps and are amplified
by the log-mel HiFT vocoder → flow mel cosine(cpu,vk)=0.961 → garbage. The LM
(shallow-per-token) is bit-correct on Vulkan (512/512 greedy tokens).
- Lever (low priority, low odds): force fp32 accumulation in the
ggml-vulkan matmul/conv path for the flow + HiFT (e.g.
GGML_PREC_F32/ coopmat-f32-accum, or a build/env knob) and re-measure the flow mel cosine and the audio ASR round-trip on a real NVIDIA Vulkan device (Kaggle P100). If cosine climbs to ≳0.999 and audio is intelligible, native Vulkan (or at least a flow+HiFT-on-Vulkan hybrid) becomes viable. Why it's likely a dead end: every op already passes test-backend-ops within tolerance, so fp32-accum may not shrink the per-op delta enough to stop the CFM+vocoder amplification; and it is deep upstream ggml-vulkan shader work. Do NOT invest unless someone independently wants native-Vulkan CV3 speed badly enough to accept the shader effort. - Groundwork already committed (branch
fix/304-cosyvoice3-se, gated, NOT merged): LM single-backend gallocr (proves the LM runs native-Vulkan), hybrid HiFT-on-CPU infra, and diagnosticsCRISPASR_COSYVOICE3_{GREEDY,DUMP_MEL, DUMP_HIFT,FORCE_GALLOCR,HIFT_ON_GPU}. Repro kernels:tools/kaggle/cv3-vulkan-{isolate,convtest}. A real self-recursion crash in thecv3_sched_*shims (CPU/Metal SIGSEGV) was fixed on that branch; the shipped release was never affected (shims are branch-only). - Default stays the shipped all-CPU route under Vulkan — correct, and the right answer unless the lever above ever pays off.