Default model-manager downloads use the published GGUF package when available; the original source/conversion instructions below remain valid for manual use.
kroko_asr is a native audio.cpp port of the free Kroko Community
Zipformer2/RNN-T models. The Kaldi-compatible filterbank, streaming
Conv2dSubsampling/ConvNeXt encoder, 19-layer Zipformer2, stateless predictor,
joiner, greedy decoder, and modified beam decoder run without ONNX Runtime.
| Field | Value |
|---|---|
| Task | asr |
| Modes | offline, native stateful streaming |
| Public free languages | German (de), English (en), Spanish (es), French (fr), Italian (it), Hebrew (he; package code IW), Dutch (nl), Portuguese (pt), Swedish (sv), Turkish (tr) |
| Input | WAV; audio.cpp converts to 16 kHz mono |
| Output | Transcript and word timestamps |
| Decoding | Greedy search; modified beam search; blank penalty; inline hotwords |
| Endpointing | Optional three-rule automatic segmentation |
| Package variants | 64-L and 128-L streaming packages |
| Native layouts | Converted safetensors and standalone GGUF |
Each Kroko package recognizes one language. Select a package whose language
matches --language; auto uses the package language. The loader normalizes
the legacy Hebrew code iw to he.
Streaming keeps the subsampling cache, Zipformer layer states, RNN-T predictor context, emitted tokens, and emission frames across chunks. Already consumed waveform is compacted while retaining the filterbank boundary overlap. Non-16-kHz input is incrementally resampled while retaining only the two source samples needed across chunk boundaries, so both buffers remain bounded. Finalization adds the same 660 ms zero tail as the original Kroko/sherpa runner to flush final punctuation and tokens.
Word starts come from the RNN-T encoder frame where the first token of the word
is emitted. One encoder frame is 40 ms (160 filterbank-hop samples times
subsampling factor 4). --words-out writes audio.cpp sample spans.
Free packages are published in
Banafo/Kroko-ASR. Their .data
container holds a JSON header, quantized encoder/decoder/joiner ONNX graphs,
and tokens.txt. Commercial/encrypted packages are intentionally rejected;
use Kroko's licensed runtime for those models.
Install converter dependencies:
python -m pip install numpy onnx safetensorsThe model manager infers a language/size-specific target such as
Kroko-DE-Community-64-L-Native, so different languages do not overwrite one
another:
python .\tools\model_manager_deprecated.py install kroko_asr_community_converted `
--source-file .\models\Kroko-ASR\Kroko-DE-Community-64-L-Streaming-001.data `
--models-root .\models\Kroko-ASR `
--overwriteUse --variant <directory-name> to override the inferred target directory.
The converter can also be called directly:
python .\tools\community_models\convert_kroko_onnx.py `
.\models\Kroko-ASR\Kroko-SV-Community-64-L-Streaming-001.data `
.\models\Kroko-ASR\Kroko-SV-Community-64-L-Native `
--overwriteThe result contains:
Kroko-SV-Community-64-L-Native/
|-- config.json
|-- model.safetensors
`-- tokens.txt
Kroko is v1-native. model_specs/kroko_asr.json is the single source of truth
for metadata, capabilities, normalized options, package installation, and the
GGUF/safetensors resource layout. The generic spec-backed loader derives model
inspection and CLI/help metadata from that contract.
The converter supports both public chunk layouts (141/128 feature frames for
64-L and 269/256 for 128-L). It dequantizes MatMulInteger tensors, recovers
folded Zipformer/downsampling constants and both exported forms of chunk-edge
scales, ignores k2 disambiguation symbols beyond the joiner vocabulary, and
writes semantic audio.cpp tensor names.
Offline safetensors transcription with word timestamps:
.\build\windows-cuda-release\bin\audiocpp_cli.exe `
--task asr --family kroko_asr `
--model .\models\Kroko-ASR\Kroko-EN-Community-128-L-Native `
--backend cpu --audio .\SAMPLES\EN_3.wav --language en `
--text-out .\outputs\kroko_en.txt `
--words-out .\outputs\kroko_en_words.json --logNative streaming:
.\build\windows-cuda-release\bin\audiocpp_cli.exe `
--task asr --mode streaming --family kroko_asr `
--model .\models\Kroko-ASR\Kroko-SV-Community-64-L-Native `
--backend cpu --audio .\speech_sv.wav --language sv `
--text-out .\outputs\kroko_sv.txt `
--words-out .\outputs\kroko_sv_words.json --logModified beam search with a blank penalty and natural-text hotwords:
.\build\windows-cuda-release\bin\audiocpp_cli.exe `
--task asr --family kroko_asr `
--model .\models\Kroko-ASR\Kroko-EN-Community-128-L-Native `
--backend cpu --threads 8 --audio .\SAMPLES\EN_3.wav --language en `
--request-option decoding_method=modified_beam_search `
--request-option num_beams=8 `
--request-option blank_penalty=0.5 `
--request-option "hotwords=security/tomorrow" `
--request-option hotwords_score=1.5Hotword phrases are separated with / or newlines. audio.cpp tokenizes the
natural text directly from the package vocabulary; no SentencePiece model is
required.
Automatic endpoint segmentation is opt-in and uses the same three default rules as sherpa-onnx:
.\build\windows-cuda-release\bin\audiocpp_cli.exe `
--task asr --mode streaming --family kroko_asr `
--model .\models\Kroko-ASR\Kroko-EN-Community-128-L-Native `
--backend cpu --audio .\speech.wav --language en `
--request-option enable_endpoint=true `
--segments-out .\outputs\kroko_segments.json| Request option | Default | Meaning |
|---|---|---|
decoding_method |
greedy_search |
greedy_search or modified_beam_search |
num_beams |
4 |
Beam hypotheses, from 1 through 64 |
blank_penalty |
0 |
Non-negative score subtracted from the blank logit |
hotwords |
empty | Slash- or newline-separated natural-text phrases |
hotwords_score |
1.5 |
Non-negative context boost per hotword token |
enable_endpoint |
false |
Enable automatic speech-segment boundaries |
rule1_min_trailing_silence_sec |
2.4 |
Endpoint timeout even without decoded speech |
rule2_min_trailing_silence_sec |
1.2 |
Endpoint silence after decoded speech |
rule3_min_utterance_length_sec |
20 |
Maximum utterance duration before an endpoint |
Request keys use the normalized v1 names directly. Hotwords require modified beam search. Greedy remains the default.
.\build\windows-cuda-release\bin\audiocpp_gguf.exe `
--input .\models\Kroko-ASR\Kroko-SV-Community-64-L-Native\model.safetensors `
--root .\models\Kroko-ASR\Kroko-SV-Community-64-L-Native `
--family kroko_asr --type q8_0 `
--output .\models\Kroko-ASR\Kroko-SV-Community-64-L-Native\Kroko-SV-Community-64-L-Q8.gguf `
--overwriteThe GGUF embeds config.json, tokens.txt, and the kroko_asr package spec.
It can therefore be moved or renamed and passed directly to --model.
.\build\windows-cuda-release\bin\audiocpp_cli.exe `
--task asr --mode streaming --family kroko_asr `
--model .\models\Kroko-ASR\Kroko-SV-Community-64-L-Native\Kroko-SV-Community-64-L-Q8.gguf `
--backend cuda --audio .\speech_sv.wav --language sv `
--words-out .\outputs\kroko_sv_q8_words.jsonConfigure either mode. This example exposes streaming SSE:
{
"host": "127.0.0.1",
"port": 8080,
"backend": "cuda",
"models": [
{
"id": "kroko-sv-stream",
"family": "kroko_asr",
"path": "models/Kroko-SV-Community-64-L-Q8.gguf",
"task": "asr",
"mode": "streaming"
}
]
}.\build\windows-cuda-release\bin\audiocpp_server.exe --config .\server.json --log
curl.exe -N http://127.0.0.1:8080/v1/audio/transcriptions `
-F "file=@speech_sv.wav" -F "model=kroko-sv-stream" `
-F "language=sv" -F "stream=true" -F "response_format=json"Request options are also forwarded by the JSON route:
$body = @{
model = "kroko-sv-stream"
audio_path = (Resolve-Path .\speech_sv.wav).Path
language = "sv"
options = @{
decoding_method = "modified_beam_search"
num_beams = 4
blank_penalty = 0.5
enable_endpoint = $true
}
} | ConvertTo-Json -Depth 4
Invoke-RestMethod -Method Post `
-Uri http://127.0.0.1:8080/v1/audio/transcriptions `
-ContentType application/json -Body $bodyThe generic /v1/tasks/run and /v1/tasks/stream result schemas carry the
model's word_timestamps. The OpenAI-compatible transcription route currently
returns its normal text/delta schema.
The complete reproducible commands, per-request multilingual ONNX comparison, 64-L and 128-L tensor-boundary parity, streaming/offline equality, standalone GGUF path test, server results, timings, and memory notes are in Kroko ASR validation.
- Only public packages with
free=trueare converted. Commercial/encrypted packages require Kroko's license and runtime. - Each package is single-language; audio.cpp does not implement Kroko's multi-model language router.
- The current CPU runtime is parity-focused and slower than the optimized ONNX Runtime reference; see the validation report.
- Vulkan and Metal require contributor testing.
Source references: