This repository contains WebSocket client samples for the Voxist ASR (Automatic Speech Recognition) service in both JavaScript and Python.
Go there.
Contact us if interested.
The scripts support both staging and production environments:
- Production:
api-asr.voxist.com(default) - Staging:
asr-staging-dev.voxist.com(with--stagingflag)
Important: Staging and production environments use different API keys. Make sure to use the correct API key for your target environment.
There are two ways to connect to the WebSocket:
wss://api-asr.voxist.com/ws?api_key=YOUR_API_KEY&lang=fr-medical&sample_rate=16000
- Request a temporary token:
curl -X 'GET' \
'https://api-asr.voxist.com/websocket?engine=voxist-rt-2' \
-H 'accept: application/json' \
-H 'X-LVL-KEY: YOUR_API_KEY'- Response:
{
"url": "wss://api-asr.voxist.com/ws?token=JWT_TOKEN"
}- Add parameters to the URL:
wss://api-asr.voxist.com/ws?token=JWT_TOKEN&lang=fr-medical&sample_rate=16000
Prefer this method for browsers and any client you do not fully control: the token is short-lived (1 hour) and does not expose your long-lived API key.
| Parameter | Required | Description |
|---|---|---|
api_key or token |
yes | Authentication. Exactly one of the two. |
lang |
no | Language / model. If omitted, the connection waits for a config message before it accepts audio — any audio sent before then is discarded. |
sample_rate |
no | Defaults to 16000. See the note under Audio Format. |
punctuation_mode |
no | Generated (default) or Dictated. See Punctuation modes. |
Send raw audio data directly to the WebSocket:
- Format: Raw PCM audio bytes — not a WAV file. Locate the
datachunk and send its contents only; the server forwards every binary frame straight to the decoder, so header bytes on the wire are decoded as if they were audio. Do not assume a fixed 44-byte header: aLIST/INFOchunk (what ffmpeg writes by default, and what most DAW exports carry) pushesdatafurther in, and skipping a fixed 44 bytes then misaligns every sample.wav.js/wav.pyin this repo walk the chunk list for exactly this reason. - Encoding: Signed 16-bit little-endian
- Channels: Mono (1 channel)
- Sample Rate: 16000 Hz
- Chunk Size: Recommended 100ms chunks (3200 bytes at 16kHz)
The server does not resample.
sample_ratein the connection URL is used for duration accounting, not for conversion, and the streaming engines are 16 kHz models. Sending 8 kHz audio does not fail — it returns a confidently wrong transcript. Convert before you stream:ffmpeg -i input.wav -ac 1 -ar 16000 -sample_fmt s16 output.wav
For optimal real-time performance:
- Timing: Send approximately 1 second of audio per second
- Chunk Interval: 100ms chunks sent every 100ms
- Buffer Management: Avoid buffering large amounts of audio
- Network Latency: Account for network delays in your timing
Example timing for 16kHz audio:
const SAMPLE_RATE = 16000;
const BYTES_PER_SAMPLE = 2;
const CHUNK_DURATION_MS = 100;
const CHUNK_SIZE = SAMPLE_RATE * BYTES_PER_SAMPLE * (CHUNK_DURATION_MS / 1000); // 3200 bytes
// Send chunk every 100ms
setInterval(() => {
const audioChunk = getAudioChunk(CHUNK_SIZE);
websocket.send(audioChunk);
}, CHUNK_DURATION_MS);Instead of (or in addition to) URL parameters, send a JSON text frame:
{
"config": {
"lang": "fr-medical",
"sample_rate": 16000,
"punctuation_mode": "Generated"
}
}This is also how you change settings on a live connection. Changing lang
transparently reconnects the session to the new engine — you do not need to open
a new socket. If you connected without a lang URL parameter, this message is
what unblocks audio processing.
Transcripts are post-processed server-side before they reach you. Two modes are available, on every language:
Generated(default) — the pipeline inserts punctuation and casing for you.Dictated— spoken punctuation commands ("point", "virgule", "à la ligne") are converted into the corresponding marks instead of being transcribed as words. This is the mode to use for dictation workflows.
Set it in the URL (&punctuation_mode=Dictated) or in a config message, and
change it mid-session with another config message:
{ "config": { "punctuation_mode": "Dictated" } }The same pipeline applies number formatting, and on the medical models, unit normalization.
Per-session term replacements are applied at the end of the text pipeline — use them for names, local jargon, or product terms the model spells differently:
{
"config": {
"user_vocabulary": [
{ "pattern": "petite soeur", "replacement": "MySys", "case_sensitive": false }
]
}
}- Literal matching only — a
patternis not a regular expression, and entries flaggedis_regexare dropped. - Up to 100 entries;
patternandreplacementare capped at 256 bytes each. - Re-send at any point to replace the whole set; send an empty array to clear it.
- Invalid entries are dropped individually rather than voiding the whole set, and dropping is silent — there is no per-entry error response.
To signal the end of audio and flush the final result, send the text frame:
Done
That is the literal four-byte string Done, sent as a text (not binary) frame:
websocket.send('Done');The server forwards it to the engine, which drains its buffer, emits the last
final message, and then closes the connection — so wait for the close
event rather than closing the socket yourself. Hanging up early truncates the
tail of your transcript.
Bound that wait. The gateway only forwards Done if its upstream engine socket
is already open; if the engine is still connecting — a cold start, or a clip
short enough to finish first — the flush is dropped and no close ever arrives.
The samples wait 30 seconds, then close themselves and report the missing tail
rather than hanging.
Breaking change (July 2026 samples update). Earlier versions of these samples sent
{"eof": 1}. That message is not part of the protocol: the server logs it as an unknown text frame and discards it. The result is a transcript missing its final segment and a socket that stays open (and metered) until the client gives up. If you copied that pattern, replace it withDone.
The WebSocket returns JSON messages with transcription results. Both partial and final results have the same format, only the type field differs:
{
"text": " Ceci est un te",
"transcript": " Ceci est un te",
"type": "partial",
"startedAt": 0,
"segment": 0,
"elements": {
"segments": [
{
"text": " Ceci est un te",
"type": "segment",
"startedAt": 0,
"segment": 0
}
],
"words": [
{
"text": "Ceci",
"type": "word",
"startedAt": 1.28,
"segment": 0
},
{
"text": "est",
"type": "word",
"startedAt": 1.8,
"segment": 0
},
{
"text": "un",
"type": "word",
"startedAt": 2.04,
"segment": 0
},
{
"text": "te",
"type": "word",
"startedAt": 2.32,
"segment": 0
}
]
}
}{
"text": " Ceci est un test",
"transcript": " Ceci est un test",
"type": "final",
"startedAt": 0,
"segment": 0,
"elements": {
"segments": [
{
"text": " Ceci est un test",
"type": "segment",
"startedAt": 0,
"segment": 0
}
],
"words": [
{
"text": "Ceci",
"type": "word",
"startedAt": 1.28,
"segment": 0
},
{
"text": "est",
"type": "word",
"startedAt": 1.8,
"segment": 0
},
{
"text": "un",
"type": "word",
"startedAt": 2.04,
"segment": 0
},
{
"text": "test",
"type": "word",
"startedAt": 2.32,
"segment": 0
}
]
}
}text: The transcribed texttranscript: Best-effort mirror oftext, added by the server-side text pipeline. It is absent when that pipeline passes a message through untouched (empty text, or an internal processing error), so readtextand treattranscriptas optionaltype:"partial"for real-time updates,"final"for completed segmentsstartedAt: Start time of the segment in secondssegment: Segment number (increments for each completed phrase/sentence)elements: Detailed breakdown with word-level timingsegments: Array of text segments with timingwords: Array of individual words with precise timestamps
Note: Read text, not transcript — see the field list above. The only
difference between partial and final results is the type field. Partial results may have incomplete words (e.g., "te" instead of "test"), while final results contain the complete, corrected transcription.
Word-level timings inside elements are produced by the acoustic model and are
not re-aligned after text post-processing, so on medical models the text may
be normalized ("15 mg") where the corresponding words entries still carry the
spoken tokens.
- Connect to WebSocket with API key or token
- (optional) Send a
configmessage if you did not passlangin the URL - Stream audio in real-time chunks (100ms recommended)
- Receive partial results for immediate feedback
- Receive final results for completed segments with detailed timing
- Send
Donewhen finished - Wait for the server to close the connection
Failures arrive as WebSocket close codes, not as JSON error messages:
| Code | Meaning | What to do |
|---|---|---|
1008 |
Unsupported language code, a model your account is not entitled to, or a per-tenant rate limit | Check lang against the supported list; back off if you are sending many sessions |
1011 |
Engine unavailable, engine timeout, or a language with no engine configured in this environment | Retry; escalate if persistent |
1013 |
Server at capacity. The close reason carries {"error":"server_overloaded","retryAfterMs":3000} |
Back off and retry after the advertised delay |
1006 (no handshake) |
Rejected during the HTTP upgrade — bad or missing credentials | Check the API key / token and the target environment |
All three scripts exit non-zero and print the code and reason on an abnormal close, so a wrapper script can tell a capacity rejection from a clean run.
Other things to watch for:
- Audio sent before the connection is configured is dropped. Pass
langin the URL, or wait after sending yourconfigmessage. - Token expiry: temporary tokens are valid for 1 hour. Long sessions should reconnect with a fresh token.
- Audio format errors do not raise an error — see the resampling note above.
The microphone script (asr-mic.js) requires SoX to be installed and available in your $PATH.
sudo apt-get install sox libsox-fmt-allbrew install soxNote: SoX is only required for the microphone script (asr-mic.js). The file-based scripts (asr-file-ws.js and asr-file-ws.py) do not require SoX.
Requires Node.js 18 or later (the microphone script uses the global fetch).
npm installDirect WebSocket connection using API key authentication with CLI parameters:
node asr-file-ws.js <API_KEY> <WAV_FILE> [LANG] [--punctuation-mode=MODE] [--staging]Examples:
# Production environment (default)
node asr-file-ws.js your-prod-api-key audio.wav fr-medical
# Staging environment
node asr-file-ws.js your-staging-api-key audio.wav fr-medical --staging
# English transcription in production
node asr-file-ws.js your-prod-api-key audio.wav en
# Dictation mode: spoken punctuation becomes real punctuation
node asr-file-ws.js your-prod-api-key audio.wav fr-medical --punctuation-mode=DictatedParameters:
API_KEY: Your Voxist API key (different for staging and production)WAV_FILE: Path to the WAV audio file (16 kHz mono 16-bit PCM; the script validates this and strips the WAV header before streaming)LANG: Language code (optional, default:fr)--punctuation-mode:Generated(default) orDictated— see Punctuation modes--staging: Use staging environment (optional)
Real-time microphone transcription using WebSocket with temporary token authentication:
node asr-mic.js <API_KEY> [LANG] [--punctuation-mode=MODE] [--staging]Examples:
# Production environment (default)
node asr-mic.js your-prod-api-key fr-medical
# Staging environment
node asr-mic.js your-staging-api-key fr-medical --staging
# English transcription in production
node asr-mic.js your-prod-api-key enParameters:
API_KEY: Your Voxist API key (different for staging and production)LANG: Language code (optional, default:fr)--punctuation-mode:Generated(default) orDictated— see Punctuation modes--staging: Use staging environment (optional)
Features:
- Real-time microphone recording and transcription
- Records headerless mono 16-bit PCM at 16 kHz, ready to stream as-is
- Temporary token authentication (more secure than direct API key in WebSocket)
- Live partial results with
[LIVE]prefix - Final results with
[FINAL]prefix - Ctrl+C stops the microphone, flushes with
Done, and waits for the last final
Requirements:
- SoX must be installed and available in PATH
- Working microphone
- Microphone permissions granted to terminal/application
Troubleshooting a silent microphone
node-audiorecorder invokes SoX with -V0, which suppresses SoX's own error
output, so a device that cannot be opened looks exactly like a device that is
simply quiet. On macOS in particular, a denied microphone permission produces
no bytes at all and no error. The client therefore detects this itself:
| Symptom | What the client does |
|---|---|
| No audio within 20 s of starting | Prints the likely causes with platform-specific steps, and exits 1 |
| SoX binary missing | Prints install instructions for your platform, and exits 1 |
| SoX exits before delivering audio | Reports the exit code and the same guidance, and exits 1 |
| Device delivers all-zero samples | Warns that the transcription will be empty, and continues |
The first run on macOS raises a permission prompt. Granting it does not
retroactively deliver audio to the SoX process that is already running, so
answer the prompt and then start the client again. If it was dismissed, grant
access under System Settings > Privacy & Security > Microphone for your
terminal application first. To see what SoX itself is complaining about, run the
capture without -V0:
sox -d -q -c 1 -r 16000 -t raw -L -b 16 -e signed-integer - > /tmp/mic-test.rawA working microphone produces 32000 bytes per second.
How it works:
- Requests a temporary WebSocket token from the API using your API key
- Adds language and sample rate parameters to the WebSocket URL
- Connects to the WebSocket using the temporary token
- Streams microphone audio in real-time
Run the setup script to create a virtual environment and install dependencies:
./setup-python.sh- Create a virtual environment:
python3 -m venv venv- Activate the virtual environment:
source venv/bin/activate # On Linux/Mac
# or
venv\Scripts\activate # On Windows- Install dependencies:
pip install -r requirements.txtpython asr-file-ws.py <API_KEY> <WAV_FILE> [LANG] [--punctuation-mode=MODE] [--staging]Examples:
# Production environment (default)
python asr-file-ws.py your-prod-api-key audio.wav fr-medical
# Staging environment
python asr-file-ws.py your-staging-api-key audio.wav fr-medical --staging
# English transcription in production
python asr-file-ws.py your-prod-api-key audio.wav en
# Dictation mode: spoken punctuation becomes real punctuation
python asr-file-ws.py your-prod-api-key audio.wav fr-medical --punctuation-mode=DictatedParameters:
API_KEY: Your Voxist API key (different for staging and production)WAV_FILE: Path to the WAV audio file (16 kHz mono 16-bit PCM; the script validates this and streams thedatachunk only)LANG: Language code (optional, default:fr)--punctuation-mode:Generated(default) orDictated— see Punctuation modes--staging: Use staging environment (optional)
All three scripts honour two optional environment variables, for self-hosted deployments and local testing:
| Variable | Overrides |
|---|---|
VOXIST_ASR_URL |
WebSocket base URL (default wss://api-asr.voxist.com) |
VOXIST_ASR_API_URL |
HTTP base URL used for the token request in asr-mic.js |
VOXIST_ASR_URL=ws://127.0.0.1:3000 node asr-file-ws.js your-api-key audio.wav frReal-time streaming models available in production:
| Code | Language |
|---|---|
fr |
French |
fr-medical |
French Medical |
fr-medicalV2-16 |
French Medical, 16-frame latency tier |
fr-medicalV2-32 |
French Medical, 32-frame latency tier (same model as fr-medical) |
fr-medicalV2-64 |
French Medical, 64-frame latency tier |
en |
English |
de |
German |
es |
Spanish |
it |
Italian |
nl |
Dutch |
pt |
Portuguese |
Region-qualified aliases (fr-FR, en-US, de-DE, nl-NL) are accepted and
collapse to their base language.
Not available for streaming: sv, pl, ja, he, tr. These codes are
recognized by the API but have no streaming engine deployed in production today
— a connection using them is closed with code 1011. Swedish is available for
offline (file upload) transcription via the REST API. Previous versions of
this README listed sv as a supported streaming language; that was incorrect.
- Format: WAV (the script streams the PCM payload, not the container)
- A file whose header declares more audio than it contains is transcribed as far as it goes, with a warning and a non-zero exit, so a partial result is never mistaken for a complete one
- Sample Rate: 16000 Hz — the server does not resample
- Channels: Mono (1 channel)
- Bit Depth: 16-bit
- Automatically configured to headerless mono 16-bit PCM at 16 kHz
- SoX handles audio capture and format conversion
- Works with any microphone supported by the system
Important: You need different API keys for staging and production environments:
- Production API Keys: Used with
api-asr.voxist.com(default behavior) - Staging API Keys: Used with
asr-staging-dev.voxist.com(with--stagingflag)
Contact Voxist support to obtain API keys for both environments.