Self-hosted German pronunciation assessment, with optional cloud services. Record a sentence and get a pronunciation score, the phoneme error rate, and each mispronounced word with expected vs. heard IPA. Includes a web app, a CLI and an HTTP API.
- Pronunciation score (0–100) and phoneme error rate for any German sentence
- Mispronunciation detection per word: expected IPA, the IPA actually heard, and a confidence
- Speech-to-text to show what the speaker said: Qwen3-ASR locally, or any OpenAI-compatible cloud API
- Reference audio (optional): hear the sentence or a single word, at normal or slow speed, via Chatterbox TTS Server
- Runs fully locally on NVIDIA GPU, Apple Silicon or CPU; cloud services are opt-in
-
Install the system packages:
# Debian / Ubuntu sudo apt install ffmpeg espeak-ng build-essential python3-dev # macOS brew install ffmpeg espeak-ng
-
Clone and start:
git clone https://github.com/RefNull/phonoscore-german.git cd phonoscore-german python3 serve.pyOn first run,
serve.pycreates.venvand installsrequirements.txt. -
When the log says
Ready, open http://localhost:8002.
First run: the first request to each service downloads the models it uses from Hugging Face, so expect a wait:
- The ASR service downloads the chosen ASR model (default Qwen3-ASR-1.7B), unless you use a cloud endpoint.
- The scoring service downloads OpenPronounce's Wav2Vec2 models and the Piper German voice.
Later starts load them from the local cache. See Models for the options.
Open http://<server>:8002. Pick or type a sentence, press Record, read it aloud, then press Score it. You can also upload a recording.
Browsers only allow the microphone on https:// or localhost. For a server on your network, either upload files or tunnel the port: ssh -L 8002:localhost:8002 <server>.
.venv/bin/python cli.py --text "Ich habe morgen einen Termin beim Arzt." # record from the microphone
.venv/bin/python cli.py --text "Guten Tag." --audio-file sample.wav # score a file| Flag | Default | |
|---|---|---|
--text |
Ich habe morgen einen Termin beim Arzt. |
Sentence the speaker should say |
--audio-file |
.wav or .ogg to score instead of recording |
|
--asr-url |
http://localhost:8001 |
ASR service |
--pronounce-url |
http://localhost:8002 |
Scoring service |
--router |
off | Send requests through llama-swap instead |
--router-url |
http://localhost:8080 |
llama-swap URL |
--samplerate |
16000 |
Microphone sample rate |
To record from the microphone, the CLI needs PortAudio (sudo apt install libportaudio2 or brew install portaudio).
# Score pronunciation
curl -F file=@recording.wav -F expected_text="Ich habe morgen einen Termin beim Arzt." \
http://localhost:8002/assess
# Transcribe (OpenAI-compatible)
curl -F file=@recording.wav -F language=de http://localhost:8001/v1/audio/transcriptions| Endpoint | Service | Form fields |
|---|---|---|
POST /assess |
scoring, port 8002 | file (.wav/.ogg), expected_text, lang (default de) |
POST /v1/audio/transcriptions |
ASR, port 8001 | file, language (default de), model (ignored) |
GET /health |
both | returns status, service, device (and model for ASR) |
GET / |
scoring, port 8002 | the web app |
/assess response:
{
"score": 82.5,
"transcription": "ich habe morgen einen termin beim arzt",
"phoneme_error_rate": 0.175,
"errors": [
{ "word": "termin", "expected_ipa": "tɛʁˈmiːn", "actual_ipa": "tɛʁˈmɪn", "confidence": 0.65 }
]
}| Field | |
|---|---|
score |
Pronunciation score, 0–100 |
transcription |
What the Wav2Vec2 acoustic model heard |
phoneme_error_rate |
Phoneme error rate; 0 is perfect |
errors[] |
Mispronounced words: word, expected_ipa, actual_ipa, confidence (0–1; 0 when OpenPronounce's phone recognizer is off) |
| Flag | Default | |
|---|---|---|
--model |
Qwen3-ASR-1.7B |
ASR model, see Models |
--asr-endpoint |
Cloud ASR API instead of a local model, see Cloud services | |
--tts |
piper |
Reference voice for scoring: piper (local) or gtts (Google) |
--device |
auto |
auto, cuda, mps or cpu |
--host |
0.0.0.0 |
Bind address |
--asr-port / --pronounce-port |
8001 / 8002 |
Service ports |
-d, --detach |
Run in the background | |
--status / --stop |
Check or stop a background run | |
-f, --logs |
Follow the log (logs/serve.log; per service in logs/asr.log, logs/pronounce.log) |
|
--python / --no-auto-setup |
Use a specific interpreter / don't create .venv |
You can also run the services separately: python asr_server.py --port 8001 and python pronounce_server.py --port 8002. Both take --host and --device; asr_server.py also takes --model and --endpoint, pronounce_server.py takes --tts.
To load models on demand, run llama-swap from this directory with .venv activated, using the included llama-swap.yaml:
llama-swap --config llama-swap.yaml --port 8080
.venv/bin/python cli.py --routerIn the web app, choose llama-swap under Settings. See the llama-swap README for how it works.
Models download from Hugging Face on first use into the Hugging Face cache (~/.cache/huggingface, or wherever HF_HOME points). For offline use, run once to fill the cache, then set HF_HUB_OFFLINE=1.
These are the built-in models (recommendations as of 29 Sep 2026). Choose one with --model or PHONOSCORE_ASR_MODEL.
--model |
Hugging Face repo | Comment |
|---|---|---|
Qwen3-ASR-1.7B (default) |
Qwen/Qwen3-ASR-1.7B-hf |
Default, based on testing |
Qwen3-ASR-0.6B |
Qwen/Qwen3-ASR-0.6B-hf |
Smaller and faster than 1.7B, but lower quality |
| any other value | a repo ID or local path | Any other Qwen3-ASR checkpoint, such as a fine-tune. For other models (Whisper and others), use --asr-endpoint. |
OpenPronounce uses three Wav2Vec2 models:
| Hugging Face repo | Used for |
|---|---|
facebook/wav2vec2-lv-60-espeak-cv-ft |
Recognizing the phones actually spoken |
jonatasgrosman/wav2vec2-large-xlsr-53-german |
German word transcription |
facebook/wav2vec2-large-960h |
Acoustic comparison with a reference recording |
It synthesizes that reference recording from the expected sentence. Choose the voice with --tts:
--tts |
Runs | Comment |
|---|---|---|
piper (default) |
Locally | Piper voice de_DE-thorsten-medium, downloaded on first use |
gtts |
Cloud | Google Translate TTS. Sends each sentence to Google; no download. |
PhonoScore runs locally by default. Each part can use a cloud service instead:
| Part | Option | |
|---|---|---|
| Speech-to-text | --asr-endpoint <url> |
Any OpenAI-compatible transcription API. No local ASR model is downloaded. |
| Reference voice | --tts gtts |
Google Translate TTS |
| Listen buttons in the web app | Settings → TTS endpoint | Any Chatterbox TTS Server, local or remote |
For --asr-endpoint, pass the API's base URL and set --model to the provider's model name. The API key goes in PHONOSCORE_ASR_API_KEY and stays on the server. For example:
export PHONOSCORE_ASR_API_KEY=sk-...
python3 serve.py --asr-endpoint https://api.openai.com/v1 --model whisper-1
python3 serve.py --asr-endpoint https://api.groq.com/openai/v1 --model whisper-large-v3A self-hosted server with the same API works too, such as a faster-whisper server on another machine.
- Scoring: OpenPronounce turns the expected sentence into German IPA with espeak-ng and recognizes the phones actually spoken with Wav2Vec2. It aligns the two and flags substitutions, deletions and insertions.
- Transcription: Qwen3-ASR, or the cloud endpoint you configured, transcribes the same recording, so you see what was said.
- The two services are independent. The client calls both.
PhonoScore supports German only for now. These parts are German-specific:
- The default language
deinasr_server.py,pronounce_server.py,cli.pyandclient.html - The OpenPronounce models and espeak-ng voice chosen for
de - The sample sentences and German TTS voice in
client.html, and the default sentence incli.py
pip install -r requirements-dev.txt
pytestThe tests run offline with mocks. No GPU or model download is needed.
MIT. One dependency, piper-tts (the local reference voice), is GPL-3.0. For all third-party software and models, see THIRD_PARTY_LICENSES.md.
