Media files often carry audio or subtitle tracks tagged und
(undetermined) — no language metadata. That breaks track cleanup: your
"keep English + Icelandic" rule can't match a track it can't identify, so
und tracks get dropped, or you have to hand-pick every track. Shrinkerr
detects the real language of und tracks and can write the correct tag
back into the file.
- What gets detected
- How to run it
- Writing tags back to the file
- Settings
- The "Unknown language" filter
- How detection works
- Confidence and fail-open
- Dependencies and image size
- Troubleshooting
| Track type | Method | Speed | Runs automatically? |
|---|---|---|---|
| Text subtitles (SRT/ASS/…) | Text analysis (langdetect) | milliseconds | Yes, at scan time |
External subtitles (sidecar .srt/.ass) |
Text analysis (langdetect) | milliseconds | On demand |
External VobSub (sidecar .idx/.sub) |
OCR (subtile-ocr) | seconds/track | On demand |
| Audio | Spoken-language ID (faster-whisper) | seconds/track | On demand; optionally before conversion |
| Image subtitles (PGS, VobSub) | OCR (tesseract) | minutes/track | On demand only |
Text-subtitle detection includes charset detection, so subtitles in legacy non-Unicode encodings — GB2312 / GBK / Big5 (Chinese), Shift-JIS (Japanese), EUC-KR (Korean), Windows-1252 / Latin-1 (Western) — are decoded correctly instead of turning into gibberish the detector can't read.
Image-subtitle OCR is multi-script: it detects Latin, Han (Chinese), Japanese, Korean, Cyrillic, and Arabic scripts and OCRs with the matching model.
Automatic (text subtitles): happens during a normal scan. An und text
subtitle gets its language filled in with no action from you.
On demand (any track): open a file's detail panel and click Detect
languages. This detects every und audio and subtitle track on that file
— including the slow image-sub OCR — applies the results, and (by default)
writes the tags into the file. Image-sub OCR takes minutes; a live status
line shows the current stage (extracting → OCR passes).
Batch: from the scanner list (e.g. filtered to Unknown language), use Detect all unknown, or select titles and use Detect selected (N). The batch runs server-side with a live N/total progress indicator that survives navigating away and back, and a Cancel button (stops after the current file).
Before conversion (audio): when the Auto-detect languages setting is
on (default), converting a file with und audio detects the language first,
so the converted output carries the correct tag. Image-sub OCR is not run
before conversion — it's too slow to gate an encode on.
By default, on-demand detection doesn't just update Shrinkerr's database — it writes the corrected language tags into the file, so the fix is real on disk and every other tool (Plex, Jellyfin, mpv, …) sees it too. No re-encode is involved:
- mkv — edited in place with
mkvpropedit: instant, no rewrite, no extra disk space, even on a 50 GB file. - mp4 / m4v / mov / ts — a fast
ffmpeg -c copyremux to a temp file, then an atomic replace. No quality loss (streams are copied, not re-encoded), but it does rewrite the container. The tag is verified present in the output before the replace. - containers with no per-track language field (avi, mpg, wmv, flv, …) — these cannot store a language tag, so the write is skipped and the track stays flagged as unknown. The only real fix is remuxing to mkv (i.e. convert the file); detection alone can't help here.
- external sidecar subs — the language lives in the filename, so the
sidecar is renamed to embed the ISO-639-2 code (
Movie.srt→Movie.eng.srt) — the convention Plex/Jellyfin/Bazarr read. For VobSub the.idxand.subare renamed together so the pair stays matched. If a same-named file already exists, the rename is skipped (no clobber).
Only tracks that were upgraded from und are touched; already-tagged tracks
are left alone. A track leaves the Unknown-language filter only once the tag
is verified in the file — if the write can't persist (unsupported container,
read-only mount, error), the track stays flagged so you never lose track of it
while the file and Plex still read und.
If Plex is configured, a library refresh is triggered afterward so Plex re-reads the corrected languages. A batch run coalesces this into one full refresh per affected library section at the end (see the setting below).
Both are in Settings → Video (shown in the audio and subtitle sections):
- Auto-detect languages for unknown tracks before converting
(default on) — when converting a file with
undtracks, detect audio language first so the output is tagged correctly. Turn off to leaveundtracks untouched through conversion. - Notify Plex on track-language change (default on) — after tags are written to a file, refresh the file's Plex library section. No effect if Plex isn't configured.
- Audio detection model (
tiny/base/small, defaulttiny) — the faster-whisper model for spoken-language ID. Larger is more accurate on non-English speech at the cost of download size, RAM, and speed. Takes effect on the next detection (no restart). - Audio / Subtitle confidence (defaults
0.6/0.7) — minimum confidence to accept a detected language. Lower tags more tracks but risks wrong tags.
The env vars below still work and override the Settings values (handy for headless/compose-only deploys):
| Env var | Default | Meaning |
|---|---|---|
SHRINKERR_LANG_DETECT_AUDIO_MIN |
0.6 |
Min confidence to accept an audio detection |
SHRINKERR_LANG_DETECT_SUB_MIN |
0.7 |
Min confidence to accept a subtitle detection |
SHRINKERR_WHISPER_CACHE |
/app/data/models |
Where the faster-whisper model is cached |
SHRINKERR_WHISPER_MODEL |
tiny |
faster-whisper model for audio ID. tiny is under-confident on non-English speech (Nordic tracks land just under the gate); set base or small for better accuracy at the cost of size/speed. |
SHRINKERR_PGS_SAMPLE_SECONDS |
1200 |
Seconds of a PGS track to OCR for detection (0 = whole track). OCR of a full movie is slow; a leading sample is enough to ID the language. Falls back to the full track if the sample has no text. |
The scanner has a dedicated Unknown language filter listing every title
that still has at least one und audio or subtitle track, with a count.
Use it to find what needs attention, then Detect-all-unknown across the
results. Ignored titles are included — an ignore rule means "don't
convert", not "hide that the audio is untagged" — so bulk detection reaches
them too.
- Text subtitles — the subtitle text is extracted (raw, without forcing
a decode that would reject non-UTF-8), the real charset is detected, and
langdetectidentifies the language. Runs inline during scanning. - Audio — a 30-second clip is extracted (16 kHz mono) and run through a faster-whisper tiny model's language identification. Several positions in the track are sampled and the most confident wins, so a track whose first sampled window is music or silence still gets identified from a better window.
- Image subtitles — the track is extracted with
mkvextract, OCR'd withpgsrip/tesseract, and the recognized text is fed to the samelangdetectstep. It OCRs with the Latin model first (covers all Latin-script languages cleanly); if the result can't be identified — the usual sign the sub is non-Latin — it re-OCRs with the CJK / Cyrillic / Arabic models.
Before any of the above, if the track's title names a language ("Traditional Chinese", "Romanian", "English SDH"), that's used directly — it's free and reliable, and it rescues forced/SDH tracks that carry too little text for content detection to work. Applies to audio and subtitle tracks alike.
In every case the recognized/identified language flows through one shared path: confidence gate → apply to the DB → write to the file → optional Plex refresh.
Detection is deliberately cautious. A result below the confidence threshold
(see the env vars above) is discarded and the track stays und — a wrong
language tag is worse than none, because it makes your keep/drop rules act
on the wrong track. Every step is fail-open: a missing model, an OCR
failure, an extraction error, or an offline model download leaves the track
und, logs a warning, and never blocks a scan or conversion.
The detection tooling is bundled in every image:
langdetect+charset-normalizer— text-subtitle detection (tiny).faster-whisper+pytesseract+pgsrip— audio + image-sub detection.tesseract-ocrand the language packs (eng,chi-sim,chi-tra,jpn,kor,rus,ara) plusmkvtoolnixand OpenCV's runtime libs.pgsriphandles PGS (Blu-ray). VobSub (DVD) OCR usessubtile-ocr, which is bundled in the NVENC images; on the portable/CPU images VobSub tracks fall back tound(PGS works everywhere). Detection is fail-open, so a missing tool never breaks a scan or conversion.
These add on the order of tens of MB to the image. The faster-whisper
tiny model (~75 MB) is not baked in — it downloads on the first audio
detection and caches in the data volume (/app/data/models), so it needs
outbound network the first time and a writable data volume.
"No new languages detected" immediately — audio detection needs the
faster-whisper model. If it returns instantly with nothing, the model
probably failed to load or download. Check the logs for [LANG-DETECT]
warnings. First use needs network + a writable /app/data/models.
An image sub stays und — OCR of a heavily-styled or low-resolution
bitmap sub can come back below the confidence threshold; that's the gate
working (better und than wrong). PGS is well-supported; VobSub is more
variable and its OCR (subtile-ocr) is only bundled in the NVENC images —
on the portable/CPU images VobSub tracks stay und.
A detected language didn't reach the file — the write is fail-open; if
mkvpropedit/the remux failed, Shrinkerr keeps the DB detection but the
file is unchanged. Check the logs for [LANG-DETECT] file write failed.
An audio track won't detect — some tracks (music-only, commentary beds,
very short) genuinely don't yield a confident spoken-language result across
any sampled window, and stay und. Lower SHRINKERR_LANG_DETECT_AUDIO_MIN
if you want to accept weaker guesses (at the risk of wrong tags).
Image-based subtitles only support Latin/CJK/Cyrillic/Arabic scripts.
Other scripts, or non-video subtitle formats, fall back to und.