easr is a local Apple Silicon CLI that turns audio and video files into
subtitle files (.srt, .vtt) and timestamp-aligned JSON.
Use it when you want local speech recognition, forced alignment, readable subtitles, and machine-friendly timing data without running a server.
Current scope:
- runtime target: macOS on Apple Silicon
- backend: MLX with Qwen3 ASR and Qwen3 ForcedAligner
- output: SRT, WebVTT, and JSON
- license: MIT
- not included yet: translation, speaker diarization, Linux/Windows support
For each supported media file, easr writes:
<name>.srtfor subtitle players and editors<name>.vttfor web video workflows<name>.jsonfor downstream tools that need segments, tokens, timestamps, language metadata, and provider metadata<name>.metrics.jsonwhen--verboseis enabled
easr accepts files, directories, and glob patterns. Directory scans are
non-recursive by default, and recursive processing is opt-in with --recursive.
- macOS on Apple Silicon
- Python
>=3.14,<3.15 ffmpegandffprobeavailable onPATHuvif you run from a source checkout- network access on first run so the models can be downloaded from Hugging Face
Install the media tools with Homebrew:
brew install ffmpegDefault provider models:
Install from PyPI:
python3.14 -m pip install "echoalign-asr-mlx[mlx]"
easr --helpRun from a source checkout:
uv sync --extra mlx
uv run --python 3.14 --extra mlx easr --helpIf you use the source checkout flow, prefix examples in this README with:
uv run --python 3.14 --extra mlx easr ...Transcribe one file:
easr ./demo.mp4Write outputs to a custom directory:
easr ./demo.mp4 --output-dir ./subtitlesProcess a directory:
easr ./mediaProcess a directory recursively:
easr ./media --recursiveProcess a glob pattern:
easr "./media/**/*.mp4" --recursiveExport token-level subtitle and JSON views:
easr ./demo.mp4 --granularity tokenShow detailed progress and write metrics:
easr ./demo.mp4 --verboseAudio:
wavmp3m4aflacaac
Video:
mp4movm4vmkvwebm
Default output directory name: outputs.
When the input is a single file, outputs are written next to that file:
/project/media/demo.mp4
/project/media/outputs/demo.srt
/project/media/outputs/demo.vtt
/project/media/outputs/demo.json
When the input is a directory or the current directory, outputs are written under that input root:
/project/media/
a.mp4
nested/b.wav
/project/media/outputs/
a.srt
a.vtt
a.json
nested/b.srt
nested/b.vtt
nested/b.json
Use --output-dir to choose another output root.
The JSON export keeps the readable transcript and the alignment data used to create subtitle views.
Common top-level fields:
source_pathprovider_namedetected_languagesegmentssource_mediagranularityitems
Each segment includes text, start/end timestamps, language metadata, optional
speaker metadata, and token timing when available. source_media includes the
prepared audio path, VAD metadata, and provider diagnostics such as processing
strategy, duration, window counts, quality pass counts, and window diagnostics.
| Option | Meaning |
|---|---|
inputs |
File, directory, or glob pattern. Defaults to the current directory. |
--recursive |
Recursively scan directory inputs. |
--output-dir PATH |
Override the default output directory root. |
--granularity sentence |
Use segment boundaries for subtitle entries and JSON items. This is the default. |
--granularity token |
Use token timing for subtitle entries and JSON items. |
--no-vad |
Disable voice activity detection preprocessing. |
--verbose |
Print detailed progress and write <name>.metrics.json. |
--version |
Show the installed package version. |
--help |
Show CLI help. |
Voice activity detection is enabled by default. easr scans the prepared audio,
finds likely speech ranges, groups them into padded chunks, and asks the
provider to process only those ranges. Final subtitle timestamps remain on the
original media timeline.
Disable VAD when you want full-duration provider processing:
easr ./demo.mp4 --no-vadIf VAD fails, easr falls back to full-duration processing. If VAD succeeds and
finds no speech, easr writes successful empty subtitle outputs.
Fish shell users can generate or install completions:
easr completion fish
easr completion install fishThe install command writes:
~/.config/fish/completions/easr.fish
Existing completion files at that path are overwritten.
- exit code
0: all discovered files processed successfully - exit code
1: no supported input was found, environment preflight failed, or at least one file failed in a batch - batch files are processed one by one
- failures are reported per file to stderr
- other files continue processing after a per-file failure
- the first run may be slower because model files are downloaded and cached
Install the media tools and make sure they are visible from the shell running
easr:
brew install ffmpeg
which ffmpeg
which ffprobeCheck that you are on Apple Silicon, using the expected Python environment, and installed the MLX runtime extra:
python3.14 -m pip install "echoalign-asr-mlx[mlx]"For a source checkout:
uv sync --extra mlx
uv run --python 3.14 --extra mlx easr --helpThis is expected when the Qwen3 model files are downloaded and the local cache is warmed. Later runs should be faster.
- Translation is not implemented.
- Speaker diarization is not implemented.
- Subtitle segmentation quality depends on model and alignment behavior.
- The public CLI does not expose provider selection.
