Repository navigation
feat: --dump-segments, so a lyrics edit doesn't re-run the ASR - #29
Merged
Merged
Conversation
`--segments` could read a segments file, but nothing could write one: `-f json` emits aligned lines, and `Segment` had `from_dict` without its mirror. So the cheap path existed and was unreachable from the CLI. `--dump-segments FILE` writes the transcription in exactly the shape `--segments` reads back. It lands *before* alignment, so a run that dies in the cheap half still leaves the expensive half on disk — that is the case the request came from. Measured on the public-domain demo (tiny model, no Demucs): 7.02s for the transcribing run, 0.04s re-running off the dump, byte-identical output. With `medium` and `--separate` the first number is minutes. The file is written one segment per line rather than with a plain `indent=`, which would put every word timing on three lines of its own — 443 lines instead of 8 for a 6-segment song. Five mutations were run against the new tests (dropping word timings, moving the dump after alignment, removing the no-op guard, breaking the hand-written brackets, re-expanding the indent). Each is caught by the test that claims to guard it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
--segmentscould read a segments file, but nothing could write one —-f jsonemits aligned lines, and
Segmenthadfrom_dictwithout its mirror. The cheappath existed and was unreachable from the CLI.
What this adds
--dump-segments FILEwrites the transcription in exactly the shape--segmentsreads back, and
Segment.to_dict/Word.to_dictto go with it.The write lands before alignment, not after: the expensive half is done, the
cheap half is the one you re-run, and a run that dies in the cheap half must not
cost the audio work. That is the case #28 came from. When
--segmentswas theinput it no-ops with a note rather than copying the input under a second name.
Measured
Real end-to-end on the public-domain demo (
tinymodel, no Demucs):Output byte-identical between the two. With
mediumand--separatethe firstnumber is minutes.
The file is one segment per line instead of a plain
indent=, which puts everyword timing on three lines of its own — 443 lines vs 8 for a 6-segment song.
Floats are left exactly as the ASR gave them; rounding would make the documented
from_dict(to_dict(x)) == xa lie for a cosmetic gain.Tests
Six new tests (3 model, 3+1 CLI). Five mutations run against them, each caught by
the test that claims to guard it:
to_dictdrops word timings--segmentsno-op testindent=Full suite: 67 passed on 3.9 and 3.14.
Version bumped to 0.4.0, matching how #6 handled a feature. Publishing is a
separate decision.
Closes #28
🤖 Generated with Claude Code