Repository navigation
docs: variable unit size places more lines, and some of them badly - #22
Merged
Merged
Conversation
`pairing` is a rounded average, so it is wrong for part of any track: at 76 lines over 64 segments the true ratio is 1.19, and a fixed 1 leaves 27 lines unplaced while a fixed 2 straddles boundaries. Choosing (lines, segment) jointly at each step raises placements 49/76 -> 66/76 and barely moves the first-line error, so on the shipped metric it reads as free. It is not. The ground truth is marked per two-line pair and the metric scores only the first line of each, which is where the extra placements do not land. Checking the second lines — each must start inside its pair's window — shows what was bought: the last verse line scores 0.000 against the hook segment alone and 0.286 once the following hook line is absorbed into the same unit, clearing the 0.25 threshold, so it lands 13.3 s late inside the hook and displaces the hook line that had been correct. A unit chosen to maximise similarity will straddle a section boundary, and such a unit needs only its tail to match; pairing=1 cannot do this because a one-line unit has no tail. Of the two checkable extra placements, one is right and one is 13.3 s wrong. On a track where the ASR merges two lines consistently it fails from the other side: deviating downward orphans the remainder onto the next segment and shifts every later unit's phase (within-0.5 s 26/33 -> 18/33, the 4x-repeated hook landing a repetition early). Restricting deviation to upward only looks safe, but only because pairing=2 with k<=2 leaves it no room — allowing k<=3 breaks that track the same way (26/33 -> 13/33). This also disposes of the reason for trying it. The four earlier attempts changed which segment a unit selects, so the plan was to change how much a unit consumes and dodge that coupling. Consumption sets how fast the segment cursor advances relative to the line cursor, so it moves idx too — one step removed, same gains-and-losses-in-adjacent-pairs result. Documentation only; no behaviour change. Harness in toryu-web scratchpad/lyrics_align_bench (exp_var_*.py, exp_second_line.py).
ijuinryukichi
added a commit
that referenced
this pull request
Jul 25, 2026
`pairing` is a rounded average, so it is wrong for part of any track: at 76 lines over 64 segments the true ratio is 1.19, and a fixed 1 leaves 27 lines unplaced while a fixed 2 straddles boundaries. Choosing (lines, segment) jointly at each step raises placements 49/76 -> 66/76 and barely moves the first-line error, so on the shipped metric it reads as free. It is not. The ground truth is marked per two-line pair and the metric scores only the first line of each, which is where the extra placements do not land. Checking the second lines — each must start inside its pair's window — shows what was bought: the last verse line scores 0.000 against the hook segment alone and 0.286 once the following hook line is absorbed into the same unit, clearing the 0.25 threshold, so it lands 13.3 s late inside the hook and displaces the hook line that had been correct. A unit chosen to maximise similarity will straddle a section boundary, and such a unit needs only its tail to match; pairing=1 cannot do this because a one-line unit has no tail. Of the two checkable extra placements, one is right and one is 13.3 s wrong. On a track where the ASR merges two lines consistently it fails from the other side: deviating downward orphans the remainder onto the next segment and shifts every later unit's phase (within-0.5 s 26/33 -> 18/33, the 4x-repeated hook landing a repetition early). Restricting deviation to upward only looks safe, but only because pairing=2 with k<=2 leaves it no room — allowing k<=3 breaks that track the same way (26/33 -> 13/33). This also disposes of the reason for trying it. The four earlier attempts changed which segment a unit selects, so the plan was to change how much a unit consumes and dodge that coupling. Consumption sets how fast the segment cursor advances relative to the line cursor, so it moves idx too — one step removed, same gains-and-losses-in-adjacent-pairs result. Documentation only; no behaviour change. Harness in toryu-web scratchpad/lyrics_align_bench (exp_var_*.py, exp_second_line.py).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The next idea after
--pairing auto: stop fixing the unit size up front and choose(lines, segment)jointly at each step.pairingis a rounded average, so it is wrong for part of any track — at 76 lines over 64 segments the true ratio is 1.19, and a fixed 1 leaves 27 lines unplaced while a fixed 2 straddles boundaries.Measured worse. Documentation only, no behaviour change.
On the shipped metric it looks free
The last column is the point. The ground truth is marked per two-line pair and the metric scores only the first line of each — which is exactly where the extra placements do not land. Second lines have a free check: each must start inside its pair's window.
What the 17 "free" placements bought
The last verse line scores 0.000 against the hook segment on its own and 0.286 once the following hook line is absorbed into the same unit — over the 0.25 threshold. The breath split then drags that verse line 13.3 s forward into the hook, and displaces the hook line that had been placed correctly.
A unit chosen to maximise similarity will straddle a section boundary, and such a unit needs only its tail to match.
pairing=1cannot do this because a one-line unit has no tail. So variable pairing manufactures the very tail-match pathology that the dedicated veto in Known limits already failed to fix.Of the two extra placements that can be checked, one is right and one is 13.3 s wrong.
It fails from the other side too
On a track where the ASR merges two lines consistently, deviating downward orphans the remaining line onto the next segment and shifts every later unit's phase: within-0.5 s 26/33 → 18/33, and the 4×-repeated hook lands a repetition early.
Restricting deviation to upward only appears to fix that — but only because
pairing=2withk ≤ 2leaves upward no room. Allowingk ≤ 3breaks the same track again (26/33 → 13/33).Why the premise was wrong
The four earlier attempts all changed which segment a unit selects, so the plan here was to change how much a unit consumes and dodge that coupling. It does not dodge it: consumption sets how fast the segment cursor advances relative to the line cursor, so it moves
idxas well — one step removed, same gains-and-losses-in-adjacent-pairs behaviour.Also disclosed
The Accuracy section now states that the headline figures score first lines only, and that the shipped configuration passes the second-line containment check (19/19 and 29/33, the four misses being the already-documented repeated hook). The tables were not hiding a second failure mode — but the distinction decides any change judged on placement count.
Verification
57 tests pass. Harness in
toryu-web/scratchpad/lyrics_align_bench:exp_var_pairing.py,exp_var_diag.py,exp_var_rows.py,exp_var_margin.py,exp_var_updown.py,exp_second_line.py. Every sweep endpoint (m=inf) reduces exactly to the shipped baseline on all four (track, model) configs — the first version did not, and that bug made the whole first sweep unreadable.