Skip to content

docs: variable unit size places more lines, and some of them badly - #22

Merged
ijuinryukichi merged 1 commit into
mainfrom
docs/variable-pairing-negative-result
Jul 25, 2026
Merged

ijuinryukichi merged 1 commit into
mainfrom
docs/variable-pairing-negative-result

Conversation

@ijuinryukichi

Copy link
Copy Markdown
Owner

The next idea after --pairing auto: stop fixing the unit size up front and choose (lines, segment) jointly at each step. pairing is a rounded average, so it is wrong for part of any track — at 76 lines over 64 segments the true ratio is 1.19, and a fixed 1 leaves 27 lines unplaced while a fixed 2 straddles boundaries.

Measured worse. Documentation only, no behaviour change.

On the shipped metric it looks free

unit size lines placed mean |err| within 0.5 s worst 2nd lines outside their GT window
fixed, from the ASR (shipped) 49/76 0.31 s 15/20 0.8 s 0 of 16
variable, k ≤ 2 66/76 0.42 s 14/20 2.0 s 2 of 18
variable, k ≤ 2, only if it wins by 0.1 66/76 0.32 s 15/20 0.8 s 1 of 18 (by 13.3 s)

The last column is the point. The ground truth is marked per two-line pair and the metric scores only the first line of each — which is exactly where the extra placements do not land. Second lines have a free check: each must start inside its pair's window.

What the 17 "free" placements bought

The last verse line scores 0.000 against the hook segment on its own and 0.286 once the following hook line is absorbed into the same unit — over the 0.25 threshold. The breath split then drags that verse line 13.3 s forward into the hook, and displaces the hook line that had been placed correctly.

A unit chosen to maximise similarity will straddle a section boundary, and such a unit needs only its tail to match. pairing=1 cannot do this because a one-line unit has no tail. So variable pairing manufactures the very tail-match pathology that the dedicated veto in Known limits already failed to fix.

Of the two extra placements that can be checked, one is right and one is 13.3 s wrong.

It fails from the other side too

On a track where the ASR merges two lines consistently, deviating downward orphans the remaining line onto the next segment and shifts every later unit's phase: within-0.5 s 26/33 → 18/33, and the 4×-repeated hook lands a repetition early.

Restricting deviation to upward only appears to fix that — but only because pairing=2 with k ≤ 2 leaves upward no room. Allowing k ≤ 3 breaks the same track again (26/33 → 13/33).

Why the premise was wrong

The four earlier attempts all changed which segment a unit selects, so the plan here was to change how much a unit consumes and dodge that coupling. It does not dodge it: consumption sets how fast the segment cursor advances relative to the line cursor, so it moves idx as well — one step removed, same gains-and-losses-in-adjacent-pairs behaviour.

Also disclosed

The Accuracy section now states that the headline figures score first lines only, and that the shipped configuration passes the second-line containment check (19/19 and 29/33, the four misses being the already-documented repeated hook). The tables were not hiding a second failure mode — but the distinction decides any change judged on placement count.

Verification

57 tests pass. Harness in toryu-web/scratchpad/lyrics_align_bench: exp_var_pairing.py, exp_var_diag.py, exp_var_rows.py, exp_var_margin.py, exp_var_updown.py, exp_second_line.py. Every sweep endpoint (m=inf) reduces exactly to the shipped baseline on all four (track, model) configs — the first version did not, and that bug made the whole first sweep unreadable.

`pairing` is a rounded average, so it is wrong for part of any track: at
76 lines over 64 segments the true ratio is 1.19, and a fixed 1 leaves 27
lines unplaced while a fixed 2 straddles boundaries. Choosing (lines,
segment) jointly at each step raises placements 49/76 -> 66/76 and barely
moves the first-line error, so on the shipped metric it reads as free.

It is not. The ground truth is marked per two-line pair and the metric
scores only the first line of each, which is where the extra placements
do not land. Checking the second lines — each must start inside its
pair's window — shows what was bought: the last verse line scores 0.000
against the hook segment alone and 0.286 once the following hook line is
absorbed into the same unit, clearing the 0.25 threshold, so it lands
13.3 s late inside the hook and displaces the hook line that had been
correct. A unit chosen to maximise similarity will straddle a section
boundary, and such a unit needs only its tail to match; pairing=1 cannot
do this because a one-line unit has no tail. Of the two checkable extra
placements, one is right and one is 13.3 s wrong.

On a track where the ASR merges two lines consistently it fails from the
other side: deviating downward orphans the remainder onto the next
segment and shifts every later unit's phase (within-0.5 s 26/33 -> 18/33,
the 4x-repeated hook landing a repetition early). Restricting deviation
to upward only looks safe, but only because pairing=2 with k<=2 leaves it
no room — allowing k<=3 breaks that track the same way (26/33 -> 13/33).

This also disposes of the reason for trying it. The four earlier attempts
changed which segment a unit selects, so the plan was to change how much
a unit consumes and dodge that coupling. Consumption sets how fast the
segment cursor advances relative to the line cursor, so it moves idx too
— one step removed, same gains-and-losses-in-adjacent-pairs result.

Documentation only; no behaviour change. Harness in toryu-web
scratchpad/lyrics_align_bench (exp_var_*.py, exp_second_line.py).
@ijuinryukichi
ijuinryukichi merged commit fda4c07 into main Jul 25, 2026
3 checks passed
@ijuinryukichi
ijuinryukichi deleted the docs/variable-pairing-negative-result branch July 25, 2026 07:51
ijuinryukichi added a commit that referenced this pull request Jul 25, 2026
`pairing` is a rounded average, so it is wrong for part of any track: at
76 lines over 64 segments the true ratio is 1.19, and a fixed 1 leaves 27
lines unplaced while a fixed 2 straddles boundaries. Choosing (lines,
segment) jointly at each step raises placements 49/76 -> 66/76 and barely
moves the first-line error, so on the shipped metric it reads as free.

It is not. The ground truth is marked per two-line pair and the metric
scores only the first line of each, which is where the extra placements
do not land. Checking the second lines — each must start inside its
pair's window — shows what was bought: the last verse line scores 0.000
against the hook segment alone and 0.286 once the following hook line is
absorbed into the same unit, clearing the 0.25 threshold, so it lands
13.3 s late inside the hook and displaces the hook line that had been
correct. A unit chosen to maximise similarity will straddle a section
boundary, and such a unit needs only its tail to match; pairing=1 cannot
do this because a one-line unit has no tail. Of the two checkable extra
placements, one is right and one is 13.3 s wrong.

On a track where the ASR merges two lines consistently it fails from the
other side: deviating downward orphans the remainder onto the next
segment and shifts every later unit's phase (within-0.5 s 26/33 -> 18/33,
the 4x-repeated hook landing a repetition early). Restricting deviation
to upward only looks safe, but only because pairing=2 with k<=2 leaves it
no room — allowing k<=3 breaks that track the same way (26/33 -> 13/33).

This also disposes of the reason for trying it. The four earlier attempts
changed which segment a unit selects, so the plan was to change how much
a unit consumes and dodge that coupling. Consumption sets how fast the
segment cursor advances relative to the line cursor, so it moves idx too
— one step removed, same gains-and-losses-in-adjacent-pairs result.

Documentation only; no behaviour change. Harness in toryu-web
scratchpad/lyrics_align_bench (exp_var_*.py, exp_second_line.py).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant