Skip to content

6.66.0 - A session can be told to correct its lip-sync, and pays for it where the media is already committed

Choose a tag to compare

@superuser404notfound superuser404notfound released this 02 Sep 11:54
· 645 commits to main since this release

Lip-sync error belongs to the viewer's chain, not to the file: a soundbar or AVR adding video-processing latency, or the reverse. It is a per-setup nudge people set once and expect a player to honour on every session, and AVFoundation offers a host nothing to set it with. AVPlayerItem carries no audio-delay control, and an HLS-streamed asset vends no AVAssetTrack for an AVAudioMix to bind to. Requested by @cmcpherson274, closing the set filed alongside #460, #461 and #462.

The setting

player.setAudioDelay(0.120)          // audio 120 ms later than video
player.audioDelaySeconds             // what is in force
LoadOptions(audioDelaySeconds: 0.120)   // or already at load

Positive presents audio later, negative earlier, clamped to +/-2 s. It holds across the rebuilds a session makes on its own (reload at position, audio-track switch, AirPlay LAN swap, background return), because it lives in the options those rebuilds replay, and reloadAtCurrentPosition(applying:) from #460 can correct it like any other tuning field.

Where the offset is applied, and why not one place earlier

The engine moves audio at the last point it still holds the timestamps. That is a different place on each route, and on both of them it is the last place rather than the obvious one.

Route Applied at A change costs
.software the delivered sample's stamp, in AudioOutput.enqueue a seek to the current position
.loopback the audio track of the fMP4 segments, fixed per muxer the session-preserving reload (#460)
.remoteBypass, audio-only nothing to move kept for the next load, no-op logged

On the software path the obvious place is the decoder's stamp, and it is wrong three times over. That buffer is what the #95 audio tap mirrors, and the tap documents its sourceTime as the SOURCE axis and feeds transcription with it. It is stamped by AudioClockAnchor, the gapless clock that reads anything under its 100 ms discontinuity threshold as container rounding, so a 50 ms nudge would be absorbed into the running sample count and simply never heard. And the caller measures its look-ahead lead off that same timestamp against the synchronizer clock. Past the enqueue funnel, only the renderer sees the shift, and both software enqueue sites are covered by one place.

On the loopback path the offset goes into the muxer, and it is fixed for that muxer's life on purpose: two offsets inside one output track are not splicable. The seam gains a gap or an overlap of exactly the change, and a change that moves audio EARLIER walks into OutputTimestampSanitizer's strictly-increasing-DTS rule, which clamps it away entirely. A new offset is therefore a new muxer.

Video is never moved on either route. Its timestamps are the axis the session reports its position on and draws its subtitles against, so moving them would take the clock and the cues along with the sound.

.remoteBypass says so instead of pretending, in selectAudioTrack's style: AVPlayer owns that media selection and the engine never sees the timestamps. An audio-only session has no video for audio to be early or late against.

Not free the way setRate is

The media between the offset and the speaker is already committed to the previous value: up to AudioLookaheadPolicy.targetLeadSeconds of decoded audio on the software path, and the fetched segments on the loopback one. So a change is brought to the playhead rather than left to arrive when that drains, and the loopback half of that was settled by measurement, not by reading:

  • Seeking to the position AVPlayer already holds is a buffer hit. It keeps the segments cut with the old offset and plays them out regardless.
  • Dropping those segments under it does not help either. It turns the hand-over into a 6 s rebuffer.
  • Asking for the producer restart beside the seek is reported as a user scrub (setNativeScrubSeek), which opens a second seek ticket aimed at the re-cut segment's start and leaves it stalled, with phase stuck at seeking for the rest of the session.

Replacing the item is the one call that makes AVPlayer let go, which is exactly what #460's reload does, and it is the spelling this issue offered as its own alternative. A live session without a DVR window has no position to return to; it keeps the value and takes it at the next seam it makes on its own.

Measured

The delivered offset, read straight off the served segments rather than judged by ear. aetherctl serve --audio-delay <ms>, then the first audio and video packet timestamps out of init.mp4 + one segment (30 fps H.264 + 44.1 kHz AAC):

                 first video pts   first audio pts   audio - video
--audio-delay 0     16.000000 s       16.021769 s      +21.8 ms   (the source's own alignment)
--audio-delay 200   16.000000 s       16.221769 s     +221.8 ms
--audio-delay -150  16.000000 s       15.871769 s     -128.2 ms

Exactly the offset asked for, in both directions, with the video timestamp unchanged in every arm.

The runtime path on both routes, via aetherctl play --switch-audio-delay <ms>[@ms]:

software   set +200 ms at t=4.90s   seek landed 4.90s   sample at 3.994s delivered at 4.194s
loopback   set -150 ms at t=7.80s   reload            position 7.80s -> 7.80s, ~0.3 s held picture

The engine states the delivered offset, not the requested one, so a host can confirm the correction arrived without a capture card: [AudioOutput] AE#464 audio delay in effect: +200 ms (sample at 3.994s delivered at 4.194s) on the software path, [MP4SegmentMuxer] AE#464 cutting seg1+ with audio delay -150 ms (-6615 ticks @ 44100/1) on the loopback one.

swift test: 2543 tests in 349 suites plus 606 XCTest cases, 0 failures.

Full diff: 6.65.0...6.66.0