Releases: superuser404notfound/AetherEngine
Release list
6.65.0 - A source that loses its audio says so, and says which silence it is
A source whose audio can neither stream-copy into fMP4 nor go through the bridge plays video-only. state reaches .playing, nothing failed by the error taxonomy's lights, and until now the only account was a log line. Requested by @cmcpherson274, completing the three asks filed alongside #460 and #461.
The fact, not the reconstruction
$audioDelivery publishes an AudioDelivery:
| Value | Meaning |
|---|---|
.none |
no session |
.noAudioInSource |
the source carries no audio track, or none was selected |
.streamCopy |
the source bitstream is muxed into fMP4 unchanged |
.bridged |
decoded and re-encoded (FLAC or E-AC-3) for the fMP4 pipeline |
.decoded |
libavcodec decodes and the engine renders it |
.droppedNoPipeline |
the source HAS audio and none of it could be delivered |
.playerManaged |
AVFoundation owns the audio (remote-HLS bypass, native audio-only host) |
player.$audioDelivery
.filter { $0 == .droppedNoPipeline }
.sink { _ in ladder.demote(reason: .silentVideo) }.droppedNoPipeline is the one a fallback ladder acts on, and it is the counterpart of PlaybackErrorKind.audioBridgeProducedNoOutput: the same user outcome from the other end of the cascade. That kind fails loudly when a bridge WAS built and then decoded nothing; this value reports one that could never be built at all. Neither ends a ladder, since re-serving the source with audio the pipeline can carry plays it.
It is deliberately not a PlaybackErrorKind. That taxonomy is terminal: publishError moves state to .error, and errorInfo is cleared by the state's own move away from it. Video-only playback is neither terminal nor an error for every host, so carrying it there would break the invariant that ties those two publishers together.
Why the old inference recipe could not simply be documented
The drop was inferable: a non-empty audioTracks paired with a nil activeAudioDecoder at .playing. That is an undocumented pairing of two publishers, and it broke in both directions.
- It read as a drop where the probe had merely failed to list the tracks, so the list published empty while the session had found audio and lost it.
- It read as healthy on the software path, and that half was outright wrong.
activeAudioDecoderwas built there from the PROBE's track list, so a session whoseAudioDecoder.openhad refused the stream, gone video-only and set its audio index to -1 still published"libavcodec AC3 → CoreAudio".
The label now comes from the host's own resolved index, which also repairs the case it got wrong the other way: the #133 live-TS by-type fallback resolves an index the engine's pick does not know about. Show $activeAudioDecoder, classify on $audioDelivery.
Derived, so it cannot drift
AudioDelivery is folded from playbackBackend, the session's effective options and the live pipeline's own classification, never assigned on its own (the videoRoute arrangement, #321). Each pipeline classifies at the point of decision: the loopback cascade at its three exits, SoftwarePlaybackHost once its decoder has had its chance. The routing bits alone can never produce .droppedNoPipeline, which is pinned by its own test: the drop is only ever reported by the pipeline that dropped, about itself.
Measured
Every case on real media, in one run each:
# native loopback
aac.mp4 audio delivery=streamCopy pipeline=Stream-copy (AAC) tracks=1
mp3.mkv audio delivery=bridged pipeline=MP3 → FLAC bridge tracks=1
silent.mp4 audio delivery=noAudioInSource pipeline=none tracks=0
# software path (--sw), and the audio-only hosts
aac.mp4 audio delivery=decoded pipeline=libavcodec AAC → CoreAudio
tone.m4a delivery=playerManaged decoder=AVPlayer (native audio host)
tone.m4a delivery=decoded (FFmpeg audio-only host, customio)
# the drop, forced on both paths
aac.mp4 --drop-audio audio delivery=droppedNoPipeline pipeline=none tracks=1
aac.mp4 --sw --drop-audio audio delivery=droppedNoPipeline pipeline=none tracks=1
The label fix, both arms of the same forced software drop:
before audio delivery=droppedNoPipeline pipeline=libavcodec AAC → CoreAudio
after audio delivery=droppedNoPipeline pipeline=none
aetherctl play --drop-audio is the new harness behind those two: it makes every audio pipeline fail, because the drop otherwise needs a source whose codec the build has no decoder for (AC-4 being the realistic one). It is loud in both the CLI and the engine log on purpose, since a forced classification read as a real one would be worse than no harness at all.
swift test: 2526 tests in 343 suites plus 606 XCTest cases, 0 failures.
Full diff: 6.64.0...6.65.0
6.64.0 - A session can be told to decode in software, and only in the direction that helps
VTCapabilityProbe.canHardwareDecode fails open by design: four classes it cannot classify keep the native path. That is the right default, and occasionally wrong, because when VideoToolbox then cannot build a decoder for what arrives, the item reaches readyToPlay and renders nothing. Until now a host that could SEE that had nothing to do about it. Requested by @cmcpherson274, alongside #460, which this composes with.
The escape was the missing half, not the detection
Detecting the symptom is a host's own job and was never the gap: AVPlayerItemVideoOutput.hasNewPixelBuffer against softwareHostFramesEnqueued already says "ready, playing, no picture". What did not exist was a way to answer it in place. The two levers that were there are both wrong for the job:
setForceSoftwarePathForTestingis process-global. On an engine shared across sessions it drags every concurrent one onto the software host with the one that needed it, which is why it is documented as test-only.- The only per-session route onto
SoftwarePlaybackHostwas presenting a customIOReaderwhose seek fails, since a forward-only VOD source is forced to software. That reaches the host by costing the source everything the host came for: no seeking, no mid-session audio switch, no title switch, andreloadAtCurrentPositiongated off, so it does not even compose with #460.
LoadOptions.preferredDecodePath is the lever, scoped to the session and costing the source nothing:
options.preferredDecodePath = .software // at load
try await player.reloadAtCurrentPosition { $0.preferredDecodePath = .software } // mid-play, #460It is not VOD-only. The #2 capability gate runs behind !options.isLive, so a live load never reaches a classification step at all, and a live session is exactly where a host has no second engine left to hand a struggling stream to.
One-way, and it does not suspend what software cannot represent
DecodePath has .automatic and .software, and no .native. Every route the engine sends to software it sends there because the native path cannot serve it (AV1 without hardware decode, VP9, a forward-only source, MVC carriage, a format VideoToolbox cannot hardware-decode), so a native preference would buy a black screen. A host's evidence is only ever "this native session is not decoding", never the reverse, and the missing case is the guarantee.
The override says which host serves the session, not what that host can do. It sits before the guards that run after the routing decision, and those still fire: a source whose only signal is IPT-PQ-c2 (Dolby Vision HEVC Profile 5, AV1 Profile 10.0) still fails with dolbyVisionUnplayableOnSoftwarePath rather than decoding as YCbCr and rendering green/purple, and a demuxed-audio live source still fails rather than playing silent. On nativeRemoteHLS there is no decode path to prefer at all, since AVPlayer plays the remote playlist and the engine decodes nothing, so the engine logs that it ignored the preference rather than letting it look applied.
Measured
Same 300 s H.264 fixture, three arms:
# load-time, no override
[AetherEngine] dispatch: codec=27 → native
# load-time, --sw
[AetherEngine] #461: host asked for the software path, overriding the native route
[AetherEngine] dispatch: codec=27 → software
# mid-play, --reload-applying decode-path=software at +10 s
t=09 state=playing cur=9.00 buf=56.00
HOSTCALL reloadAtCurrentPosition(applying: preferredDecodePath=software) at +10000 ms (t=9.90s)
[AetherEngine] #460: reload applying preferredDecodePath
[AetherEngine] #461: host asked for the software path, overriding the native route
[AetherEngine] dispatch: codec=27 → software
#460 correction applied, session resumed at 10.81s from 9.90s (state=playing)
A native session became a software session at its own playhead, with no restart. On a live load the override reaches the same routing decision, measured on an HLS fixture; the live software path itself is what aetherctl live --sw has always exercised and is unchanged here.
aetherctl play --sw now drives the real option instead of the process-global test hook, so the harness exercises what a host calls. live --sw and dvr keep the hook, since those harnesses run several sessions and want every one of them on the software host. New correction key --reload-applying decode-path=software|automatic.
swift test: 2507 tests in 340 suites plus 606 XCTest cases, 0 failures.
Full diff: 6.63.0...6.64.0
6.63.0 - An option a session can be corrected on is changed without restarting the item
Every in-place rebuild this engine offered a host was tied to a selection (an audio track, a subtitle track, a disc title) or replayed the session's own options verbatim (reloadAtCurrentPosition()). An option a host needed to CORRECT mid-session had no in-place answer at all: it had to load the source again, and the viewer watched the picture drop and rebuild. Requested by @cmcpherson274, who read the reload path in 6.60.0 and priced the alternative honestly enough to name what it actually costs.
A fresh load is not the same rebuild
The cost is not the teardown. reloadAtCurrentPosition() is itself a from-scratch load(), so both pay the same one, both keep the native host where the load allows it, and the audio pick rides the load's own override either way. What a host-issued load() cannot reach is the two fields the reload sets from inside the engine, subtitleSessionCarryover and isLiveRejoin. Without them the id-exact external-subtitle registry, every mid-session addExternalSubtitleTrack, the host's explicit subtitle authority (subtitles explicitly OFF included) and the live rejoin contract are wiped and re-derived by auto-selection.
reloadAtCurrentPosition(applying:) is that same reload with the options it replays taken from the host:
try await player.reloadAtCurrentPosition { $0.httpHeaders["Authorization"] = "Bearer \(fresh)" }The closure is seeded with what the session is CURRENTLY running on, which is not always what was passed to load: the engine rewrites its own routing fields on a reroute. The change is installed into the session before the rebuild, which is what makes the internal reopens that follow (an audio switch, a background reload) replay it instead of reverting to the load-time value, and it is also what covers the custom-source branch, which reads those fields one at a time and never takes a struct.
An option that names the session is not one it can be corrected on
isLive, audioOnly, nativeRemoteHLS and sequentialOrigin each open the source on a different pipeline, and the engine writes the last two itself (the #154 / #168 remote-HLS reroute, the probe's no-video fallback). A change to one of them is a different item, not a correction of this one, so it is refused by name through AetherEngineError.loadIdentityNotCorrectable rather than half-applied. That refusal and sessionNotReloadable both run before any teardown, so a refused correction leaves the session playing untouched, and both are all-or-nothing: a correction refused for one field installs none of it.
sessionReloadRefusal answers "would a reload rebuild anything" without attempting one. reloadAtCurrentPosition() returns silently when there is nothing to rebuild, and a host correcting a session had no way to tell that silence from a rebuild that ran except by timing one out.
Measured against a header-logging origin
New CLI lever --reload-applying <key>=<value> with --reload-applying-at <ms>, driving the correction on a session that is already playing. With --header "X-Auth: stale" at load and --reload-applying header.X-Auth=fresh at +10 s, over a 300 s H.264/AAC source:
REQ t=120.59 auth=stale range=bytes=0-33554431
REQ t=128.84 auth=stale range=bytes=17579980-50348495
HOSTCALL reloadAtCurrentPosition(applying: httpHeaders[X-Auth]) at +10000 ms (t=9.90s)
[AetherEngine] #460: reload applying httpHeaders, audioBridgeMode
REQ t=130.79 auth=fresh range=bytes=0-33554431
REQ t=130.87 auth=fresh range=bytes=17334168-50348495
#460 correction applied, session resumed at 10.90s from 9.90s (state=playing)
The corrected header reaches the rebuilt session's source IO, and the transport runs straight through: the buffer goes 56 s to 60 s across the rebuild with no rebuffer. The other arm, --reload-applying is-live=true, prints correction refused: LoadOptions.isLive names the session and the session plays on without a tick of interruption.
Where a correction lands
A URL source rebuilds through the full load, so every field applies. A custom IOReader source rebuilds through the narrower reopen that keeps the retained reader, which replays loadedOptions field by field: the routing, probe-budget, live-join, deinterlace and subtitle-preparation fields apply there, and the four load itself consumes (preferredAudioLanguages, externalSubtitles, maxConcurrentSourceRequests, autoplay) are installed and take effect at the next load. Documented in docs/api.md rather than left for a host to find out.
LoadOptions field coverage is pinned by a test, so a field added to the struct fails the suite until someone decides which of the two halves it belongs in.
swift test: 2499 tests, 338 suites, 0 failures.
Full diff: 6.62.3...6.63.0
6.62.3 - The run that answers a placement is the one that opens at its seam
A placement is a claim about bytes in AVPlayer's timeline. Two of the ways this engine checked that claim turned out to answer a different question than the one they were asked, and both are visible with the picture as witness. Reported by @rrgomes, who retested AE#418 round 6 on 6.59.1 across three seek bursts on the same asset and brought back a placement whose composition was kept over bytes that never arrived.
#418 seg735 placed (advertised 2940.479s, worth -28.028s, ...): axis -12.889s -> -40.835s
#418 seg735 opened no run of its own to measure; keeping the composed axis -40.835s
[HLSSegmentProducer] seg-738.m4s partial at teardown (8321003 B) discarded, not adopted
...
#418 seg800 placed on base -12.856s, not -40.835s (residual +27.979s): axis -40.835s -> -12.856s
The axis stood 28 s from the picture for 4.5 s, and the placement that put it there had never reached the item at all.
The run that answers a placement is the one that opens where its segment begins
Round 4 identified a placement's run by asking which run was NEW, against a baseline of what the item held when the placement was recorded. During a seek burst that is a different question with a different answer: a later seek's run is new by every baseline test, and the placement's own run, opening below the playhead, looks like backfill. Reproduced on tc-bf-cues-lie.mkv over a throttled origin (play --picture-probe --seek-every 1 --seek-count 4 --seek-pattern 70,53,71,54):
seg11 placed (advertised 44.000, worth -1.000): axis -9.000 -> -10.000, seam item 53.000
sample 1-4 ranges=[53.083-70.035] its own run, one lead above the seam, refused four times
sample 5 ranges=[74.208-86.099] a later seek's run, adopted: axis -10.000 -> -31.208
The picture reads -10.125 for the rest of that session, so 21 s of axis came out of a run that had nothing to do with the placement, while the run that did belong to it had been in hand for a second. On a 60 s-drought fixture the same test adopted a run 41.667 s away.
A placed segment's first sample goes to its advertised start read through the base its timeline carries, so the run that answers a placement is the one that opens at the seam the composition predicts, give or take the presentation lead the reading exists to measure. A timeline AVPlayer threw away carries no base, so the same segment then opens at its advertised start itself, and that is the second admissible answer (measured: a seek 35 s out of the buffer opens [52.000-75.969] for an advertised 52.000, base 0.000, and the 10.3 s correction it produces is right). The identity picks the run; the base is still read off it, residual and all, which is what keeps round 4's corrections working.
One consequence worth naming: the placement round 5 called unmeasurable, and paid a whole design for, opens on its seam to the millisecond. It was readable all along.
Bytes nobody holds never moved the axis
opened no run of its own to measure covered a placement that happened and cannot be read AND a placement that never happened, and kept the composition for both. A placement with no reading is now kept only while its request is still being answered, or when AVPlayer holds what it placed, and rolled back otherwise:
#418 segN never reached AVPlayer's timeline (nothing held at item Xs); rolling the composed axis Ys back to Zs
The local server reports what became of every media segment request for this, with a 503 counted as a retry rather than an answer, and the wait covers the #93 slow-serve window so a deep re-aim is never mistaken for a placement that never landed.
One reading is a sample, not a measurement
Round 4 accepted that a reading may be wrong once because the next placement undoes it. Round 6 then let one reading set the presentation-lead coefficient for the whole session, so it inherited the right to be wrong without the correction that made it safe. Readings resolve the base to whole frames, and where a lead is one frame that noise is a whole unit of coefficient: nine readings on the reporting asset came back 0, +-1 and +-2 frames, and one stray swung a settled session the full clamp, three leads in a step. The session now holds the median of its readings, an even count holding what the odd one before it settled on, and only a reading taken from the placement's own run teaches it at all.
Measured
play --picture-probe reads the axis off AVPlayer's own video output, so capErr is the error a host placing a cue at sourceTime would make. Ticks outside a fifth of a second, two runs per arm, 6.62.2 against 6.62.3:
| arm | 6.62.2 | 6.62.3 |
|---|---|---|
70,53,71,54 (wrong adoption) |
30 of 52, max 21.150 | 0 of 52, max 0.075 |
45,99,46,98 (wide drought) |
16 of 47, max 52.133 | 0 of 48, max 0.075 |
80,42,81,43 (deep re-aim) |
34 of 48, max 27.833 | 2 of 48 |
65,60,70,58 (round 6 burst) |
0 of 52 | 0 of 52 |
60,96,61,97, 65,58,66,57, 60,53,61,54 |
0 | 0 |
The remaining ticks on the deep re-aim arm are one seam-crossing tick per run, present in both builds: below the seam the picture is still the old epoch's. The unchanged arms are the point as much as the fixed ones, round 6's compose to the same axes they did.
swift test: 2485 tests, 336 suites, 0 failures.
Full diff: 6.62.2...6.62.3
6.62.2 - Every live item's axis is stated by the playlist it loads
A live item's zero is the first segment ITS playlist listed. That rule arrived with AE#446 round 4 and it is right, but the engine only ever heard it from one place. Reported by @cmcpherson274, who measured the axis on his own stack on 6.60.0, on a tvOS simulator and on an Apple TV 4K gen 3, and noticed that the item a session STARTS with reads a number it should not have.
#446 this item's playlist began 0.05s into the session
That playlist begins at 0.00 s. The reading is the reconstruction, not the playlist.
A statement that was gated on an unrelated condition
6.57.0 introduced the statement: the build that serves a playlist knows which segment it listed first, so it can say where the item that loads it begins, and a statement outranks a reconstruction. But it was recorded only on the builds that also carried an EXT-X-START rejoin placement, which is a different question about a different feature. Every other live attach stayed on the reconstruction:
- the session's own first item,
- the #130 media fallback, whose own comment says the window may have slid since the failed master attempt,
- the #35 readiness-gate reloads,
- an AirPlay hop,
- and the rejoin branch whose target had been evicted, which arms no placement and therefore had none.
The reconstruction is a difference between two independently sampled quantities, the segment cache's resident floor and the item's own reported seekable start:
max(0, producerFloorSession - (itemSeekableStart + shift))It is only as good as the older of its two samples, and it is latched for the item's whole life on the first tick that produces any number at all. On 6.57.0 it read 0 for an item whose playlist began 6.76 s into the session, which put a correctly placed item through a correcting seek and left every published number 6.76 s from the picture. The 0.05 s above is the same construction with a small number.
Every live build states the axis now
MEDIA-SEQUENCE already carries the fact by index; the build now also records the seconds it stands for. The statement is armed once per item attach, through the one funnel every attach passes (NativeAVPlayerHost.load, which every in-place swap delegates to), so a swap path added later inherits it instead of having to remember to arm at its own call site. First build wins, because an item's zero is the FIRST playlist it loaded and later builds of a sliding window state a smaller offset against the very same content.
The error is measured rather than argued
The statement reports what the reconstruction it replaces would have said, on the same line, so the difference is visible in a field log instead of inferable from two logs disagreeing. On the live-only freeze leg:
[+ 0.19s] #446 the reconstruction this replaces had nothing to say yet, so it would have been
latched from a later sample
[+ 0.19s] #454 the playlist this item loaded begins 0.00s into the session
[+118.41s] #446 the reconstruction this replaces would have said 92.02s, 0.00s off the axis the
manifest states
[+118.41s] #454 the playlist this item loaded begins 92.02s into the session
The start item is the case. At the instant the manifest states its axis the reconstruction has no reading at all, because the item has not reported a range yet, so it could only ever have latched from a sample taken later, by which point the window has moved. That is where a nonzero reading on a playlist beginning at zero comes from. On the rejoin item, where the placement branch already stated the axis, the two agree exactly.
Nothing else on that leg moves. Before and after, judged in segments:
VERDICT: live-freeze position held
across the swap: consumed through seg29, resumed into seg28,29,30,31,
never fetched 0 segment(s) in between,
re-fetched 2 segment(s) below it (8.0s, lookback allows 3)
2478 tests / 335 suites green.
Scope
No public API change. An item whose playlist begins at the producer's first segment states 0.00, so a cold live join carries the same number it did before, from a source that cannot be wrong about it.
6.62.1 - An audio track with no decoder costs the probe nothing
An ATSC 3.0 channel tuned through Jellyfin sat at containerOpened for most of a minute with no picture, no error and no way out but backing out. Reported downstream as Sodalite#100 by @classicjazz, whose report arrived with the codec already identified and the failing warning quoted, which is most of the distance to the loop condition below.
One stream can hold the whole probe open
The channel's audio is AC-4, and nothing in this build decodes it: FFmpeg's libavcodec/allcodecs.c names neither AC-4 nor MPEG-H 3D Audio, so there is no build flag to turn on, and Apple's stack has no AC-4 format constant either.
That alone would only mean silence. What it cost was time. has_codec_parameters fails an audio stream with no sample rate, and that value can only come from the container or from opening a decoder. With no decoder compiled in, try_decode_frame gives up on the first packet (found_decoder goes negative), but the outer find_stream_info loop keeps reading regardless, because its only exit is every stream satisfying has_codec_parameters:
for (i = 0; i < ic->nb_streams; i++) {
if (!has_codec_parameters(st, NULL))
break;So one undecodable stream spends the entire 50 MB / 60 s budget and then fails open with the track missing anyway. A file absorbs that as read-ahead. A live source cannot be read ahead of, so it is spent in wall-clock seconds.
The fix
Such a stream is reclassified to AVMEDIA_TYPE_ATTACHMENT for the duration of the probe and restored immediately after, which is the lever the attached-picture fix already uses (#75): has_codec_parameters asks nothing of an attachment, so the pass reaches its "all info found" exit as soon as the real streams are known. The restore has to follow find_stream_info rather than be skipped as redundant, since it writes its internal context back over codecpar on the way out. What a caller ends up with is exactly the state a full-budget probe would have left it, an audio stream whose parameters never resolved, reached without paying for it.
Measured on a synthetic 120 s transport stream built inside the test, since the reported source needs an ATSC 3.0 tuner:
| bytes read by the open | |
|---|---|
| video only | 262,144 |
| video + AC-4, before | 1,572,864 |
| video + AC-4, after | 262,144 |
The 1,572,864 is the 60 s cap landing halfway through a 120 s stream, which is where the reported minute came from.
Deliberately narrow
Four exclusions, and the last one is the reason the predicate asks about decoders rather than about parameters:
AV_CODEC_ID_NONEis a stream still being identified, which is what the probe exists for.- A codec whose parameters the container already carries is never chased in the first place.
- Video is out of scope: the native path decodes formats libavcodec was not built with, ProRes being the standing example.
- A live MPEG-TS AAC stream reaches the probe with
sample_rate == 0and is repaired downstream. Parking it would take the audio off every live channel the engine plays. Having a decoder is what separates the two cases, and a test pins them together.
One correction to the issue
The tracking issue said the pass "always runs until probesize, max_analyze_duration or EOF" on a transport stream, because the early exit is skipped for AVFMTCTX_NOHEADER formats. The first half is true, but mpegts.c clears that flag once every program has received a PMT, under a comment saying exactly why. A healthy channel does take the early exit. This was never a property of live TS; it was one stream holding the door open.
A live-specific probe budget was implemented and reverted. On the same fixture it produced the same reach as the parking does, so there was no case where it demonstrably helped on its own, and shipping it anyway would have been an unproven behaviour change on every live open. LoadOptions.probesize and LoadOptions.maxAnalyzeDuration remain available to any host that wants a tighter live budget today.
No API change
Nothing public changed, hence a patch. A host that wants to name the problem can already read the codec off audioTracks, which is what Sodalite does: it refuses such a channel before a tuner is opened, using the codec Jellyfin names in PlaybackInfo, and says so in words instead of spinning.
Full changelog: 6.62.0...6.62.1
6.62.0 - A panel parked in HDR is asked again once frames are running
@DrHurt reported on #459 that an Apple TV set to HDR output with Match Dynamic Range off and Match Frame Rate on labels every HDR10+/DV session "HDR -> SDR" in Stats for Nerds, and has for months. It is not a labelling slip. The engine genuinely could not tell that panel was in HDR, and the same boolean is the master-versus-media routing gate, so those sessions were also served the bare media playlist with no HDR signaling and no native subtitle renditions.
Both terms of the answer were dead for that panel
currentPanelIsHDR() answers from UIScreen.currentEDRHeadroom, and on tvOS that is not a readout of the panel's mode. It is a transition artifact: raised around a dynamic-range switch and decayed back to 1.00 within about twenty seconds while the panel keeps presenting HDR. panelPresentsHDR therefore trusts it only as a positive and falls back on panelProvenToEngageHDR, a latch meaning "a criteria write has demonstrably driven this display into HDR before".
The flaw is where both are sampled. Every reading is taken around the criteria write, which is the one moment a panel already parked in HDR has nothing to report:
- no dynamic-range transition, so the live reading is 1.00, and
- the latch is armed only by such a reading, so the proof can never be earned either.
So the fallback that exists to survive a decayed headroom cannot help the one configuration that never raises it in the first place. That panel reads SDR forever. It does not reproduce on a box set to 4K SDR with Match Dynamic Range on, where a real switch happens and the headroom rises, which is why this sat for months.
The signal exists, it was read three seconds too early
Device measurement (Apple TV 4K 3rd gen, tvOS 26.5, HDR10 panel, one title played twice):
output fixed to 4K HDR, t+2.5s: headroom 1.20
output fixed to 4K HDR, t+20s: headroom 1.00
output fixed to 4K SDR, t+2.5s: headroom 1.00
The rise comes with HDR content reaching the screen, not with an HDMI mode switch, and it separates exactly the two setups AVPlayer.eligibleForHDRPlayback conflates: a panel parked in HDR from a rate-only panel parked in SDR, both of which report HDR-eligible.
DisplayCriteriaController.probePanelDuringPlayback() samples that window: 250 ms apart, 12 s cap, through the existing observeHeadroom funnel so one reading latches the proof for every later load in the process. Three properties worth naming:
- It samples during steady playback rather than across a mode switch, so it is less exposed to a switch transient than the load-time reads it supplements.
- It does not re-route the running session. The proof it latches makes the next load route correctly on its own, which keeps the change to one published property plus a latch.
- A window that closes with no reading is logged with its max headroom and sample count. That is the one thing no log line could report before: whether a given panel raises the headroom at all.
Still open by construction: the first HDR load of a process on such a panel routes media-direct, because no proof can exist before any playback has happened. Its label corrects itself within seconds and every later load is right. Persisting the latch across launches, cleared by a -11848/-11868 display rejection (the authoritative negative the latch currently lacks), would close that too and is deliberately not in this release.
One funnel for the published format
AetherEngine.presentedVideoFormat now produces videoFormat both at load time and after a late proof, so the two cannot drift apart. It carries the HDR10+ upgrade across: T.35 detection fires while the label still reads SDR, and handleHDR10PlusDetected only upgrades a label that already reads .hdr10, so republishing the bare effective format would have relabelled a proven HDR10+ session "HDR10+ -> HDR10", trading one wrong arrow for a quieter one.
Not attempted
The second half of #459 asks for an "SDR -> HDR" indicator. It is not answerable with today's tvOS API: SDR content composites no HDR so the headroom never rises for it, isDisplayCriteriaMatchingEnabled reports Match Dynamic Range and Match Frame Rate through one combined bit, and AVDisplayManager still exposes no mode read-back (re-checked against the tvOS 26.5 SDK headers). "Box fixed to HDR" and "box on SDR, would switch up for HDR content" are indistinguishable from inside the app while SDR content plays.
Hosts
A host that clamps the engine's videoFormat itself should stop: the engine publishes what the panel presents, and a host-side gate on isDisplayCriteriaMatchingEnabled cannot answer that question and will override this fix. Sodalite's companion removal is Sodalite#101.
swift test: 2459 tests, 333 suites, 0 failures. New suite Issue459PanelProofDuringPlaybackTests (14 tests).
Full notes: CHANGELOG.md - 6.61.0...6.62.0
6.61.0 - An audio language is a reason to serve the master
6.60.0 declared the muxed audio track in the master playlist, which is the only place AVFoundation reads a track language from on an HLS asset. It did not also route there. @htrung14 reported back on #458 that plain SDR H.264 still read "Not Specified", and that is exactly right: the fix landed only on sources that were already being served a master for some other reason.
The routing decision had three reasons, and a language was not one
resolveUseMasterPlaylist served a master for an HDR/DV source on a panel ready for it, for native subtitle renditions (#15), and for tvOS HEVC (#187, where a bare media playlist leaves the device with no HEVC signaling and the item fails with tracks count=0). Everything else went media-direct, and a media playlist has nowhere to put an EXT-X-MEDIA tag. So an SDR H.264 source with no subtitle track kept the original symptom, and so did SDR HEVC on iOS and macOS, where the AE#187 flag is not set.
The 6.60.0 notes described the master-routing paths as "SDR on any panel, HDR / DV on HDR-ready panels". That was too generous: SDR reached a master only through the two forced cases above.
hasAudioRendition is now a fourth reason, the same shape as hasNativeSubs:
- It forces the master only where routing is safe, so an HDR source on an unready panel still routes media-direct. A label is not worth a -11848.
- It is gated on the audio having actually reached the variant, so an audio cascade that fell through to video-only cannot force a master for a group its segments do not carry.
#130still wins: a PQ/HLG variant with noFRAME-RATEis unloadable regardless.
Measured with aetherctl serve --no-dv on a cmn-tagged SDR H.264 fixture (the --no-dv arm matters: with DV mode available, effectiveDvMode inflates sourceIsHDR and a Mac routes to a master for the wrong reason):
6.60.0: serving on .../media.m3u8 useMaster=false videoRange=sdr
6.61.0: serving on .../master.m3u8 useMaster=true videoRange=sdr audioLang=zho
The serving log line now names the language it advertises, so a device log answers "did the rendition go out" without fetching the manifest.
ICU drops the 639-3 codes CLDR aliases
Also from @htrung14: Locale.Language(identifier: "cmn").languageCode?.identifier(.alpha3) is nil, so Mandarin tagged with its ISO 639-3 code hit the fail-closed branch and lost its LANGUAGE, while yue, nan, wuu, hak, gan and hsn resolved. ICU has no alpha3 entry for the 639-3 members CLDR aliases to a macrolanguage.
Locale.canonicalLanguageIdentifier(from:) resolves the alias. It runs behind the direct route and only where that came back empty, so no tag that resolves today changes its answer: canonicalization alone would move no to nob and tl to fil, and both keep nor and tgl. cmn, arb, pes, swh, uzn and kmr now resolve; yue and nan keep their own codes rather than collapsing into zho.
The fallback is gated on a well-formed BCP-47 primary subtag (RFC 5646 2.2.1, two or three ASCII letters), because canonicalization also turns English into en and Japanese into jpn. Free text in a language field is a track name, not a tag, and guessing at it is how a commentary track ends up labelled as a language.
Compatibility
No API change. The behaviour change is the routing one, and it is broad: a consumer whose library is largely SDR H.264 with tagged audio should expect most sources to move from the media playlist to the master. That manifest carries CODECS, RESOLUTION, FRAME-RATE, VIDEO-RANGE, BANDWIDTH and #EXT-X-INDEPENDENT-SEGMENTS, none of which a media-direct session announced. The master is fetched from the engine's own loopback server, so it costs no origin round trip, and MasterFallbackDecision still reverts to the media playlist once per session if AVPlayer rejects it. AirPlay is unaffected: its playlist decision keys on whether the served source is HDR, not on master-ness.
swift test: 2445 tests, 332 suites, 0 failures.
Full diff: 6.60.0...6.61.0
6.60.0 - A muxed audio track is named by the master, not by its mdhd
AVKit labelled the engine's audio track "Not Specified". Reported by @htrung14 (#458), who brought a patch that wrote ISO 639-2/T into the fMP4 audio stream's mdhd, which is where a progressive .mp4 carries it. The symptom is real; the mdhd is not what decides it.
The measurement
Two Matroska fixtures (language=ger and language=fas), served by the engine with aetherctl serve, read back through a real AVPlayerItem on macOS 26:
| what the engine served | what AVFoundation reported |
|---|---|
mdhd only, init.mp4 audio TAG:language=deu |
AVAssetTrack.languageCode = nil, mediaSelectionGroup(.audible) = nil |
plus #EXT-X-MEDIA:TYPE=AUDIO ... LANGUAGE="deu" |
1 audible option, displayName German, extLangTag deu, locale de |
control: the identical mdhd in a progressive .mp4 |
languageCode = deu, option German |
The control row is what makes the first two readable: the probe does see an mdhd when the asset is progressive. For an HLS asset the track language comes from the master playlist, and the engine's master declared no audio rendition at all, because it muxes its one audio track into the variant. There was nothing there to name.
The master declares what it muxes
#EXT-X-MEDIA:TYPE=AUDIO,GROUP-ID="aud",NAME="German",LANGUAGE="deu",DEFAULT=YES,AUTOSELECT=YES
#EXT-X-STREAM-INF:...,AUDIO="aud"
The rendition carries no URI, which is how RFC 8216 4.3.4.2.1 says an audio track that lives inside the variant is declared. AVPlayer exposes it as a one-option audible AVMediaSelection group and labels it from LANGUAGE, not from NAME: NAME="Deutsch" still displays as German, because AVKit localises the language itself. LANGUAGE takes ISO 639-2/T directly, including languages with no alpha-2 form (yue reads back as Cantonese, tgl as Filipino), so no second conversion is needed anywhere.
MP4SegmentMuxer writes the mdhd as well, since the track is that language whoever reads it and a segment pulled out with ffprobe should say so. It is simply not the half that moves the label.
Resolving the source label
AudioLanguageMap is ICU rather than a table. Locale.Language(identifier:).languageCode?.identifier(.alpha3) resolves every language ICU knows plus BCP-47 subtags (pt-BR to por, zh-Hans to zho) and returns nil for xyz, English, Director Commentary or a stray track title, so it doubles as the validity gate. What ICU does not resolve is the twenty ISO 639-2/B bibliographic codes, and Matroska writes exactly those (ger, fre, cze, dut, per, may), so those twenty are mapped first and everything else goes through ICU. A fixed table is always the languages someone thought of; fas is a fixture here because it is one of the ones that would have been missed.
Anything that resolves to nothing writes nothing, so an untagged source produces a byte-identical master to the one it produced before. The rendition is also advertised only when the audio actually reached the variant: an audio cascade that fell through to video-only cannot name a group its segments do not carry.
Compatibility
No API change; AudioLanguageMap and HLSSegmentProvider.masterAudioRendition are internal. Hosts that read the audible media selection group will now find one option where an untagged asset previously offered none, which is the point. Note that the loopback master is only served on the master-routing paths (SDR on any panel, HDR / DV on HDR-ready panels); a media-direct session has no master and is unchanged.
swift test: 2440 tests, 332 suites, 0 failures.
Full diff: 6.59.1...6.60.0
6.59.1 - How much a presentation lead counts is a property of the source
6.56.7 established that a VOD placement composes onto a BASE that sits one presentation lead under the axis, measured it on a fixture whose reorder depth is two frames, and shipped it as arithmetic for every source. It is not arithmetic. Reported by @rrgomes (#418), whose retest split his run by the lead each placement applied and showed the two groups separating cleanly, which is what made the quantity legible enough to reproduce here.
How much a lead counts belongs to the source
Scripts/timecode-fixture.sh now writes a third clip, tc-bf1.mkv, identical to tc-bframes.mkv but for -bf 1: one frame of reorder instead of two, everything else the same. With tc-drought.mkv that is a ladder of three reorder depths differing in nothing else, and the same three-placement burst arm on each, three runs apiece, aetherctl play --picture-probe reading the axis off the picture while the engine reads it out of loadedTimeRanges. The two instruments agreed in all nine runs:
| clip | gate lead | base a composition lands on | the burst arm's three axes |
|---|---|---|---|
tc-drought.mkv |
0.000 | the axis | -9.000 -18.000 -23.000 |
tc-bf1.mkv |
0.042 (one frame) | the axis | -9.000 -18.000 -23.000 |
tc-bframes.mkv |
0.083 (two frames) | one lead below the axis | -9.000 -18.083 -23.166 |
So a source whose gating sample is presented one frame after it is decoded places its composition ON the axis, and one presented two frames after it places a whole lead below. 6.56.7 applied the third row to all of them.
What that cost
On a one-frame-reorder source, which is what the reporting asset is at 23.976 fps, every composition landed two frames from where the reading found it. A placement that opens a run of its own is measured and corrected, so the visible symptom was a correction that fired on every single such placement with the same residual, +0.042 three times across 1450 s of media in the reporter's own log. A placement that opens no run of its own cannot be corrected, and kept the error: on the fixture the unmeasurable third placement was held at -23.042 against a picture reading -23.000.
The fix: read the coefficient off the placement that checks it
A session now starts with no coefficient and composes without one, then takes it from the first placement it can read back. Nothing is assumed before something is measured, and the reading that sets the axis is the same reading that states what a lead was worth:
#418 seg13 placed (advertised 52.000s, worth -9.000s, lead 0.083s x0.00 unmeasured): axis shift -9.000s -> -18.000s from item 61.000s
#418 seg13 says a lead counts 1.00x on this source (axis -9.000s, base measured -9.083s, lead 0.083s)
#418 seg13 placed on base -9.083s, not -9.000s (... residual -0.083s): axis -18.000s -> -18.083s
#418 seg12 placed (advertised 48.000s, worth -5.000s, lead 0.083s x1.00): axis shift -18.083s -> -23.166s from item 66.166s
Every placed line prints the coefficient it used, and x0.00 unmeasured says the session has not read one yet. A reading is allowed to be wrong once by design (6.56.6 accepts that on the grounds the next placement undoes it), so what one teaches the next composition is bounded at two leads.
Compatibility
No API change. HLSVideoEngine.axisShift(after:placing:presentationLead:coefficient:) and placementBase(axis:presentationLead:coefficient:) are internal.
Full diff: 6.59.0...6.59.1