Skip to content

Latest commit

 

History

History
267 lines (223 loc) · 14.5 KB

File metadata and controls

267 lines (223 loc) · 14.5 KB

Headless Rust streaming API

openjoc-api provides the current high-level embeddable interface. It remains experimental, with decoder and renderer semantics bounded by the current capability and limitation contracts.

Lifecycle

use openjoc_api::{OpenJocConfig, OpenJocPacket, OpenJocSession};

let mut session = OpenJocSession::new(OpenJocConfig::default())?;
let status = session.push_packet(OpenJocPacket {
    data: complete_access_unit,
    pts_samples: Some(0),
    discontinuity: false,
    preroll: false,
})?;
while let Some(frame) = session.receive_frame() {
    consume_interleaved_f32(frame.interleaved_f32, frame.sample_rate);
}
let _ = session.drain()?;
while let Some(frame) = session.receive_frame() {
    consume_interleaved_f32(frame.interleaved_f32, frame.sample_rate);
}
# Ok::<(), Box<dyn std::error::Error>>(())

The session is serially accessed. Separate sessions are independent and may run concurrently. Immutable configuration is copied into the session; no process-global decoder, layout, SOFA object, or error buffer exists.

Input contract

OpenJocPacket is borrowed for the duration of push_packet. data must be one complete General E-AC-3 JOC access unit: ordered I0/D0..D7 programme sets covering six cumulative audio blocks (1,536 samples), within the bounded maximum. A long six-block syncframe is the common one-set form; legal 1/2/3 block syncframes are grouped by cumulative timing and retain per-substream synthesis state. This matches the JOC metadata, Base/RB alignment, and E-AC-3 inheritance contract. The session does not retain the compressed input buffer. CMAF and the existing legacy AC-3 Annex-J combination remain D0-only; CMAF also requires six blocks per syncframe.

Arbitrary byte fragmentation, file paths, MP4/Matroska demuxing, and multiple AUs in one call are intentionally outside this first contract.

Output and layout

The canonical format is interleaved IEEE-754 f32. Each OpenJocPcmFrame owns its vector and may be retained by a Rust caller. It reports sample rate, sample count, sample-domain PTS, render mode, layout name, and ordered semantic channel labels. Speaker layouts use the repository's canonical public presets; 2.0 is FL, FR, 5.1 is FL, FR, FC, LFE, Ls, Rs, 22.2 is the canonical 24-channel FL, FR, FC, LFE1, ... , BtFR order, and binaural reports Left Ear, Right Ear even when its virtual layout is multichannel.

Physical speaker sessions may also use SpeakerLayout::custom(...) and OpenJocConfig::with_speaker_layout(...). The custom layout keeps the caller's speaker array order as PCM/channel order, validates finite spherical geometry, and keeps LFE channels outside the spatial projector. The JSON/CLI form is documented in custom speaker layouts; it is advanced functionality and does not widen downstream host/device channel layout support.

For binaural sessions, BinauralConfig::builtin_generic("7.1.4") selects the default offline SADIE II D1/KU100 HRTF. Use BinauralConfig::builtin(BuiltinHrtf::SadieD2Kemar, "7.1.4") for D2/KEMAR, and BinauralConfig::from_sofa_bytes(...) for an explicit user SOFA; strict SOFA validation and fail-closed coverage behavior are unchanged.

output_info() is available before the first packet. Sample rate is None until the first AU establishes the stream format.

Time, latency, drain, and seek

PTS uses the decoded sample domain. If a first packet has PTS P, output for logical sample n reports P + n; the PTS is not silently moved by the filterbank or final linked-gain delay. If initial packets omit PTS, the first later packet with PTS anchors the segment by subtracting the input samples already decoded. Frames returned before that anchor remain untimestamped; frames returned afterward, including delayed PCM from earlier packets, use the inferred origin. Later supplied PTS must match the sample-count continuation; omitted PTS does not clear an established anchor. An origin or expected packet PTS outside the signed 64-bit range is rejected before decoding that packet. An unrepresentable output-frame PTS, including during drain, returns a render error rather than wrapping or clamping. Reset, flush, or a discontinuity starts a new segment. This complete-AU API permits late anchoring; the FFmpeg packet-stream wrapper retains its stricter untimed-segment contract.

Speaker output reports a 609-sample delay: the 577-sample QMF/Base-RB delay plus the admitted 32-sample causal speaker-stage block. Binaural output reports 577 samples for built-in or 48 kHz custom HRIRs because it does not use the speaker FinalLinkedGain stage. For non-48 kHz custom SOFA, latency_samples also includes the resampler's common causal filter delay; source Data.Delay stays in the HRIR and is not added to the reported value. These are public synchronization contracts; dialnorm and offline static normalization add zero audio-sample latency. This makes availability delay explicit without forcing callers to reverse- engineer it from frame counts.

  • drain() flushes QMF/reconstruction state and the direct SOFA FIR tail.
  • flush() discards pending PCM and resets stream-derived state while keeping configuration and prepared SOFA data.
  • reset() has the same reusable-session semantics and is the intended seek or discontinuity boundary.
  • A packet with discontinuity = true performs the stream reset before decode.
  • preroll = true is accepted to prime decoder state; this first ABI does not suppress the delayed frame automatically.

The output queue is bounded. A caller must receive pending PCM before pushing another packet; OpenJocStatus::OutputPending is returned otherwise.

Seekable WAV output

openjoc_wave::WaveWriter starts at the sink's current byte position, including nonzero positions in a file or Cursor. Basic and extensible writers patch sizes relative to that origin and place any odd-byte PCM padding at the actual audio-data end. finish() returns the sink positioned immediately after the WAV, including padding.

The caller must own the span being written: existing bytes in that span are overwritten, not inserted or shifted. Prefix bytes and any existing suffix beyond the completed WAV stay untouched; the writer does not truncate the sink. Offset arithmetic is checked before header/data writes and finalization.

Policies

DrcPolicy maps directly to the existing E-AC-3 InternalBasePolicy and supports disabled, line, RF, and custom boost/cut. DRC changes program dynamics; it is not a final volume or loudness control. DownmixPolicy supports auto, Lo/Ro, and Lt/Rt for stereo output. No CLI enum is reused as a public library type.

OpenJocConfig::dialnorm selects the decoder/program calibration policy: DialnormMode::Default (calibrated default behavior) is the default and is recommended for normal playback/decoding. Digital explicitly selects encoded digital program-level calibration. Analog uses a unity dialnorm factor and is an advanced compatibility/diagnostic policy; it is not a recommended louder-output or mastering mode. Dialnorm is separate from DrcPolicy; DRC remains encoded dynamic-range metadata processing.

The selected dialnorm program scalar is applied once to the complete decoded program before speaker projection, FinalLinkedGain, or SOFA convolution. FinalLinkedGain is internal renderer headroom behavior, not a user mastering control. OpenJocSession never performs file-export peak normalization or any other file-oriented output transform; applications may apply their own final gain policy after receiving PCM. The CLI's --normalize-peak is a separate offline convenience: one static sample-peak gain applied after renderer processing, not DRC, dialnorm, limiting, compression, LUFS, or true-peak normalization. The streaming API does not perform file-export peak normalization or spool a complete program for a file-level transform.

BinauralConfig accepts a complete in-memory SimpleFreeFieldHRIR SOFA buffer, a virtual speaker layout, and an explicit LFE policy. The session does not retain a filesystem path. The public API currently uses direct convolution; partitioned convolution is deferred to a later ABI extension.

Experimental listener orientation

Canonical 22.2 is supported with 22 non-LFE virtual sources in the 24-channel layout. Its directions are read from the shared scene topology (including FL/FR at ±52.5°); LFE1/LFE2 are excluded from HRIR preparation. Built-in D1/D2 identity-pose HRIRs and PCM match the static path. Tested pose coverage does not guarantee arbitrary poses or device latency.

OpenJocSession::new_with_listener_orientation_pull(config, pull_samples) is an explicit, opt-in device-independent 3DoF path. The default constructor and all existing callers keep static listener orientation. The opt-in constructor requires binaural mode and accepts 1..=256 output samples per pull. A ListenerOrientationPreparer can be cloned to a worker thread; it resolves every non-LFE virtual speaker's HRIR for the complete pose off the render path. Submit its immutable update between render/pull calls:

use openjoc_api::{
    BinauralConfig, ListenerOrientation, OpenJocConfig, OpenJocSession, RenderMode,
};

let mut config = OpenJocConfig::default();
config.render_mode = RenderMode::Binaural;
config.binaural = Some(BinauralConfig::builtin_generic("7.1.4"));
let mut session = OpenJocSession::new_with_listener_orientation_pull(config, 128)?;
let preparer = session.listener_orientation_preparer().expect("enabled");
let epoch = session.listener_orientation_stream_epoch().expect("enabled");
let orientation = ListenerOrientation::new(0.0, 0.0, 0.0, 1.0)?;
let update = preparer.prepare(orientation, epoch, 1)?;
let receipt = session.apply_prepared_listener_orientation(update)
    .map_err(|failure| failure.error)?;
drop(receipt.retired_kernels); // Release/recycle away from a real-time callback.
// After push_packet, call receive_binaural_frame repeatedly. A new pose may be
// applied between returned chunks; each frame owns its interleaved f32 samples. An accepted target may wait for an active 240-sample fade to finish before its transition starts.
# Ok::<(), Box<dyn std::error::Error>>(())

The quaternion is finite, scale-normalized (x, y, z, w) and represents an active listener-local-to-scene rotation. Speaker axes are +Y forward, +X right and +Z up. Each fixed world-space speaker direction is transformed by the inverse rotation before HRIR lookup. Sequence numbers must increase within the stream epoch; reset advances that epoch, so an update prepared before reset is rejected without consuming it. The acceptance receipt reports the accepted and any superseded sequence plus retired FIR buffers; callers should release those buffers on a control thread.

Pull mode retains at most one 1,536-sample projected virtual-speaker access unit before binauralization. The queue count is queryable through pending_binaural_input_samples. Apply before a pull to make the pose eligible for the next not-yet-rendered chunk; if a 240-sample shared crossfade is already active, the latest accepted target waits until it completes (and one waiting target may be superseded). After drain, reconstruction and FIR tails also arrive as bounded chunks. This changes when existing virtual-speaker PCM is binauralized; it does not change speaker projection, gain, output channel order, HRIR tap count, or sample rate. Identity-orientation regression output is bit-exact with the static path. Non-identity HRIR preparation can still reject uncovered directions, and overlong HRIR pairs reject the whole update; no taps are truncated. This API has no sensor/device adapter or hard-real-time, perceptual, or end-to-end latency guarantee. See known limitations for measured host-only bounds.

For a runnable synthetic direct-FIR trajectory and waveform comparison, use:

cargo run -p openjoc-api --example listener_orientation_trajectory -- target/orientation-evidence/trajectory

It writes listener-orientation-trajectory.wav, a fixed-identity comparison WAV, and a summary with sequence-to-sample receipts. The example synthesizes independent tones for the 7.1.4 virtual speakers; it is not a JOC access unit, decoder integration, sensor capture, or headphone/device test. The recorded run reports an exact-zero identity prefix comparison and a nonzero yaw-interval RMS delta, demonstrating that the prepared direction updates change the waveform while leaving the preceding identity segment unchanged.

For frontend parity audits, OpenJocConfig::effective_config_descriptor() and effective_config_fingerprint() expose the normalized session-boundary fields. Custom-layout descriptors include channel roles in output order and fixed/named route vectors sorted by identity, with length-framed identities and exact IEEE-754 gain bits. This intentionally changes existing custom-layout fingerprints; ordinary preset configurations retain their previous descriptors/fingerprints. trace_access_units() records each grouped AU's exact byte length, SHA-256, sample-domain PTS, rate, and independent/dependent frame counts.

Errors and status

Statuses are numeric and non-error lifecycle outcomes: NeedMoreInput, FrameAvailable, OutputPending, and EndOfStream. Rust errors are typed OpenJocError values. The C adapter maps them to numeric status codes and an instance-owned diagnostic message.

Malformed packets, format changes, timestamp discontinuities, profile changes, and render failures are not silently converted into mismatched PCM.

Custom-layout transport boundary

Direct Rust sessions preserve arbitrary validated custom speaker names and geometry. The FFmpeg bridge (and C openjoc_stream_decoder) instead requires each name to have an existing OpenJOC-to-AVChannel mapping. For example Ls/SiL map to SL and Rs/SiR to SR; literal FFmpeg spellings are not automatically accepted. Mapped identities must be unique, and each explicit LFE/full-range role must agree with its mapped channel. Unsupported names, alias collisions, or conflicting roles fail at construction before audio is submitted (InvalidConfig in Rust; OPENJOC_STATUS_INVALID_ARGUMENT in the C stream API).

Representable custom definitions take precedence over the preset string and use ordered FFmpeg CUSTOM channels, including when their name resembles a preset or their channel count differs. PCM order is preserved; the bridge does not claim a predefined layout or transmit the custom speaker angles. Direct C openjoc_decoder retains the more general Rust-session layout support.