openjoc-api provides the current high-level embeddable interface. It remains
experimental, with decoder and renderer semantics bounded by the current
capability and limitation contracts.
use openjoc_api::{OpenJocConfig, OpenJocPacket, OpenJocSession};
let mut session = OpenJocSession::new(OpenJocConfig::default())?;
let status = session.push_packet(OpenJocPacket {
data: complete_access_unit,
pts_samples: Some(0),
discontinuity: false,
preroll: false,
})?;
while let Some(frame) = session.receive_frame() {
consume_interleaved_f32(frame.interleaved_f32, frame.sample_rate);
}
let _ = session.drain()?;
while let Some(frame) = session.receive_frame() {
consume_interleaved_f32(frame.interleaved_f32, frame.sample_rate);
}
# Ok::<(), Box<dyn std::error::Error>>(())The session is serially accessed. Separate sessions are independent and may run concurrently. Immutable configuration is copied into the session; no process-global decoder, layout, SOFA object, or error buffer exists.
OpenJocPacket is borrowed for the duration of push_packet. data must be
one complete General E-AC-3 JOC access unit: ordered I0/D0..D7 programme sets
covering six cumulative audio blocks (1,536 samples), within the bounded
maximum. A long six-block syncframe is the common one-set form; legal 1/2/3
block syncframes are grouped by cumulative timing and retain per-substream
synthesis state. This matches the JOC metadata, Base/RB alignment, and E-AC-3
inheritance contract. The session does not retain the compressed input buffer.
CMAF and the existing legacy AC-3 Annex-J combination remain D0-only; CMAF
also requires six blocks per syncframe.
Arbitrary byte fragmentation, file paths, MP4/Matroska demuxing, and multiple AUs in one call are intentionally outside this first contract.
The canonical format is interleaved IEEE-754 f32. Each OpenJocPcmFrame
owns its vector and may be retained by a Rust caller. It reports sample rate,
sample count, sample-domain PTS, render mode, layout name, and ordered semantic
channel labels. Speaker layouts use the repository's canonical public presets;
2.0 is FL, FR, 5.1 is FL, FR, FC, LFE, Ls, Rs, 22.2 is the canonical
24-channel FL, FR, FC, LFE1, ... , BtFR order, and binaural reports
Left Ear, Right Ear even when its virtual layout is multichannel.
Physical speaker sessions may also use SpeakerLayout::custom(...) and
OpenJocConfig::with_speaker_layout(...). The custom layout keeps the caller's
speaker array order as PCM/channel order, validates finite spherical geometry,
and keeps LFE channels outside the spatial projector. The JSON/CLI form is
documented in custom speaker layouts; it is
advanced functionality and does not widen downstream host/device channel
layout support.
For binaural sessions, BinauralConfig::builtin_generic("7.1.4") selects the
default offline SADIE II D1/KU100 HRTF. Use
BinauralConfig::builtin(BuiltinHrtf::SadieD2Kemar, "7.1.4") for D2/KEMAR,
and BinauralConfig::from_sofa_bytes(...) for an explicit user SOFA; strict SOFA
validation and fail-closed coverage behavior are unchanged.
output_info() is available before the first packet. Sample rate is None
until the first AU establishes the stream format.
PTS uses the decoded sample domain. If a first packet has PTS P, output for
logical sample n reports P + n; the PTS is not silently moved by the
filterbank or final linked-gain delay. If initial packets omit PTS, the first
later packet with PTS anchors the segment by subtracting the input samples
already decoded. Frames returned before that anchor remain untimestamped;
frames returned afterward, including delayed PCM from earlier packets, use the
inferred origin. Later supplied PTS must match the sample-count continuation;
omitted PTS does not clear an established anchor. An origin or expected packet
PTS outside the signed 64-bit range is rejected before decoding that packet.
An unrepresentable output-frame PTS, including during drain, returns a render
error rather than wrapping or clamping.
Reset, flush, or a discontinuity starts a new segment. This complete-AU API
permits late anchoring; the FFmpeg packet-stream wrapper retains its stricter
untimed-segment contract.
Speaker output reports a 609-sample
delay: the 577-sample QMF/Base-RB delay
plus the admitted 32-sample causal speaker-stage block. Binaural output reports
577 samples for built-in or 48 kHz custom HRIRs because it does not use the
speaker FinalLinkedGain stage. For non-48 kHz custom SOFA, latency_samples
also includes the resampler's common causal filter delay; source Data.Delay
stays in the HRIR and is not added to the reported value. These are public
synchronization contracts; dialnorm and offline static
normalization add zero audio-sample latency.
This makes availability delay explicit without forcing callers to reverse-
engineer it from frame counts.
drain()flushes QMF/reconstruction state and the direct SOFA FIR tail.flush()discards pending PCM and resets stream-derived state while keeping configuration and prepared SOFA data.reset()has the same reusable-session semantics and is the intended seek or discontinuity boundary.- A packet with
discontinuity = trueperforms the stream reset before decode. preroll = trueis accepted to prime decoder state; this first ABI does not suppress the delayed frame automatically.
The output queue is bounded. A caller must receive pending PCM before pushing
another packet; OpenJocStatus::OutputPending is returned otherwise.
openjoc_wave::WaveWriter starts at the sink's current byte position, including
nonzero positions in a file or Cursor. Basic and extensible writers patch
sizes relative to that origin and place any odd-byte PCM padding at the actual
audio-data end. finish() returns the sink positioned immediately after the
WAV, including padding.
The caller must own the span being written: existing bytes in that span are overwritten, not inserted or shifted. Prefix bytes and any existing suffix beyond the completed WAV stay untouched; the writer does not truncate the sink. Offset arithmetic is checked before header/data writes and finalization.
DrcPolicy maps directly to the existing E-AC-3 InternalBasePolicy and
supports disabled, line, RF, and custom boost/cut. DRC changes program
dynamics; it is not a final volume or loudness control. DownmixPolicy
supports auto, Lo/Ro, and Lt/Rt for stereo output. No CLI enum is reused as a
public library type.
OpenJocConfig::dialnorm selects the decoder/program calibration policy:
DialnormMode::Default (calibrated default behavior) is the default and is
recommended for normal playback/decoding. Digital explicitly selects
encoded digital program-level calibration. Analog uses a unity dialnorm
factor and is an advanced compatibility/diagnostic policy; it is not a
recommended louder-output or mastering mode. Dialnorm is separate from
DrcPolicy; DRC remains encoded dynamic-range metadata processing.
The selected dialnorm program scalar is applied once to the complete decoded
program before speaker projection, FinalLinkedGain, or SOFA convolution.
FinalLinkedGain is internal renderer headroom behavior, not a user mastering
control. OpenJocSession never performs file-export peak normalization or any
other file-oriented output transform; applications may apply their own final
gain policy after receiving PCM. The CLI's --normalize-peak is a separate
offline convenience: one static sample-peak gain applied after renderer
processing, not DRC, dialnorm, limiting, compression, LUFS, or true-peak
normalization. The streaming API does not perform file-export peak
normalization or spool a complete program for a file-level transform.
BinauralConfig accepts a complete in-memory SimpleFreeFieldHRIR SOFA buffer,
a virtual speaker layout, and an explicit LFE policy. The session does not
retain a filesystem path. The public API currently uses direct convolution;
partitioned convolution is deferred to a later ABI extension.
Canonical 22.2 is supported with 22 non-LFE virtual sources in the 24-channel layout. Its directions are read from the shared scene topology (including FL/FR at ±52.5°); LFE1/LFE2 are excluded from HRIR preparation. Built-in D1/D2 identity-pose HRIRs and PCM match the static path. Tested pose coverage does not guarantee arbitrary poses or device latency.
OpenJocSession::new_with_listener_orientation_pull(config, pull_samples) is
an explicit, opt-in device-independent 3DoF path. The default constructor and
all existing callers keep static listener orientation. The opt-in constructor
requires binaural mode and accepts 1..=256 output samples per pull. A
ListenerOrientationPreparer can be cloned to a worker thread; it resolves
every non-LFE virtual speaker's HRIR for the complete pose off the render path.
Submit its immutable update between render/pull calls:
use openjoc_api::{
BinauralConfig, ListenerOrientation, OpenJocConfig, OpenJocSession, RenderMode,
};
let mut config = OpenJocConfig::default();
config.render_mode = RenderMode::Binaural;
config.binaural = Some(BinauralConfig::builtin_generic("7.1.4"));
let mut session = OpenJocSession::new_with_listener_orientation_pull(config, 128)?;
let preparer = session.listener_orientation_preparer().expect("enabled");
let epoch = session.listener_orientation_stream_epoch().expect("enabled");
let orientation = ListenerOrientation::new(0.0, 0.0, 0.0, 1.0)?;
let update = preparer.prepare(orientation, epoch, 1)?;
let receipt = session.apply_prepared_listener_orientation(update)
.map_err(|failure| failure.error)?;
drop(receipt.retired_kernels); // Release/recycle away from a real-time callback.
// After push_packet, call receive_binaural_frame repeatedly. A new pose may be
// applied between returned chunks; each frame owns its interleaved f32 samples. An accepted target may wait for an active 240-sample fade to finish before its transition starts.
# Ok::<(), Box<dyn std::error::Error>>(())The quaternion is finite, scale-normalized (x, y, z, w) and represents an
active listener-local-to-scene rotation. Speaker axes are +Y forward, +X
right and +Z up. Each fixed world-space speaker direction is transformed by
the inverse rotation before HRIR lookup. Sequence numbers must increase within
the stream epoch; reset advances that epoch, so an update prepared before reset
is rejected without consuming it. The acceptance receipt reports the accepted
and any superseded sequence plus retired FIR buffers; callers should release
those buffers on a control thread.
Pull mode retains at most one 1,536-sample projected virtual-speaker access
unit before binauralization. The queue count is queryable through
pending_binaural_input_samples. Apply before a pull to make the pose eligible for the next not-yet-rendered chunk; if a 240-sample shared crossfade is already active, the latest accepted target waits until it completes (and one waiting target may be superseded). After drain, reconstruction and FIR tails also arrive as bounded chunks. This changes when existing virtual-speaker PCM is
binauralized; it does not change speaker projection, gain, output channel
order, HRIR tap count, or sample rate. Identity-orientation regression output
is bit-exact with the static path. Non-identity HRIR preparation can still
reject uncovered directions, and overlong HRIR pairs reject the whole update;
no taps are truncated. This API has no sensor/device adapter or hard-real-time,
perceptual, or end-to-end latency guarantee. See known limitations
for measured host-only bounds.
For a runnable synthetic direct-FIR trajectory and waveform comparison, use:
cargo run -p openjoc-api --example listener_orientation_trajectory -- target/orientation-evidence/trajectoryIt writes listener-orientation-trajectory.wav, a fixed-identity comparison
WAV, and a summary with sequence-to-sample receipts. The example synthesizes
independent tones for the 7.1.4 virtual speakers; it is not a JOC access unit,
decoder integration, sensor capture, or headphone/device test. The recorded
run reports an exact-zero identity prefix comparison and a nonzero yaw-interval
RMS delta, demonstrating that the prepared direction updates change the
waveform while leaving the preceding identity segment unchanged.
For frontend parity audits, OpenJocConfig::effective_config_descriptor() and
effective_config_fingerprint() expose the normalized session-boundary fields.
Custom-layout descriptors include channel roles in output order and fixed/named
route vectors sorted by identity, with length-framed identities and exact IEEE-754
gain bits. This intentionally changes existing custom-layout fingerprints;
ordinary preset configurations retain their previous descriptors/fingerprints.
trace_access_units() records each grouped AU's exact byte length, SHA-256,
sample-domain PTS, rate, and independent/dependent frame counts.
Statuses are numeric and non-error lifecycle outcomes: NeedMoreInput,
FrameAvailable, OutputPending, and EndOfStream. Rust errors are typed
OpenJocError values. The C adapter maps them to numeric status codes and an
instance-owned diagnostic message.
Malformed packets, format changes, timestamp discontinuities, profile changes, and render failures are not silently converted into mismatched PCM.
Direct Rust sessions preserve arbitrary validated custom speaker names and geometry.
The FFmpeg bridge (and C openjoc_stream_decoder) instead requires each name to
have an existing OpenJOC-to-AVChannel mapping. For example Ls/SiL map to SL
and Rs/SiR to SR; literal FFmpeg spellings are not automatically accepted.
Mapped identities must be unique, and each explicit LFE/full-range role must
agree with its mapped channel. Unsupported names, alias collisions, or conflicting
roles fail at construction before audio is submitted (InvalidConfig in Rust;
OPENJOC_STATUS_INVALID_ARGUMENT in the C stream API).
Representable custom definitions take precedence over the preset string and use
ordered FFmpeg CUSTOM channels, including when their name resembles a preset or
their channel count differs. PCM order is preserved; the bridge does not claim a
predefined layout or transmit the custom speaker angles. Direct C
openjoc_decoder retains the more general Rust-session layout support.