Skip to content

Latest commit

 

History

History
335 lines (263 loc) · 16.9 KB

File metadata and controls

335 lines (263 loc) · 16.9 KB

Audio-V external validation corpus

Purpose

The Audio-V Validation Corpus measures how Oracle Engine rules behave on legally acquired reference recordings whose dataset identity and processing history are documented. It does not require private studio masters, artist relationships, or recordings created by the project maintainer. It also does not turn a spectral heuristic into proof of a file’s history.

The corpus establishes the tested population, controlled transformations, false-positive and false-negative rates, confidence intervals, and limitations behind each published claim.

Evidence tiers

  1. Synthetic regression catches deterministic implementation regressions. It is never counted as real-world calibration.
  2. Public external references come from datasets whose license permits reproducible transformation and corpus redistribution, while remaining excluded from the desktop package.
  3. External challenge references are separately acquired research or test material governed by their original terms. They remain local and never count toward the public-reference target.
  4. Controlled derivatives are generated by Audio-V from an eligible reference. Their transformation history is exact, but they never increase the independent-reference count.

Initial five-source plan

The versioned registry at validation/real-world/external-datasets.json is the authority for URLs, versions, access rules, licenses, intended use, and limitations.

Priority Dataset Corpus role Public target? Access
1 Slakh2100 Redux Reproducible 44.1-kHz/16-bit FLAC transformation baseline Yes, CC BY 4.0 Direct, approximately 104 GB
2 MUSAN Music, speech, and noise false-positive controls Yes, CC BY 4.0 Direct, approximately 11 GB
3 MAESTRO v3 Natural recorded-piano challenge population No, CC BY-NC-SA Direct, approximately 101 GB compressed
4 EBU SQAM Standardized codec and audio-system test material No, local testing under EBU terms Manual terms review
5 MUSDB18-HQ Multi-genre, full-bandwidth academic challenge population No, educational/academic and per-track terms Authorized direct download, approximately 22.7 GB

Slakh2100 and MUSAN can advance the reproducible public pilot. MAESTRO, EBU SQAM, and MUSDB18-HQ broaden real-world testing without being silently relicensed or shipped through Audio-V.

Reference eligibility

An external recording is eligible only when:

  • its exact dataset, version, source identity, official URL, and license are recorded;
  • the operator obtained it from the named dataset rather than an unverified mirror;
  • its SHA-256, acquisition date, format, sample rate, bit depth, and channel layout are recorded;
  • every related source and derivative stays in one development, calibration, test, or challenge partition;
  • restricted material is marked nonredistributable and challenge-only; and
  • the recording is not called an “original studio master” merely because it is stored in a lossless container.

CC0 1.0 and CC BY 4.0 references may satisfy the public target when attribution is preserved. NonCommercial, ShareAlike, academic-use, approval-gated, and test-material collections remain separately identified local challenge sets.

Ground-truth model

Ground truth records facts:

  • external dataset and source-recording identity;
  • integrity state;
  • known transformation graph;
  • codec and encoding parameters;
  • sample-rate and word-length changes;
  • intentional filtering;
  • clipping and level changes;
  • channel transformations; and
  • the Oracle outcome allowed by the disclosed evidence.

For transformed cases, the reference recording is the starting point of a controlled experiment. Audio-V does not claim that it is the artist’s studio master.

Every derivative retains the independent-reference ID, source group, dataset ID, transformation recipe, tool version and arguments, and SHA-256.

Leakage prevention

Development, calibration, and test partitions are assigned at the independent source-group level. Every derivative of one recording stays in the same partition. MAESTRO performances of the same composition share one group. Restricted datasets are challenge-only.

Thresholds may be explored with development data, frozen with calibration data, and evaluated once per candidate Oracle Engine version against locked test and challenge sets. Derivatives of a development reference may never appear in calibration or test.

Milestones

  • Pilot threshold: 25 independent public external references.
  • Target threshold: 50 independent public external references.
  • Controlled transformations: at least 10 cases per eligible reference.
  • Required partitions: development, calibration, and test.
  • Required evidence: native controls, lossy-to-lossless transcodes, upsampling, intentional low-pass confounders, clipping variants, channel variants, and deterministic corruption fixtures.

The state is external references pending before import, pilot building below 25 public references, pilot ready from 25–49 with every required partition, and target corpus ready from 50 with every required partition. None of those states alone means probability calibration.

Current measured milestone

Corpus 0.5.0 contains 85 independent references: 20 public MUSAN edge controls, 30 public Slakh2100 Redux controlled sources, 10 local EBU SQAM challenge references, 15 local MAESTRO composition-unique piano challenge references, and 10 local MUSDB18-HQ full-bandwidth mixtures balanced between the official train and test partitions. The public population has 30 development, 10 calibration, and 10 held-out test source groups and reaches the numeric 50-reference target.

Oracle Engine 0.10.0-oracle-v12 evaluated 2,295 generated cases:

  • 1,735 scored cases and 560 observational-only origin-positive cases;
  • 1,522/1,735 accepted scored outcomes (87.72% case-weighted, descriptive);
  • 55/85 source groups with every scored case accepted;
  • zero lossy-origin false advisories across 975 eligible negative cases;
  • zero upsample false advisories across 1,495 eligible negative cases;
  • zero explicit positive advisories across 360 controlled lossy derivatives;
  • zero explicit positive advisories across 120 controlled upsample derivatives;
  • 540/540 safe MUSAN edge-control abstentions; and
  • 110/110 accepted scored MUSDB18-HQ challenge outcomes, with 160 additional positive-origin transformations retained as observational-only.

This is a useful, intentionally uncomfortable result. It demonstrates strong specificity and conservative abstention on the disclosed population, but it also establishes a severe positive-detection coverage gap. Reaching the source count target does not make the detector probability-calibrated or authoritative. The current aggregate scorecard is published at validation/real-world/scorecards/0.10.0-oracle-v12-corpus-0.5.0.json; earlier scorecards remain immutable historical measurements.

Derivative case counts are reported for debugging and recipe-level comparison, but they are not treated as independent observations. Headline uncertainty is calculated over independent source groups, and a group passes only when every scored generated case for that source passes.

Challenge sources labeled negative-only can measure false origin advisories, but they cannot establish detector sensitivity because their own prior processing history is not controlled. Lossy-transcode and upsample derivatives made from those references are therefore retained as observational-only cases: Oracle Engine still analyzes them and the scorecard exposes the result, but they are excluded from acceptance totals rather than being mislabeled as failures. Deterministic integrity, channel, and clipping cases remain scored. All exact-binomial intervals are calculated over the explicitly named independence unit.

Operator workflow

First write and inspect the source plan:

npm run corpus:plan

For a directly hosted archive, use the resumable segmented downloader. It preallocates the registry-pinned byte count, uses bounded HTTP range workers, records completed segments in ignored local state, and refuses to finish unless the published checksum matches:

npm run corpus:download -- \
  --dataset slakh2100 \
  --concurrency 8 \
  --parts 48

Concurrency is capped at 16 so corpus acquisition cannot create an unbounded network workload. An interrupted command resumes completed segments when it is run again with the same part count.

For Slakh, validate the archive and selectively extract a deterministic candidate pool. This reads the complete archive index but extracts only Redux mix.flac entries from the train, validation, and test directories; stems and the omitted duplicate-MIDI directory are excluded:

npm run corpus:extract -- \
  --dataset slakh2100 \
  --file "/external/datasets/downloads/slakh2100_flac_redux.tar.gz" \
  --root "/external/datasets/slakh2100-selected" \
  --candidate-limit 120

MAESTRO uses the same command with its ZIP archive. Audio-V reads the official column-oriented metadata, groups repeat performances by composition, extracts the metadata plus a deterministic candidate pool, and leaves all other WAV and MIDI files in the archive:

npm run corpus:extract -- \
  --dataset maestro-v3 \
  --file "/external/datasets/downloads/maestro-v3.0.0.zip" \
  --root "/external/datasets/maestro-selected" \
  --candidate-limit 60

MUSDB18-HQ also uses the selective ZIP path. Audio-V balances the selection between the official train and test partitions and extracts only mixture.wav plus any archive-supplied readme/license records; isolated source stems are not copied into the validation workspace. The current official archive contains no readme/license file, so import requires a locally saved, hashed copy of the official Zenodo record as terms evidence:

npm run corpus:extract -- \
  --dataset musdb18-hq \
  --file "/external/datasets/downloads/musdb18hq.zip" \
  --root "/external/datasets/musdb18-hq-selected" \
  --candidate-limit 40

Alternatively, download a dataset from the official URL recorded in the generated plan. Extract it outside the tracked repository and preview the deterministic selection:

Before extraction, verify the downloaded byte count and the publisher checksum when one is available. Audio-V always calculates SHA-256 and writes an ignored local acquisition receipt:

npm run corpus:verify-download -- \
  --dataset musan \
  --file "/external/datasets/downloads/musan.tar.gz"

If an exact-byte mirror was used only as the transfer transport, include --transport-url. The archive is accepted solely by the official registry byte count and checksum, while the receipt preserves the actual chain of custody instead of implying it came directly from the publisher.

npm run corpus:import:dataset -- \
  --dataset slakh2100 \
  --root "/external/datasets/slakh2100_flac_redux" \
  --limit 30 \
  --corpus-version 0.3.0 \
  --dry-run true

Import the selected references:

npm run corpus:import:dataset -- \
  --dataset slakh2100 \
  --root "/external/datasets/slakh2100_flac_redux" \
  --limit 30 \
  --corpus-version 0.3.0

npm run corpus:import:dataset -- \
  --dataset musan \
  --root "/external/datasets/musan" \
  --limit 20 \
  --corpus-version 0.3.0

Research-only challenge sources require an explicit terms acknowledgement and a local evidence file (for example, the publisher-supplied license or saved access approval). Audio-V hashes the evidence and attaches it to every imported reference; the path itself is not published:

npm run corpus:import:dataset -- \
  --dataset maestro-v3 \
  --root "/external/datasets/maestro-selected/maestro-v3.0.0" \
  --limit 15 \
  --corpus-version 0.4.0 \
  --terms-accepted true \
  --terms-evidence "/external/datasets/maestro-license.txt"

EBU SQAM is registry-pinned to the EBU QC download byte count and remains a challenge-only R&D source. Its machine-readable record states that downloading the material accepts noncommercial use except as an R&D tool. Preserve that record as the --terms-evidence file and do not package or republish the audio. Audio-V maps the handbook's numbered ranges into alignment, artificial, single-instrument, vocal, speech, solo-instrument, vocal/orchestra, orchestra, and pop strata before deterministic challenge selection.

The importer uses hard links when the dataset and project are on the same filesystem, then falls back to copying. Use --storage copy to force an independent local copy. It probes every selected recording, records SHA-256, assigns a leakage-safe partition, and refuses duplicate dataset/source identities. MUSAN imports also locate and hash the nearest component LICENSE file so its mixed-source attribution evidence stays attached to each selected recording.

Every import requires an explicit --corpus-version. A changed reference population must receive a new semantic version so a published scorecard can never silently acquire a different meaning. Additional imports intentionally belonging to the same unpublished iteration may repeat that current version.

Large corpora may live entirely outside the repository. Set AUDIO_V_VALIDATION_CORPUS_DIRECTORY to a directory containing copies of corpus.json, recipes.json, and external-datasets.json; imported references, derivatives, and generated manifests will then remain there. This is the preferred configuration when the datasets reside on a large external or network volume. Each source contributes one deterministic window of at most 60 seconds, providing enough decoded material for stable spectral evidence while bounding derived storage and evaluation time.

Restricted sources require an explicit acknowledgment after the operator reviews their terms:

npm run corpus:import:dataset -- \
  --dataset maestro-v3 \
  --root "/external/datasets/maestro-v3.0.0" \
  --limit 15 \
  --terms-accepted true

Finally generate controlled derivatives and evaluate Oracle Engine:

npm run corpus:generate
npm run validate:real-world
npm run corpus:publish-scorecard

Audio-V’s bundled LGPL FFmpeg engine does not include libmp3lame. AAC and Opus transformations use the bundled engine. To include MP3 recipes, set AUDIO_V_CORPUS_MP3_FFMPEG_PATH to a rights-reviewed FFmpeg executable exposing libmp3lame. Missing MP3 capability is recorded as skipped evidence rather than silently substituted.

Evaluation contract

Scorecards report:

  • independent references and source groups;
  • cases by dataset, source category and subcollection stratum, partition, recipe, and transformation class;
  • accepted and observed Oracle classifications;
  • false-positive, false-negative, and inconclusive counts;
  • sensitivity, specificity, and inconclusive rate;
  • exact two-sided 95% binomial confidence intervals; and
  • corpus, recipe, dataset-registry, and Oracle Engine versions.

Production language remains “possible,” “bandwidth limited,” or “inconclusive” until the target corpus and challenge evaluation justify a stronger claim. Synthetic fixtures and multiple derivatives of one reference never inflate the independent-reference count.

Origin candidate promotion

validation/real-world/origin-promotion-policy.json is the machine-readable gate for any proposed lossy-history or upsample detector. Development fits a candidate, calibration selects its operating threshold, and test is opened once for the promotion decision. EBU SQAM, MAESTRO, and MUSDB18-HQ independently veto non-generalizing false advisories; MUSAN is an all-negative edge-abstention population. Source group—not derivative file—is the headline independence unit.

A model score may not be displayed as probability without separate probability calibration evidence. No origin candidate may produce Failed, promote a file to Clear, or bypass abstention. A failed gate preserves Oracle production behavior and publishes the negative result. Once a held-out population is opened, it may not be reused as an untouched promotion set for a retuned candidate. For the statistical reason finite validation cannot make calibration self-evident, see Metrics of Calibration for Probabilistic Predictions.

Repository and package boundaries

validation/
  fidelity-corpus.json
  real-world/
    corpus.json
    external-datasets.json
    recipes.json
    schemas/
    scorecards/              tracked aggregate results without source audio
    .audio/                 ignored external references and derivatives
    generated/              ignored generated case manifest
build/
  corpus-acquisition-plan-latest.json
  real-world-validation-latest.json
  real-world-scorecard-latest.json

External audio is deliberately excluded from Git and from every DMG, installer, and portable build. The repository contains the dataset registry, official links, checksums where published, selection rules, transformation recipes, and aggregate scorecards—not the source recordings.