This file captures two things:
- what Tagminder already does today (implemented checks/capabilities)
- what is still outstanding (prioritized backlog)
Scope is intentionally offline, dataset-only checks using fields already retained (see tagminder.toml [cleanup].keep_columns).
Guiding principles:
- No online lookups required.
- Prefer checks that are high-signal, actionable, and cheap to compute over 800k+ rows.
- Where possible, phrase checks as: Condition → Why it matters → Suggested remediation.
-
Track sequence anomaly reporting exists
- Implemented by
scripts/reports/93-report-track-sequence-anomalies-by-album.py. - Covers missing/invalid/duplicate/gapped track numbering behavior at album/disc scope.
- Implemented by
-
Missing critical tags reporting exists
- Implemented by
scripts/reports/94-report-missing-critical-tags-by-album.py. - Includes ReplayGain coverage checks used for album-level quality visibility.
- Implemented by
-
ReplayGain coverage consistency
- Implemented via
scripts/reports/94-report-missing-critical-tags-by-album.pyby including ReplayGain fields in critical-column checks.
- Implemented via
-
Duplicate detection reports exist
- Implemented by
scripts/reports/96-report-duplicate-tracks-all.pyandscripts/reports/97-report-duplicate-albums.py. - Supports duplicate discovery at both track and album levels.
- Implemented by
-
Chunk/full semantics aligned (idempotency-first)
- Step 18 behavior now preserves deterministic outcomes regardless of processing mode.
-
All synthetic outcomes persisted to decisions table
- Synthetic assignments are written to
_USR_disambiguation_decisions.
- Synthetic assignments are written to
-
Synthetic rows not written to main disambiguated lookup table
- Policy enforced: do not insert synthetic IDs into
contributors_unified_disambiguated.
- Policy enforced: do not insert synthetic IDs into
-
Decision provenance captured
_USR_disambiguation_decisionsincludesdecision_source(with migration support).
-
Primary user guide added and linked
docs/user-guide.mdcreated and cross-linked with README for navigation.
-
Import performance guidance expanded
- Multi-drive concurrent ingest guidance documented, including anti-thrashing notes.
-
Export behavior made explicit
- README documents that export is intentionally serialized (one file at a time).
-
No persistent SQLite WAL sidecars after TUI exit
- Implemented in the TUI quit path (best-effort WAL checkpoint/truncate + switch back to
journal_mode=DELETE). - Why: for users treating the staging DB as the metadata master/backup, leftover WAL sidecars are confusing and can feel like an unclean shutdown.
- Implemented in the TUI quit path (best-effort WAL checkpoint/truncate + switch back to
-
Cross-database metadata sync by
track_uuid- Implemented via
scripts/export/98-sync-metadata-by-track-uuid.py. - Applies selected metadata updates from source DB to target DB by
track_uuid, updates only changed values, increments__sqlmodded, and writes target changelog entries.
- Implemented via
-
Synthetic MBID retirement workflow
- Implemented via
scripts/pipeline/23-retire-synthetic-mbids.py. - Uses normalized contributor name + context matching (no name-only auto replacement), dry-run by default, and apply mode for confirmed replacements.
- When applying, synthetic->real replacements are propagated across all
musicbrainz_*idcolumns inalib,__sqlmoddedis incremented,changelogis written, and_USR_disambiguation_decisionsis updated.
- Implemented via
Core identity fields: track, tracknumber, disc, discnumber.
- Track/disc numbering contradictions (beyond existing sequence anomaly report)
- Flag when
trackandtracknumberdisagree materially. - Flag when
discis present butdiscnumberis missing (or vice versa), or non-numeric. - Why: breaks ordering, multi-disc grouping, and player UI expectations.
- Flag when
Applies to: year, date, releasedate, originaldate, originalyear, originalreleasedate, recording_date, recordingstartdate, recordingenddate, performancedate.
-
Cross-field contradictions
- Examples:
yearconflicts with year extracted fromreleasedate;originalyear>year. - Why: breaks chronology, "original vs reissue" logic, and library browse.
- Examples:
-
Invalid or placeholder dates
- Flag clearly invalid values, impossible ranges, placeholders (e.g.,
0000, epoch-like placeholders if you use them). - Why: placeholders pollute sorting and can be mistaken for real data.
- Flag clearly invalid values, impossible ranges, placeholders (e.g.,
-
Recording range sanity
- Flag when
recordingenddate<recordingstartdate, or when ranges exist but the mainrecording_dateconflicts.
- Flag when
Applies to: work, movement, part, composer, conductor, orchestra, ensemble, performer, movementname (if retained).
-
Work/movement coherence
- Flag when
movementis present butworkis missing. - Flag when
workis present butcomposeris missing.
- Flag when
-
Movement field drift
- If both
movementandmovementnameexist in your retained dataset, flag divergence (one present without the other, or conflicting values). - Why: classical browsing depends on stable work/movement semantics.
- If both
-
Ensemble/orchestra/performer overlap anomalies
- Flag when identical entities appear across multiple of
ensemble,orchestra,performerin ways that violate your chosen tagging model.
- Flag when identical entities appear across multiple of
Applies to: lyrics, unsyncedlyrics, explicit.
-
Unsynced lyrics leftover
- If your policy is "move
unsyncedlyrics→lyricswhenlyricsempty", flag rows whereunsyncedlyricsremains populated after the cleanup stage.
- If your policy is "move
-
Explicit value domain consistency
- Flag mixed representations (e.g.,
0/1mixed withClean/Explicitor vendor-specific codes).
- Flag mixed representations (e.g.,
Applies to: genre, style, mood, theme.
-
Over-broad genre residue
- Track counts of generic buckets (e.g.,
Pop,Pop/Rock,Jazz,Classical) post-enrichment. - Why: tells you whether enrichment/normalization is paying off.
- Track counts of generic buckets (e.g.,
-
Style-in-genre leakage
- If you sometimes merge style→genre, flag records where
styleis empty butgenrelooks like it contains many style tokens (or vice versa).
- If you sometimes merge style→genre, flag records where
Applies to: isrc, upc, barcode, asin, catalog, catalognumber (if present), musicbrainz_*, acoustid_*, discogs_*, songkong_id, roonid, itunesalbumid, itunesartistid.
-
ISRC collisions
- Flag when the same
isrcappears across materially differenttitle/artistcombinations. - Why: indicates tag collisions or mis-assignment.
- Flag when the same
-
Discogs URL/ID inconsistencies
- Flag when
discogs_release_urlis present butdiscogs_release_idis missing (or malformed), and similarly for master release.
- Flag when
Applies to: __md5sig, __length_seconds, __file_size_bytes, plus ReplayGain fields.
- Audio-stream content hashing for formats without embedded MD5
- Consider implementing an audio-stream MD5 (or similar stable digest) for file formats that don't natively provide an embedded content hash.
- Implementation approach: Rust Symphonia decoder wrapped via PyO3, producing a deterministic digest over decoded PCM frames.
- Why: enables "exact duplicates by audio content" detection beyond FLAC/WavPack embedded MD5 coverage.
- Prefer generating a report with:
- counts per issue type
- top examples (sample
__path+ key fields) - a stable issue key so you can track "fixed vs remaining" over time
- For date checks, decide whether you allow partial dates (
YYYY,YYYY-MM) and treat them consistently.
- Synthetic-retirement logic should remain opt-in and review-first (no automatic replacements).
- Cross-database sync should always be changeloged and idempotent across reruns.