Skip to content

Commit e539856

Browse files
Jamestthclaude
andauthored
Restructure example notebooks and add cross-table merge support (#45)
* Restructure example notebooks and add cross-table merge support Split the monolithic examples/vowl_usage_patterns_demo.ipynb (91 cells) into three focused, self-contained notebooks under their own folders: - 1_core_tutorial/ setup, running a validation, understanding results - 2_multiple_sources/ validating one contract across multiple sources - 3_real_databases/ server-side validation with Testcontainers Each notebook resolves the shared dataset paths on its own (repo-root walk-up), imports what it needs, and writes generated artifacts to a local outputs/ folder. Section numbering/titles are scoped per notebook, and doc links in README.md and docs/usage-patterns.md are updated to the new paths. Also lands the cross-table-merge annotated-output work: a subquery-projected referential check now merges onto its anchor table instead of becoming a residue (src/vowl/validation/result.py), with expanded tests and expected outputs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Fix markdown formatting to satisfy prettier lint Apply `prettier --write` to the three files the lint CI job flagged: table column-padding and `*emphasis*` → `_emphasis_` normalization. No content changes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Deprecate consolidated failed-rows output in favour of annotated output Mark get_consolidated_output_dfs() and output_mode="failed_rows"/"both" as deprecated, steering users to get_annotated_output() / output_mode="annotated". - get_consolidated_output_dfs() now emits a DeprecationWarning and delegates to a private _get_consolidated_output_dfs() so internal callers (save() in failed_rows/both mode) reuse the grouping without warning. - save() warns on the deprecated paths: implicit default (upcoming flip to "annotated"), explicit "failed_rows", and the failed-rows half of "both". "annotated" stays silent. All warnings use stacklevel=2. - ValidationConfig.output_mode docstring notes the default will change. - Tests assert the warnings fire and that the private helper stays silent; internal/golden callers switched to the private helper. - CHANGELOG Deprecated entry; README/getting-started/known-issues updated. - Notebook TOC anchor-link fixes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Split core tutorial into basic + advanced usage notebooks Rename 1_core_tutorial/core_tutorial.ipynb to 1_basic_tutorial/basic_tutorial.ipynb and move the two denser sections (Explicitly Defined Adapter incl. PooledAdapter, and Filtering Rows Before Validation) into a new 4_advanced_usage/advanced_usage.ipynb so the basic tutorial stays focused on the everyday workflow. - Renumber the basic tutorial sections and fix its Contents/anchors - Link the "cross-table check that merges" note out to known-issues instead of duplicating the mechanics - Update cross-references in README, docs, examples/README, and the multiple-sources / real-databases notebooks - Replace non-ASCII typographic symbols (em-dashes, arrows, ellipses) with ASCII equivalents in both notebooks - Re-execute both notebooks top-to-bottom so outputs and execution counts are consistent Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Tighten annotated-output prose in known-issues Clarify which non-mergeable checks produce residues vs. appear only in summary.json, and replace em-dash asides with plainer punctuation. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Regenerate example outputs from a clean re-run Delete and regenerate the outputs/ artifacts for the basic tutorial and multiple-sources notebooks so nothing is stale. Re-executing both notebooks top-to-bottom also removes an orphaned join-output CSV that the current multi-source run no longer produces. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Drop deprecated consolidated-output section from basic tutorial Remove the "Consolidated Failed Rows (per table)" section, which demonstrates the deprecated get_consolidated_output_dfs() accessor. The annotated output is now the recommended per-table view. Re-execute the notebook so execution counts stay contiguous. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Drop deprecated get_consolidated_output_dfs from README reference Remove the get_consolidated_output_dfs() row from the ValidationResult method table and trim it from the save() deprecation note, matching its removal from the basic tutorial. The docs (getting-started, known-issues) keep their deprecation notices for users still on the legacy accessor. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
1 parent 6dec980 commit e539856

51 files changed

Lines changed: 7287 additions & 6346 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎CHANGELOG.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,9 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
77

88
## [Unreleased]
99

10+
### Deprecated
11+
- `ValidationResult.get_consolidated_output_dfs()` and the `output_mode="failed_rows"` / `"both"` save modes (the legacy consolidated failed-rows CSVs) are deprecated in favour of annotated output (`get_annotated_output()` / `output_mode="annotated"`). Calling them now emits a `DeprecationWarning`. The `save()` default `output_mode` is still `"failed_rows"` but will change to `"annotated"` in a future minor release — pass `output_mode` explicitly to pin the behaviour you want. Annotated output supersedes the consolidated view: it returns your full tables with failing rows flagged in place via a per-row `check_info` column, plus per-check residues.
12+
1013
## [0.0.4] - 2026-07-06
1114

1215
### Fixed

‎README.md‎

Lines changed: 6 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -494,7 +494,6 @@ The `validate_data` function returns a powerful `ValidationResult` object that p
494494
| **`display_full_report(max_rows=5)`** | Prints summary + shows failed rows (convenience method) | `self` (chainable) |
495495
| **`save(output_dir=".", prefix="vowl_results", output_mode=None, check_info=None)`** | Saves enhanced CSV and summary JSON to disk. `output_mode` can be `"failed_rows"`, `"annotated"`, or `"both"`; `check_info` shapes the annotated `check_info` column (`"names"`/`"summary"`/`"full"`) | `self` (chainable) |
496496
| **`get_output_dfs(checks=None)`** | Returns per-check failed rows as `{check_id: DataFrame}` | Dict[str, DataFrame] |
497-
| **`get_consolidated_output_dfs(checks=None)`** | Deduplicates failed rows across checks, grouped by table | Dict[str, DataFrame] |
498497
| **`get_annotated_output(checks=None, check_info=None)`** | Returns full in-scope tables with a `check_info` column (JSON array of objects) marking failed rows | Dict[str, Dict[str, DataFrame]] |
499498
| **`.passed`** (property) | Boolean indicating if all checks passed | `True`/`False` |
500499

@@ -505,7 +504,7 @@ The `validate_data` function returns a powerful `ValidationResult` object that p
505504
It returns a nested dict with two reserved keys:
506505

507506
- **`"annotated"`** — a `{schema: table}` dict where each table is your full in-scope data plus a `check_info` column. Every original row is present; `check_info` is `null` for rows that passed everything and holds a JSON array of objects describing the failing check(s) otherwise.
508-
- **`"residues"`** — failed rows for checks that _cannot_ be merged onto a single table (cross-table, aggregation, and column-subset checks). Single-table contracts produce none. Residues are **per-check** (one entry per non-mergeable check, keyed `"<schema>::<check_name>"`) and carry the **same `check_info` column** as the annotated tables (a single-element JSON array, shaped by the same preset) plus `tables_in_query` — so everything `get_annotated_output()` returns is read the same way.
507+
- **`"residues"`** — failed rows for checks that _cannot_ be merged onto a single table (aggregation and column-subset checks, plus cross-table checks whose failed rows carry columns from more than the anchor table). Single-table contracts produce none. Residues are **per-check** (one entry per non-mergeable check, keyed `"<schema>::<check_name>"`) and carry the **same `check_info` column** as the annotated tables (a single-element JSON array, shaped by the same preset) plus `tables_in_query` — so everything `get_annotated_output()` returns is read the same way. (A cross-table check _can_ merge onto its home schema if you shape its failed-rows query to project only that schema's columns — see [Known Issues: Annotated Output](docs/known-issues.md#annotated-output-not-all-checks-can-be-merged).)
509508

510509
The **`check_info`** parameter (`"names"` default, `"summary"`, or `"full"`) shapes each array element. Every preset emits a JSON **array of objects** so consumers parse uniformly via `item["check_name"]`; they differ only in how many keys each object carries:
511510

@@ -568,7 +567,7 @@ clean = annotated[annotated["check_info"].isna()].drop(columns=["check_info"])
568567

569568
</details>
570569

571-
When a check spans more than one table (cross-table, aggregation, or column-subset checks), its failed rows can't be folded onto a single annotated table, so they surface under `"residues"` instead. Residues are **per-check** — one entry per non-mergeable check, keyed `"<schema>::<check_name>"`, each carrying its own failed rows plus the same `check_info` column the annotated tables use (a single-element JSON array) and `tables_in_query`:
570+
Aggregation checks, column-subset checks, and bare-JOIN cross-table checks can't be folded onto a single annotated table, so their failed rows surface under `"residues"` instead. (A cross-table check whose failed-rows query projects only its home schema's columns _is_ merged onto that schema — see the note above.) Residues are **per-check** — one entry per non-mergeable check, keyed `"<schema>::<check_name>"`, each carrying its own failed rows plus the same `check_info` column the annotated tables use (a single-element JSON array) and `tables_in_query`:
572571

573572
#### Residues
574573

@@ -604,7 +603,7 @@ Residue `'demo_employee_payroll::phone_number_exists_in_master_list'`: 2 failed
604603

605604
</details>
606605

607-
> For the full eligibility rules and worked examples of each non-mergeable category, see [Known Issues: Annotated Output](docs/known-issues.md#annotated-output-not-all-checks-can-be-merged). The [usage patterns notebook](examples/vowl_usage_patterns_demo.ipynb) walks through these examples end-to-end.
606+
> For the full eligibility rules and worked examples of each non-mergeable category, see [Known Issues: Annotated Output](docs/known-issues.md#annotated-output-not-all-checks-can-be-merged). The [Basic Tutorial notebook](examples/1_basic_tutorial/basic_tutorial.ipynb) walks through these examples end-to-end.
608607

609608
The `save()` method also supports annotated output via `output_mode`:
610609

@@ -619,6 +618,8 @@ result.save(output_mode="annotated", check_info="summary")
619618
result.save(output_mode="both")
620619
```
621620

621+
> **Deprecation:** `output_mode="failed_rows"` / `"both"` (the legacy failed-rows CSVs) are deprecated in favour of `"annotated"`. They still work but emit a `DeprecationWarning`. The `save()` default is currently `"failed_rows"` and will change to `"annotated"` in a future minor release — pass `output_mode` explicitly to pin the behaviour you want.
622+
622623
You can also set the output mode globally via `ValidationConfig`:
623624

624625
```python
@@ -700,7 +701,7 @@ result.save() # uses the configured output_mode
700701
701702
# Part 3 · Usage Patterns
702703
703-
> **Interactive demo:** Try the [usage patterns notebook](examples/vowl_usage_patterns_demo.ipynb) for a hands-on walkthrough of the examples below.
704+
> **Interactive demo:** Try the [example notebooks](examples/) for a hands-on walkthrough of the examples below — start with the [Basic Tutorial](examples/1_basic_tutorial/basic_tutorial.ipynb).
704705
705706
The patterns are grouped from most common to most advanced:
706707

‎docs/getting-started.md‎

Lines changed: 9 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -127,12 +127,12 @@ The `validate_data` function returns a powerful `ValidationResult` object that p
127127

128128
### Core Methods
129129

130-
| Method/Property | What It Does | Returns |
131-
| ------------------------------------------------- | -------------------------------------------------------------------------- | ---------------------- |
132-
| **`print_summary()`** | Prints high-level statistics (pass/fail counts, success rate, performance) | `self` (chainable) |
133-
| **`show_failed_rows(max_rows=5)`** | Displays sample of failed rows in console. Use `max_rows=-1` for all rows. | `self` (chainable) |
134-
| **`display_full_report(max_rows=5)`** | Prints summary + shows failed rows (convenience method) | `self` (chainable) |
135-
| **`save(output_dir=".", prefix="vowl_results")`** | Saves enhanced CSV and summary JSON to disk | `self` (chainable) |
136-
| **`get_output_dfs(checks=None)`** | Returns per-check failed rows as `{check_id: DataFrame}` | `Dict[str, DataFrame]` |
137-
| **`get_consolidated_output_dfs(checks=None)`** | Deduplicates failed rows across checks, grouped by table | `Dict[str, DataFrame]` |
138-
| **`.passed`** (property) | Boolean indicating if all checks passed | `True`/`False` |
130+
| Method/Property | What It Does | Returns |
131+
| ------------------------------------------------- | ----------------------------------------------------------------------------------------------------- | ---------------------- |
132+
| **`print_summary()`** | Prints high-level statistics (pass/fail counts, success rate, performance) | `self` (chainable) |
133+
| **`show_failed_rows(max_rows=5)`** | Displays sample of failed rows in console. Use `max_rows=-1` for all rows. | `self` (chainable) |
134+
| **`display_full_report(max_rows=5)`** | Prints summary + shows failed rows (convenience method) | `self` (chainable) |
135+
| **`save(output_dir=".", prefix="vowl_results")`** | Saves enhanced CSV and summary JSON to disk | `self` (chainable) |
136+
| **`get_output_dfs(checks=None)`** | Returns per-check failed rows as `{check_id: DataFrame}` | `Dict[str, DataFrame]` |
137+
| **`get_consolidated_output_dfs(checks=None)`** | _Deprecated_ — use `get_annotated_output()`. Deduplicates failed rows across checks, grouped by table | `Dict[str, DataFrame]` |
138+
| **`.passed`** (property) | Boolean indicating if all checks passed | `True`/`False` |

0 commit comments

Comments
 (0)