Skip to content

Commit c4e2933

Browse files
A BigQuery profile estimate reserves for an escalation query only where the probe could actually issue one (#302)
* Add constants for value domain limitations in base adapter * Refactor BigQuery adapter to enhance estimate composition with finer breakdown and reserved queries tracking * Enhance cost gate with OverCeilingError handling and reserve composition for improved estimate transparency * Refactor profile to reuse value domain constants from base adapter * Add unit tests for value domain and reserve behavior in profiling estimates * Add safety_spine test for BigQuery adapter profiling to ensure reserve narrowing aligns with billing estimates * Update docs * Cut release 1.6.4
1 parent a340b78 commit c4e2933

15 files changed

Lines changed: 532 additions & 73 deletions

File tree

‎.claude-plugin/plugin.json‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,7 +1,7 @@
11
{
22
"name": "dex",
33
"description": "Analytics engineering for Claude Code and any agent: data warehouse exploration, dbt transformation and semantic modeling, and schema-drift maintenance on dbt.",
4-
"version": "1.6.3",
4+
"version": "1.6.4",
55
"author": {
66
"name": "Exmergo, Inc.",
77
"email": "support@exmergo.com",

‎CHANGELOG.md‎

Lines changed: 41 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -9,6 +9,8 @@ tag releases both in lockstep, so entries below are keyed by the engine version.
99

1010
## [Unreleased]
1111

12+
## [1.6.4] - 2026-08-13
13+
1214
### Added
1315

1416
- **A verified join with a catastrophic orphan rate is now a finding, not
@@ -168,6 +170,45 @@ tag releases both in lockstep, so entries below are keyed by the engine version.
168170
profiled still needs a connection to build its sample, so it reaches the
169171
opener at the bottom of `cluster` and raises there.
170172

173+
### Fixed
174+
175+
- **A BigQuery profile estimate reserves for an escalation query only where the
176+
probe could actually issue one, and says how much of itself is reserve**
177+
([#299]). Since 1.6.0 the estimate held three per-table 10 MB floors rather
178+
than two: the value-domain probe added in that release
179+
([#203]) took a reserve beside the near-unique and composite ones. The reserve
180+
scales with object count rather than data size, so on a warehouse of many
181+
small tables the release moved a 12-object `explore map` estimate by 125.8 MB
182+
in one step and started refusing a nightly refresh that had run for months.
183+
184+
Three of the reserves were provably unspendable rather than merely unlikely,
185+
and are now dropped. A view has no row count (BigQuery reports none, and the
186+
profiling aggregate's own `COUNT(*)` is read per batch and never written back
187+
to the object), and all three probes return early without one, so a view held
188+
three floors no run could ever spend. Nested and repeated columns get no
189+
approximate distinct in the aggregate batch, and every probe's eligibility
190+
starts from one, so a table of nothing else can no more escalate than a blob
191+
column already excluded from the scan; the composite reserve was already
192+
conditioned on having two columns, but counted columns that can never join a
193+
pair. And a value domain needs at least one distinct value within a tenth of
194+
the non-null rows, so no column of a table below ten rows can qualify. The
195+
thresholds that imply that floor moved next to `ValueDomainSample` so the
196+
estimator and the probe cannot drift apart. Nothing else was narrowed: the
197+
reserve is dropped only where a probe's own guard already rules the query out,
198+
because an estimate that reserves for a query that cannot run is merely loose,
199+
while one that skips a query that can is the defect [#107] closed.
200+
201+
That leaves the common case, a warehouse of small flat tables, reserving
202+
exactly as much as before, which is the second half of this. The confirm
203+
handshake and the over-ceiling refusal now split the estimate into measured
204+
dry-run scan and held reserve, in prose and in `reserved_bytes` /
205+
`reserved_queries`. A refusal previously said only "raise the budget or narrow
206+
the work" about a number that could be three quarters reserve, leaving the
207+
operator to reconstruct the split from `.dex/spend.jsonl` afterwards to find
208+
out whether the estimate grew because the warehouse did or because dex added a
209+
probe. Over-ceiling refusals reach the connector's own description of the
210+
estimate for the first time; before this only the confirmation payload did.
211+
171212
## [1.6.3] - 2026-08-10
172213

173214
### Changed

‎packages/dex-core/src/exmergo_dex_core/adapters/base.py‎

Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -23,6 +23,7 @@
2323
from dataclasses import dataclass
2424
from datetime import date, datetime, time
2525
from decimal import Decimal
26+
from math import ceil
2627
from typing import Protocol, runtime_checkable
2728

2829
from ..envelope import Paradigm
@@ -123,6 +124,24 @@ class ValueDomainSample:
123124
total_distinct: int
124125

125126

127+
# A column's value domain is reported only when its distinct count clears BOTH
128+
# bars: small absolutely (the cap) and small relative to the table (the
129+
# fraction), so a tiny table's near-key column does not qualify on the absolute
130+
# count alone. Deliberately conservative: this codebase's general posture is to
131+
# under-report rather than over-report, and a false negative here just costs one
132+
# more `explore query`.
133+
#
134+
# They live beside the sample type rather than in `explore.profile` because a
135+
# metered adapter has to reserve for the probe *before* profiling runs, and the
136+
# two sides would drift apart the moment one of them moved. VALUE_DOMAIN_MIN_ROWS
137+
# is what the fraction implies rather than a threshold of its own: a domain needs
138+
# at least one distinct value, so a table below it can never clear the fraction
139+
# and never produces a probe worth reserving for.
140+
VALUE_DOMAIN_CAP = 25
141+
VALUE_DOMAIN_MAX_FRACTION = 0.10
142+
VALUE_DOMAIN_MIN_ROWS = ceil(1 / VALUE_DOMAIN_MAX_FRACTION)
143+
144+
126145
@dataclass(frozen=True)
127146
class QueryResult:
128147
"""The result of one firewall-approved agent query, columnar.

‎packages/dex-core/src/exmergo_dex_core/adapters/bigquery.py‎

Lines changed: 138 additions & 31 deletions
Original file line numberDiff line numberDiff line change
@@ -17,6 +17,7 @@
1717
from __future__ import annotations
1818

1919
from collections.abc import Callable
20+
from dataclasses import dataclass
2021
from typing import Any
2122

2223
from ..config import BigQueryTarget
@@ -25,6 +26,7 @@
2526
from ..guards.cost_guard import CostGate, OverCeilingError
2627
from ..guards.sql_guard import assert_select_only
2728
from .base import (
29+
VALUE_DOMAIN_MIN_ROWS,
2830
ColumnAggregate,
2931
ColumnMeta,
3032
ObjectMeta,
@@ -63,6 +65,25 @@
6365
_NESTED_FIELD_TYPES = {"RECORD", "STRUCT", "JSON", "GEOGRAPHY", "RANGE", "INTERVAL"}
6466

6567

68+
@dataclass(frozen=True)
69+
class _EstimateComposition:
70+
"""What a profile estimate is made of: measured dry-run scan versus
71+
escalation reserve.
72+
73+
The two halves answer different questions and a caller deciding whether to
74+
raise a budget needs them apart. The scan half is a dry-run of statements
75+
that will certainly run; the reserve half is headroom held for probes that
76+
may never be issued (see :meth:`BigQueryAdapter.profile_estimate`). Quoting
77+
only the sum is what left a refused nightly run unable to tell a warehouse
78+
that had grown from a release that had added a reserve (issue #299).
79+
"""
80+
81+
scan_bytes: float
82+
scan_queries: int
83+
reserved_bytes: float
84+
reserved_queries: int
85+
86+
6687
def _regexp_predicate(qcol: str, pattern: str) -> str:
6788
# Raw string literal; REGEXP_CONTAINS matches substrings, so the shared
6889
# patterns' anchors make it a full match.
@@ -130,6 +151,10 @@ def __init__(
130151
self._tables: dict[str, Any] = {}
131152
self._resolved_datasets: list[str] | None = None
132153
self._notes: dict[str, list[str]] = {}
154+
# What the last profile estimate was made of, so the handshake and the
155+
# over-ceiling refusal can attribute the number they quote. Per command,
156+
# like `_notes`: one command builds at most one profile estimate.
157+
self._composition: _EstimateComposition | None = None
133158

134159
# --- capabilities ---------------------------------------------------------
135160

@@ -684,18 +709,18 @@ def profile_estimate(
684709
what BigQuery actually bills, and an unfloored estimate would send the
685710
agent into a ladder of budget rejections.
686711
687-
The total also reserves one floor per table for each of the two
688-
possible escalation queries (``exact_distinct_counts`` for a
689-
near-unique column, ``distinct_combination_counts`` for a composite
690-
key candidate): whether either actually runs depends on the aggregate
691-
batch's own approximate results, which do not exist yet at estimate
692-
time, so there is no way to dry-run them here. Reserving their floor
693-
unconditionally keeps this total a ceiling profiling will not exceed,
694-
rather than a number a run that does escalate blows past -- the exact
695-
gap this estimator used to leave open (issue #107). Skipped only when
696-
a table is provably empty (``row_count == 0``); an unknown row count
697-
(views, whose count is never known before the aggregate that reveals
698-
it) still reserves, since escalation is not ruled out.
712+
The total also reserves one floor per table for each of the three
713+
escalation queries a profile may still issue after that batch
714+
(:meth:`exact_distinct_counts`, :meth:`value_domain_counts`,
715+
:meth:`distinct_combination_counts`). Whether any of them runs depends
716+
on the batch's own approximate results, which do not exist yet at
717+
estimate time, so there is no way to dry-run them here. Holding their
718+
floor keeps this total a ceiling profiling will not exceed rather than
719+
a number a run that does escalate blows past, which is the gap this
720+
estimator used to leave open (issue #107). See
721+
:meth:`_escalation_reserve` for what narrows each one, and
722+
:attr:`_composition` for how the reserve is reported apart from the
723+
scan it rides with.
699724
700725
Blob-type columns are excluded from the batches the same way
701726
``explore.profile.profile`` excludes them from the scan itself
@@ -704,6 +729,9 @@ def profile_estimate(
704729

705730
blob_paths = include_blobs or set()
706731
per_table: dict[str, float] = {}
732+
scan_bytes = 0.0
733+
scan_queries = 0
734+
reserved_queries = 0
707735
for identifier in identifiers:
708736
meta, columns = self.table_metadata(identifier)
709737
if self._unqueryable(identifier):
@@ -732,21 +760,68 @@ def profile_estimate(
732760
sample_percent=sample_percent,
733761
)
734762
try:
735-
total += max(self._dry_run(sql), float(_MIN_BILLED_BYTES))
763+
batch = max(self._dry_run(sql), float(_MIN_BILLED_BYTES))
736764
except self._api_exceptions.BadRequest:
737765
self._note(
738766
identifier,
739767
"could not estimate an aggregate scan (dry-run failed); "
740768
"the object is skipped",
741769
)
742-
if scan_columns and meta.row_count != 0:
743-
total += float(_MIN_BILLED_BYTES) # exact_distinct_counts
744-
total += float(_MIN_BILLED_BYTES) # value_domain_counts
745-
if len(scan_columns) >= 2:
746-
total += float(_MIN_BILLED_BYTES) # distinct_combination_counts
770+
continue
771+
total += batch
772+
scan_bytes += batch
773+
scan_queries += 1
774+
reserved = self._escalation_reserve(meta, scan_columns)
775+
total += reserved * float(_MIN_BILLED_BYTES)
776+
reserved_queries += reserved
747777
per_table[identifier] = total
778+
self._composition = _EstimateComposition(
779+
scan_bytes=scan_bytes,
780+
scan_queries=scan_queries,
781+
reserved_bytes=reserved_queries * float(_MIN_BILLED_BYTES),
782+
reserved_queries=reserved_queries,
783+
)
748784
return sum(per_table.values()), per_table
749785

786+
def _escalation_reserve(
787+
self, meta: ObjectMeta, scan_columns: list[ColumnMeta]
788+
) -> int:
789+
"""How many escalation queries to hold a billing floor for on one table.
790+
791+
A reserve is dropped only where the probe's own guard already rules the
792+
query out from metadata alone, never as a guess about what is likely: an
793+
estimate that reserves for a query that cannot run is merely loose, while
794+
one that skips a query that can is the defect issue #107 closed. So each
795+
condition below mirrors one in ``explore.profile``, and moving one
796+
without the other is the bug to watch for.
797+
798+
- Nothing at all without a row count. All three probes return early on a
799+
falsy one, and a BigQuery view never has one: ``_object_meta`` nulls it
800+
and the aggregate's own ``COUNT(*)`` is read per batch, never written
801+
back to the object. So a view provably cannot escalate, and the reserve
802+
it used to hold was money no run could spend.
803+
- Nothing for columns BigQuery cannot count distinctly. Nested and
804+
repeated fields get no approximate distinct in the aggregate batch, and
805+
every probe's eligibility starts from one, so they can no more trigger
806+
an escalation than a blob column already excluded from the scan.
807+
- No value domain below :data:`VALUE_DOMAIN_MIN_ROWS` rows, which is what
808+
the probe's row-relative fraction implies once a domain needs at least
809+
one distinct value.
810+
- No composite probe below two countable columns, since a combination
811+
needs two members. This was already conditioned, but on the raw column
812+
count, which counts columns that can never join a pair.
813+
"""
814+
815+
countable = [c for c in scan_columns if not self._is_nested(c.data_type)]
816+
if not countable or not meta.row_count:
817+
return 0
818+
reserved = 1 # exact_distinct_counts
819+
if meta.row_count >= VALUE_DOMAIN_MIN_ROWS:
820+
reserved += 1 # value_domain_counts
821+
if len(countable) >= 2:
822+
reserved += 1 # distinct_combination_counts
823+
return reserved
824+
750825
def query_estimate(self, sql: str) -> float:
751826
"""The dry-run byte estimate for one firewall-approved query, floored to
752827
what BigQuery will actually bill (the per-referenced-table minimum), so
@@ -761,33 +836,65 @@ def describe_estimate(
761836
"""The bytes-scanned handshake payload: names the per-query billing
762837
floor baked into every number here, so a small-table estimate reads as
763838
a trustworthy ceiling instead of one the actual bill will exceed
764-
(issue #107). The escalation-reserve clause only applies to a
765-
multi-table call (``per_table`` set, i.e. a profile-shaped estimate);
766-
a single ad-hoc query or cluster sample has no such reserve to explain.
839+
(issue #107).
840+
841+
Deliberately a pure function of its arguments. What a *profile* estimate
842+
is made of is answered by :meth:`profile_reserve` instead, because this
843+
method also describes estimates that carry no reserve (an ad-hoc query,
844+
a mid-command verify checkpoint) and it cannot tell which it was handed.
767845
"""
768846

769-
note = (
770-
f"BigQuery bills at least {_MIN_BILLED_BYTES:,} bytes (10 MB) per "
771-
"query; every number here already reflects that floor"
772-
)
773-
if per_table:
774-
note += (
775-
", including a reserve for the escalation queries a "
776-
"multi-table profile may still add after its initial scan"
777-
)
778847
data: dict[str, object] = {
779848
"estimated_bytes": estimate,
780849
"hint": (
781850
"review the estimate, then re-run with --confirm --budget "
782851
"<bytes> (the ceiling in bytes; 10000000000 is 10 GB, about "
783852
"$0.06 on-demand)"
784853
),
785-
"notes": [note],
854+
"notes": [
855+
f"BigQuery bills at least {_MIN_BILLED_BYTES:,} bytes (10 MB) "
856+
"per query; every number here already reflects that floor"
857+
],
786858
}
787859
if per_table:
788860
data["per_table_bytes"] = per_table
789861
return data
790862

863+
def profile_reserve(self, estimate: float) -> dict | None:
864+
"""How much of ``estimate`` is escalation reserve rather than measured
865+
scan, as a sentence plus the two numbers behind it. ``None`` when this
866+
command priced no profile, so there is nothing to attribute.
867+
868+
The confirm handshake and the over-ceiling refusal both reach for this,
869+
which is the point of it existing: a refusal that quotes a number
870+
without saying what it is made of leaves the operator to reconstruct
871+
the split from the spend ledger by hand, which is what issue #299
872+
reported doing. Only the command-level handshake asks, because only
873+
that estimate is the one :meth:`profile_estimate` built.
874+
875+
The remainder is described rather than the recorded scan half quoted,
876+
because a caller may have added statement estimates of its own to the
877+
total (``explore query`` prices an auto-profile and its statements in
878+
one handshake) and those are dry-run figures too.
879+
"""
880+
881+
composition = self._composition
882+
if composition is None or not composition.reserved_queries:
883+
return None
884+
return {
885+
"note": (
886+
f"{composition.reserved_bytes:,.0f} bytes of this estimate is "
887+
f"escalation reserve: {composition.reserved_queries} queries at "
888+
f"BigQuery's {_MIN_BILLED_BYTES:,}-byte per-query minimum, held "
889+
"for probes a profile may add after its aggregate scan and may "
890+
"never issue. The remaining "
891+
f"{estimate - composition.reserved_bytes:,.0f} bytes is dry-run "
892+
"scan"
893+
),
894+
"reserved_bytes": composition.reserved_bytes,
895+
"reserved_queries": composition.reserved_queries,
896+
}
897+
791898
def _min_billed_floor(self, sql: str) -> float:
792899
"""BigQuery bills at least ``_MIN_BILLED_BYTES`` per table a query
793900
references. The floor for one query is that minimum times its distinct

‎packages/dex-core/src/exmergo_dex_core/command_args.py‎

Lines changed: 25 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -18,7 +18,7 @@
1818
from .adapters.base import Adapter
1919
from .config import CONFIG_FILE
2020
from .envelope import Cost
21-
from .guards.cost_guard import ConfirmationRequiredError, CostGate
21+
from .guards.cost_guard import ConfirmationRequiredError, CostGate, OverCeilingError
2222
from .results import ConfirmationRequest, Result
2323
from .storage import DEX_DIR
2424

@@ -117,8 +117,24 @@ def billed_handshake(
117117
gate = cost_gate(adapter)
118118
if gate is None:
119119
return
120+
# What a connector can say about the composition of the estimate it just
121+
# priced, asked once and used by both refusals below. Only this handshake
122+
# asks: it is the one whose estimate the connector's profile pricing built,
123+
# and a mid-command phase (verify probes) prices something else entirely.
124+
reserve = getattr(adapter, "profile_reserve", None)
125+
reserve = reserve(estimate) if reserve is not None else None
120126
try:
121127
gate.preflight_command(estimate)
128+
except OverCeilingError as exc:
129+
# A refusal is the message an operator acts on and the one that carries
130+
# the least. Without this, an estimate padded with escalation reserve
131+
# reads as work the warehouse is about to do, and the only way to tell
132+
# the two apart is to reconstruct the split from the spend ledger
133+
# afterwards (issue #299). Re-raised rather than mutated so the
134+
# exception stays a plain immutable refusal, carrying the same cost.
135+
if reserve is None:
136+
raise
137+
raise OverCeilingError(f"{exc}. {reserve['note']}", cost=exc.cost) from exc
122138
except ConfirmationRequiredError as exc:
123139
# The payload speaks the connector's unit. An adapter that knows more
124140
# than the raw magnitude (Snowflake's credit translation, its
@@ -139,6 +155,14 @@ def billed_handshake(
139155
}
140156
if per_table:
141157
data["per_table_bytes"] = per_table
158+
if reserve is not None:
159+
# Prose and flat keys both: the note is what a person reads when
160+
# deciding whether raising the budget buys work or headroom, and
161+
# the keys are so a host branching on the split does not have to
162+
# parse the sentence to find it.
163+
data["reserved_bytes"] = reserve["reserved_bytes"]
164+
data["reserved_queries"] = reserve["reserved_queries"]
165+
data["notes"] = [*data.get("notes", []), reserve["note"]]
142166
if notes:
143167
data.setdefault("notes", [])
144168
data["notes"] = [*data["notes"], *notes]

0 commit comments

Comments
 (0)