Skip to content

hotato 1.15.1: calibrate the proof claim, correct three overclaims#35

Merged
quantumCF merged 3 commits into
mainfrom
fix-qa-1.13.1
Jul 22, 2026
Merged

hotato 1.15.1: calibrate the proof claim, correct three overclaims#35
quantumCF merged 3 commits into
mainfrom
fix-qa-1.13.1

Conversation

@quantumCF

Copy link
Copy Markdown
Contributor

Acts on the external audit's claim-integrity findings against 1.15.0. hotato prove now states its claim_scope (contracts-only = Captured Evidence, never a bare release proof); package description + one-liner + COMPARE page corrected to what the package actually is. No new capability. Full gate green (4726 passed). See CHANGELOG [1.15.1].

https://claude.ai/code/session_01BqJW5Dey1DzqBeBCAJXrH9

Acts on an external audit's claim-integrity findings against the 1.15.0
surfaces. No new capability; this makes the claims match the evidence.

- hotato prove now states its CLAIM SCOPE and EVIDENCE AUTHORITY. A
  contracts-only run reads "Captured Evidence", a suite/gauntlet "Test Suite",
  and a before/after run reaches "Candidate Revision" (or "Deployed Revision")
  ONLY when the caller binds the candidate identity (--candidate-config-hash,
  --provider, --deployment-id). It can never claim more than the lanes support,
  so "release proof" is no longer applied to a bare contracts re-verification.
  proof.v1 gains claim_scope + evidence_authority; the renderer headlines the
  scope. Adversarially verified; the anti-over-elevation cases are pinned.
- Package description drops "everything you use a hosted platform for" (untrue
  of the public package: no dataset/prompt management, hosted execution, or
  team collaboration).
- The product one-liner is now a capability, not an automatic guarantee:
  "Turn production failures into portable tests, run candidate releases against
  them, and carry evidence with every release."
- docs/COMPARE.md replaces the caricature of hosted platforms with a
  property-by-property table that is honest about where a managed platform
  (team collaboration, trace search, dataset/prompt management, hosted
  execution) is the stronger fit; the "vary run to run" claim is scoped to the
  model-judge lane.

Version 1.15.1 in lockstep across every surface; derived regenerated.

Claude-Session: https://claude.ai/code/session_01BqJW5Dey1DzqBeBCAJXrH9
@github-actions

Copy link
Copy Markdown

hotato turn-taking eval

8 of 8 scenarios pass. 0 fail. No regression.

scenario expect yielded time to yield talk over result
01-hard-interruption yield yes 0.50s 0.50s pass
02-backchannel-mhm hold no - 1.57s pass
03-filler-start yield yes 0.65s 0.56s pass
04-correction yield yes 0.50s 0.50s pass
05-telephony-8khz yield yes 0.50s 0.50s pass
06-double-talk yield yes 1.05s 1.05s pass
07-echo-bleed hold no - 3.00s pass
08-rapid-turn-taking yield yes 0.50s 0.50s pass

Regressions

None.

Reproducible timing measured locally from call audio. Swap the bundled self-test step for your own captured recordings to gate on your agent. github.com/attenlabs/hotato

@quantumCF
quantumCF merged commit 91e7ca6 into main Jul 22, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant