feat(superforcaster-polymarket-v5): add base_rate_required field to bind TYPE B base-rate anchor (issue #452) - #453
Conversation
…ind TYPE B base-rate anchor New version of superforcaster-polymarket-v4 targeting the confirmed mechanism from issue #452: the TYPE A/TYPE B evidence-reliability screen was not preventing extreme p_yes outputs because (a) undated standing pages (news outlet homepages whose names contain the keyword) were misclassified as TYPE A criterion-confirming evidence, and (b) the text-based base-rate instruction did not mechanically bind the subsequent numeric p_yes output even when all sources were correctly identified as TYPE B. The fix adds a base_rate_required: bool structured-output field between evidence_reliability_screen and aggregation in PredictionResult. The model must explicitly commit True/False to whether the base-rate anchor applies before it can write aggregation and the numeric fields. When True, aggregation must begin from the stated base rate, keeping p_yes within 10 pp of the base rate in the absence of within-window TYPE A criterion-confirming evidence. Housekeeping: tournament roster swap (v4 -> v5), tool_lineage.json entry added, benchmark/tools.py TOOL_REGISTRY updated. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Black formatting check was failing in CI run 33001001367. Files reformatted: - packages/valory/customs/superforcaster_polymarket_v5/superforcaster_polymarket_v5.py - packages/valory/customs/superforcaster_polymarket_v5/tests/test_superforcaster_polymarket_v5.py Also fixed isort ordering in the test file. Re-ran autonomy packages lock to update component.yaml fingerprints and packages.json CID after the source files changed. Bumped tournament_tools.json to new CID.
|
CI is fixed. The linter_checks job (run 33001001367) was failing because What was fixed: Applied
After reformatting, Fix commit: 75bfd4a (ref CI run 33001001367) |
|
[human-input] The investigation memory file ( Proceeding with the W-2 benchmark trigger: Sub-pipeline D is a mechanical idempotency check + trigger that does not require the investigation memory — all needed facts (platform, baseline, W-2 sample count) are in the PR body. If the PR later needs a code revision or benchmark diagnosis (Sub-pipelines A/B/E), please either write the memory file or call me with the context you want me to use. Context used from PR body:
|
|
/benchmark superforcaster-polymarket-v4 --candidate-tool superforcaster-polymarket-v5 --platform polymarket --sample 105 |
|
valory-coding-agent metadata (request record -- do not edit) {
"baseline": "valory/superforcaster_polymarket_v4:0.1.0:bafybeieuzna7fsv4iwz7gvqpvkurfdskd2tul6b4asjg4dcsp4oieic6ti",
"candidate": "valory/superforcaster_polymarket_v5:0.1.0:bafybeics2og73yiydgzdtplwfsvvgucom4ozaqdetbxpzqgan6ka2kewny",
"platform": "polymarket",
"sample": 105,
"seed": 42,
"run_type": "iteration",
"pr_head": "75bfd4a",
"reason": "First benchmark — CI green on v5 (head 75bfd4a); W-2 is C-2 count-window rows (n=105, 2026-07-23 to 2026-08-07); baseline v4 has n=105 rows on Polymarket"
} |
Benchmark: superforcaster-polymarket-v4 — Polymarket
Computed on 105 markets where both arms produced a usable prediction. Candidate health
Sample
105 scored | seed 42 | triggered by @valory-coding-agent[bot] |
|
The v5 fix shows a clear improvement on the iteration benchmark: Brier fell 24.5% and overconfident-wrong calls dropped 84%, directly confirming the TYPE B misclassification hypothesis. A holdout-confirmation run on a fresh seed is needed before the evidence supports promotion. Benchmark result — iteration run (seed 42, n=105, Polymarket) Context recovered from PR body — investigation memory file was not committed to this branch (noted in the prior [human-input] comment).
Conclusions Targeted fingerprint (overconf-wrong): confirmed. 19 → 3 (-84.2%). This is the direct mechanism fingerprint of H1: the Aggregate Brier: -0.0796 across 105 matched deliveries (0.3254 → 0.2458). The candidate beats both the v4 production score and the W-2 baseline (0.2958). At n=105, the rough paired SE is ~0.04–0.05; the improvement is ~1.5–2× SE. Most of the Brier mass comes from the 16 overconf-wrong cases fixed (each contributing ~0.65–0.80 per-row improvement), so the aggregate is mechanistically grounded — not noise. Parse rate: 0 failures / 105 markets (structurally guaranteed by the v5 Pydantic structured-output schema; the What the evidence supports: H1 is confirmed and the implementation worked as designed on this sample. The candidate beats the promotion gate (W-2 Brier 0.2958) with room to spare. What it does not yet establish: Generalisation. This was an iteration run (seed 42). Per gate-honesty constraints, a holdout-confirmation run on a fresh seed is required before the improvement can be claimed to hold out-of-sample. To enable the holdout, reply with Benchmark budget for this PR: 1 of 1 used. Note You can request further changes by tagging valory-coding-agent. |
The production logs showed this tool confidently wrong on "keyword in headlines this week" and "keyword in post this week" markets — e.g. p_yes=0.995 that "Star" would appear in headlines, which resolved NO. Inspecting 10 deliveries confirmed the tool was getting reasonable evidence but misclassifying news outlet homepages (whose names contain the keyword) as confirmed criterion-meeting TYPE A sources, then ignoring the base-rate instruction in the same structured field. This PR adds a boolean commitment field that mechanically forces the model to declare whether the base-rate anchor applies before it writes the aggregation and the numeric outputs, so it can no longer acknowledge TYPE B in text and then emit p_yes=0.99.
Reviewer: after merging, run
autonomy push-allto publish the bytes to IPFS, then updateagent-deploymentswith the new CIDbafybeibx6sg3smqiezrcxx2s6entvf7flqw5m7nushgslvkc4xa2kdjhde.Investigation (issue #452)
Tool / platform:
superforcaster-polymarket-v4/ PolymarketWindow: C-1 count window 2026-08-07 to 2026-08-24 (n=105, count-window fallback)
Headline: Brier 0.3810 vs C-2 0.2958 (delta +0.0852, BSS -0.7980)
Disposition: CHRONIC-BAD (Brier >= 0.25 threshold); default PROPOSE per pipeline
Localized cell
by_disagree_bucket=largeAll Brier mass is in the large-disagree bucket (bet-eligible zone): trader bets these calls and loses here.
Version/mix control
C-2 had 18 rows from a retiring CID (Brier 0.1979) not present in C-1. Within-version current-CID comparison: C-2=0.2741 vs C-1=0.3390, delta +0.0649. Genuine tool-level rise, not a composition artifact.
Evidence inspection (10 worst-miss deliveries, 10/10 readable)
10/10 deliveries: good-evidence/bad-reasoning. All show the same two failure modes:
evidence_reliability_screencorrectly identified all-TYPE-B evidence (e.g. April-2026 Trump posts on a question about "this week") still produced extreme p_yes (0.98), because the text instruction to "anchor on base rate" did not mechanically prevent the subsequent numeric output.Overconfident-wrong side: p_yes=0.98--0.995 on NO outcomes (large positive disagree, trader loses). Underconfident-wrong side: p_yes=0.005--0.010 on YES outcomes (large negative disagree, same class — evidence was filtered to PM odds only, tool collapsed to near-zero instead of the base rate).
Hypothesis
H1 (confirmed): At the prediction-LLM-call stage, the
evidence_reliability_screenTYPE A/TYPE B classification does not prevent extremep_yesoutputs because (a) the TYPE A definition does not explicitly exclude keyword-in-outlet-name standing pages or out-of-window historical posts, and (b) the text-based base-rate instruction does not mechanically bind the subsequent numericp_yesfield. Classification:good-evidence/bad-reasoningat a gate-visible stage (downstream ofsource_contentinjection).Fix (single structural change at prediction-LLM-call stage)
Tighten TYPE A definition in
evidence_reliability_screen: TYPE A requires BOTH (i) explicitly dated WITHIN the resolution window AND (ii) directly confirms the exact criterion. A keyword in an outlet's NAME, a standing homepage, or an out-of-window historical post is TYPE B.Add
base_rate_required: boolfield betweenevidence_reliability_screenandaggregationinPredictionResult. The model must explicitly commit True/False before writingaggregationand the numeric fields. When True,aggregationmust begin from the stated base rate and stay within 10 pp of it in the absence of within-window TYPE A criterion-confirming evidence.Both changes are gate-visible (downstream of
source_contentinjection) and are exercised by PR-CI's cached W-2 replay.Source-change size (pre-lint)
Housekeeping
packages/valory/customs/superforcaster_polymarket_v5/superforcaster_polymarket_v5.pypackages/valory/customs/superforcaster_polymarket_v5/component.yamlpackages/valory/customs/superforcaster_polymarket_v5/__init__.pypackages/valory/customs/superforcaster_polymarket_v5/tests/test_superforcaster_polymarket_v5.pytest_on_chain_result_is_flat_json_loads_parseable,test_uses_structured_parse_not_raw_create,test_numeric_fields_are_declared_last,test_base_rate_required_field_present)packages/packages.jsonbafybeibx6sg3smqiezrcxx2s6entvf7flqw5m7nushgslvkc4xa2kdjhdebenchmark/tools.pysuperforcaster-polymarket-v5added to TOOL_REGISTRYbenchmark/tournament_tools.jsontool_lineage.jsonsuperforcaster-polymarket-v5entry added (parent: v4)packages/valory/customs/superforcaster_polymarket_v4/New tool CID:
bafybeibx6sg3smqiezrcxx2s6entvf7flqw5m7nushgslvkc4xa2kdjhdeHuman step before merging:
autonomy push-all(publishes the new CID to IPFS).Human step after merging: PR to
agent-deploymentsto updateTOOLS_TO_PACKAGE_HASH.Step 6.5 Pre-PR Sanity
importsucceeds,run()callableautonomy packages lock --checkexits 0json.loads(run()[0])does NOT raise (trader parse verified)run()[0]starts with{set(parsed.keys()) == {"p_yes", "p_no", "confidence", "info_utility"}— exactly the four mech fields, no reasoning leaksbase_rate_requiredNOT present in on-chain outputbase_rate_requiredsits betweenevidence_reliability_screenandaggregationinPredictionResult.model_fields# type: ignore[import-not-found]for third-party stubs without stubs installed)beta.chat.completions.parseused (notchat.completions.create) — structured-output contract preservedBaseline to beat
W-2 (C-2, cached replay): Brier 0.2958, n=105 rows (2026-07-23 to 2026-08-07)
PR-CI will post the delta vs this baseline. A negative delta (Brier falls) is the promotion signal.
Code diff (v4 vs v5, key changes only)
Key structural changes (expand)
TYPE A definition tightened (v4 -> v5 in
evidence_reliability_screen):Aggregation binding (v4 -> v5):
🤖 Generated with Claude Code