Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 23 additions & 0 deletions runs/stage3-rerun/SUMMARY.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
# Stage 3 code-001 fleet rerun summary

Base commit: `5709d66` (#246 terminal toolset auto-probe).

Pre-run gate: each Linux node was updated with `git pull origin main`, `/work/agent-codebench` was reset from `fixtures/season-001/code-001/target-repo`, repo dependencies were installed where needed, and the shipping bench gate passed (`npm test` exit 0 with 4 tests; `npm run report` exit 1 with the expected `metrics.retries` TypeError) before wrapper execution. Android/Termux nodes were not run because `/work` is unavailable.

Wrapper command shape: `HERMES_NODE=<node> HERMES_EVENT_FAMILY=code bash adapters/wrappers/hermes-mission-wrapper.sh tasks/season-001/code-001-typescript-regression-v2.yaml runs/stage3-rerun/code-001-<agent> <agent>`. No `HERMES_TOOLSETS` / `HERMES_EXEC_TOOLSET` override was supplied; terminal was auto-derived from `hermes tools list`.

## Gate results

- soonwook (vps6): status=completed; toolsets=file,terminal; toolsets_source=tools_list_exec; parse_fallback=0; hermes_status=250; post_test=0; post_report=0; changed=3 files
- sogyo (vps1): status=completed; toolsets=file,terminal; toolsets_source=tools_list_exec; parse_fallback=0; hermes_status=250; post_test=0; post_report=0; changed=3 files
- nosuk (vps2): status=completed; toolsets=file,terminal; toolsets_source=tools_list_exec; parse_fallback=0; hermes_status=250; post_test=0; post_report=0; changed=3 files
- dungae (vps0): status=completed; toolsets=file,terminal; toolsets_source=tools_list_exec; parse_fallback=0; hermes_status=250; post_test=0; post_report=0; changed=3 files
- bangtong (vps3): status=completed; toolsets=file,terminal; toolsets_source=tools_list_exec; parse_fallback=0; hermes_status=250; post_test=0; post_report=0; changed=2 files
- jingun (vps8): status=completed; toolsets=file,terminal; toolsets_source=tools_list_exec; parse_fallback=0; hermes_status=250; post_test=0; post_report=0; changed=3 files
- seoseo (vps4): status=completed; toolsets=file,terminal; toolsets_source=tools_list_exec; parse_fallback=0; hermes_status=250; post_test=0; post_report=0; changed=3 files
- yukson (vps5): status=completed; toolsets=file,terminal; toolsets_source=tools_list_exec; parse_fallback=0; hermes_status=250; post_test=0; post_report=0; changed=3 files
- gwakga (vps7): status=completed; toolsets=file,terminal; toolsets_source=tools_list_exec; parse_fallback=0; hermes_status=0; post_test=0; post_report=0; changed=3 files
- gongyung: NOT-RUN — Android/Termux: /work unavailable for code-001 bench
- daegyo: NOT-RUN — Android/Termux: /work unavailable for code-001 bench

All 9 Linux runs passed the pre-scoring toolset cohort gate: `toolsets=file,terminal`, `toolsets_source=tools_list_exec`, `parse_fallback=0`. No `probe_*_file` fallback node was admitted. `hermes_status=250` appears on 8 nodes after parseable mission output and is treated the same class as the known shutdown quirk; artifacts remain parseable and validated. gwakga returned `hermes_status=0`.
35 changes: 35 additions & 0 deletions runs/stage3-rerun/code-001-bangtong/adapter-bootstrap.log
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
[hermes-adapter] result-packet.yaml — validation OK
[hermes-adapter] trace.yaml — validation OK
[hermes-adapter] evidence-bundle.yaml — validation OK
[hermes-adapter] manifest.yaml — validation OK
[hermes-adapter] WARNING: capability matrix requires evidence kind "artifact_hash" for status "completed" but the packet does not include it.

=== Hermes Adapter Run Complete ===
Run ID: run-code-001-bangtong-2026-06-13T03-49-24-bc3859
Workflow ID: wf-code-001-bangtong-2026-06-13T03-49-24-bc3859
Task: code-001 (TypeScript regression fix with targeted tests)
Agent: bangtong
Runtime: hermes 0.16.0
Mode: orchestrator
Event family: code
Model: gpt-5.x (openai)
Status: completed
Exit code: 0
Workers: 3
Steps: 3
Contradictory: false
Run dir: /root/agent-olympics/runs/stage3-rerun/code-001-bangtong
Duration: 0s
Validate: PASSED
Adapter version: 1.0.0
Publishable: false

=== Adapter Metadata ===
Adapter: hermes
Envelope versions: 1, 2
Event families: ops, code, smoke, node, wiki, general, coord
Modes: orchestrator, coordinator, simulation
Redaction rules: 4
Evidence kinds: 10
Default timeout: 900s

35 changes: 35 additions & 0 deletions runs/stage3-rerun/code-001-bangtong/adapter.log
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
STDOUT [hermes-adapter] result-packet.yaml — validation OK
STDOUT [hermes-adapter] trace.yaml — validation OK
STDOUT [hermes-adapter] evidence-bundle.yaml — validation OK
STDOUT [hermes-adapter] manifest.yaml — validation OK
STDERR [hermes-adapter] WARNING: capability matrix requires evidence kind "artifact_hash" for status "completed" but the packet does not include it.
STDOUT
STDOUT === Hermes Adapter Run Complete ===
STDOUT Run ID: run-code-001-bangtong-2026-06-13T03-49-24-bc3859
STDOUT Workflow ID: wf-code-001-bangtong-2026-06-13T03-49-24-bc3859
STDOUT Task: code-001 (TypeScript regression fix with targeted tests)
STDOUT Agent: bangtong
STDOUT Runtime: hermes 0.16.0
STDOUT Mode: orchestrator
STDOUT Event family: code
STDOUT Model: gpt-5.x (openai)
STDOUT Status: completed
STDOUT Exit code: 0
STDOUT Workers: 3
STDOUT Steps: 3
STDOUT Contradictory: false
STDOUT Run dir: /root/agent-olympics/runs/stage3-rerun/code-001-bangtong
STDOUT Duration: 0s
STDOUT Validate: PASSED
STDOUT Adapter version: 1.0.0
STDOUT Publishable: false
STDOUT
STDOUT === Adapter Metadata ===
STDOUT Adapter: hermes
STDOUT Envelope versions: 1, 2
STDOUT Event families: ops, code, smoke, node, wiki, general, coord
STDOUT Modes: orchestrator, coordinator, simulation
STDOUT Redaction rules: 4
STDOUT Evidence kinds: 10
STDOUT Default timeout: 900s
STDOUT
6 changes: 6 additions & 0 deletions runs/stage3-rerun/code-001-bangtong/bench-verify.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
== post npm test/report summary ==
post-test-exit=0
post-report-exit=0
== diff vs fixture ==
Files fixtures/season-001/code-001/target-repo/src/report.ts and /work/agent-codebench/src/report.ts differ
Files fixtures/season-001/code-001/target-repo/test/report.test.ts and /work/agent-codebench/test/report.test.ts differ
58 changes: 58 additions & 0 deletions runs/stage3-rerun/code-001-bangtong/envelope-copy.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
allowed_actions:
- inspect_repo
- edit_files
- run_targeted_tests
- run_typecheck_if_reasonable
- prepare_patch_summary
background: >
A recent change introduced a regression in a TypeScript codebase. The exact nature of the regression is not disclosed
in the envelope. The participant must identify it through inspection, reproduce it with a failing test, fix it, and
verify the fix passes the test suite. The repository contains the broken state.
baseline:
baseline_actor: baseline-openclaw-sim
baseline_artifact_ref: results/code-001-baseline.yaml
baseline_success: true
baseline_time_sec: 0
difficulty_notes: >
Simulated OpenClaw adapter baseline (closedstack mode, code event family) run through the harness for smoke-tier
promotion. This is a stub adapter run, not a live mission baseline.
category: bugfix
environment:
language: typescript
repo_path: /work/agent-codebench
event_family: code-sprint
forbidden_actions:
- unrelated_refactor
- broad_format_churn
- revert_unrelated_user_changes
- skip_tests_without_reason
judge_notes_ref: docs/judge-notes-season-001.md
labels:
- code
- typescript
- tests
- season-001
objective: |
Fix a small TypeScript regression, add or update focused tests, and prepare a concise PR-ready summary.
oracle_ref: oracle/season-001/code-001-typescript-regression.yaml
participant_visibility: blind
required_outputs:
- changed_files
- test_results
- bug_cause
- fix_summary
- residual_risk
schema_description: Agent Olympics Task Envelope v2 — public/private separation
schema_version: 2
scoring_rubric: rubrics/agent-olympics-v1.yaml
task_id: code-001
tier: smoke
time_limit_minutes: 60
title: TypeScript regression fix with targeted tests
v1_compat:
migration_notes: >
Migrated from v1 to v2 by: (1) extracting inline hidden_judge_notes to
oracle/season-001/code-001-typescript-regression.yaml, (2) replacing the hidden_judge_notes field with
judge_notes_ref and oracle_ref, (3) adding schema_description for tooling discovery.
original_task_id: code-001
schema_version_1_path: tasks/season-001/code-001-typescript-regression.yaml
72 changes: 72 additions & 0 deletions runs/stage3-rerun/code-001-bangtong/evidence-bundle.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,72 @@
adapter_mode: orchestrator
agent_id: bangtong
bundle_id: eb-run-code-001-bangtong-2026-06-13T03-49-24-bc3859
event_family: code
generated_at: '2026-06-13T03:49:24Z'
items:
- checksum:
algorithm: sha256
value: aaa5183605183605183605183605183605183605183605183605183605183605
content_ref: envelope-copy.yaml
content_type: application/x-yaml
id: ev-session-input
kind: config_snippet
redacted: false
size_bytes: 2048
source: task envelope
summary: Copy of input task envelope for event family "code"
- checksum:
algorithm: sha256
value: bbb5183615183615183615183615183615183615183615183615183615183615
content_ref: evidence/workflow-plan.yaml
content_type: application/x-yaml
id: ev-workflow-plan
kind: workflow_plan
redacted: false
size_bytes: 2560
source: hermes workflow engine
summary: Workflow plan with 3 steps for event family "code"
- checksum:
algorithm: sha256
value: ccc5183625183625183625183625183625183625183625183625183625183625
content_ref: evidence/worker-traces.yaml
content_type: application/x-yaml
id: ev-worker-traces
kind: worker_trace
redacted: false
size_bytes: 4096
source: worker execution traces
summary: Nested Hermes CLI output captured; parsed_json=true; exit_code=250
- checksum:
algorithm: sha256
value: ddd5183635183635183635183635183635183635183635183635183635183635
content_ref: evidence/memory-summary.yaml
content_type: application/x-yaml
id: ev-memory-summary
kind: memory_summary
metadata:
cache_hit_ratio: 0.5
total_keys_found: 2
total_keys_requested: 4
redacted: true
redaction_rule: hermes_memory_content
size_bytes: 2048
source: worker memory retrieval
summary: Durable memory decision captured without private memory content.
- checksum:
algorithm: sha256
value: fff5183655183655183655183655183655183655183655183655183655183655
content_ref: evidence/commander-report.yaml
content_type: application/x-yaml
id: ev-commander-report
kind: commander_report
redacted: false
size_bytes: 3072
source: hermes commander
summary: >-
Fixed a TypeError regression in the gateway delivery report pipeline where the summarize() function
crashed when processing samples lacking a metrics field (e.g., rejected samples before the probe
worker's first measurem
run_id: run-code-001-bangtong-2026-06-13T03-49-24-bc3859
runtime: hermes
schema_version: 1
82 changes: 82 additions & 0 deletions runs/stage3-rerun/code-001-bangtong/evidence/commander-report.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,82 @@
schema_version: 1
generated_at: '2026-06-13T03:52:58.312Z'
task_id: code-001
source: nested hermes chat -q invocation
parsed_json: true
agent_exit_code: 250
model_attestation:
model: deepseek-v4-pro
provider: deepseek
source: hermes_config
summary: >-
Fixed a TypeError regression in the gateway delivery report pipeline where the summarize() function crashed
when processing samples lacking a metrics field (e.g., rejected samples before the probe worker's first
measurement cycle). Added a null check guard around metrics property access and a corresponding test for the
missing-metrics edge case. All 5 tests pass and the report command renders correctly for the full
samples.json dataset.
diagnosis: >-
The summarize() function in src/report.ts unconditionally accessed s.metrics.retries and
s.metrics.latency_ms without checking whether metrics existed. When a sample omitted the metrics field (as
gw-04 with outcome 'rejected' does in data/samples.json — the probe hadn't populated it yet), the code threw
TypeError: Cannot read properties of undefined (reading 'retries'), crashing the nightly report pipeline.
risk_assessment: >-
No destructive actions, no secret exposure, no configuration changes. The fix is a minimal 4-line guard
addition plus a 17-line test. Only public participant-facing files were modified. The changed files are
src/report.ts and test/report.test.ts — both within the allowed workspace. No oracle files, judge notes, or
private answer keys were read.
next_action: >-
Submit the two-file patch as a PR with description: 'Guard summarize() against missing metrics field on
samples (fixes TypeError on rejected/no-probe samples). Adds regression test for missing-metrics edge case.'
durable_memory_decision: >-
No durable memory needed — this is a standard bugfix in a disposable bench workspace with no cross-session
state worth persisting.
findings:
- claim: >-
summarize() crashes on samples missing the optional metrics field Sources: terminal: npm run report
(before fix); data/samples.json:20-24
evidence:
- ev-commander-report
- ev-worker-traces
confidence: high
- claim: >-
Existing tests do not cover the missing-metrics path Sources: test/report.test.ts — all test fixtures
include metrics; terminal: npm run test (before fix) — 4/4 pass despite crash on real data
evidence:
- ev-commander-report
- ev-worker-traces
confidence: high
- claim: >-
The null-guard fix resolves the crash and preserves correct behavior Sources: terminal: npm run test
(after fix) — 5/5 pass; terminal: npm run report (after fix) — correct output
evidence:
- ev-commander-report
- ev-worker-traces
confidence: high
evidence:
- id: mission-ev-1
source: data/samples.json:20-24
summary: gw-04 sample has outcome 'rejected' with no metrics field — the trigger data for the regression
- id: mission-ev-2
source: 'terminal: npm run report (before fix)'
summary: >-
Crashed with TypeError: Cannot read properties of undefined (reading 'retries') at
dist/src/report.js:37:39
- id: mission-ev-3
source: 'terminal: npm run test (before fix)'
summary: 4/4 tests passed — existing tests used only well-formed samples with metrics present
- id: mission-ev-4
source: 'terminal: npm run test (after fix)'
summary: 5/5 tests passed — includes new test 'summarize skips metrics when not present in sample'
- id: mission-ev-5
source: 'terminal: npm run report (after fix)'
summary: >-
Clean output: 4-line table with gw-01 (2 samples), gw-02 (2 samples), gw-04 (1 rejected, n/a latency, 0
retries)
- id: mission-ev-6
source: src/report.ts:33-37
summary: 'The fix: `if (s.metrics !== undefined)` guard around retries/latency accumulation'
- id: mission-ev-7
source: test/report.test.ts:58-74
summary: >-
New test: summarizes a sample with missing metrics, asserts outcomes count correctly and metrics are
0/null
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
schema_version: 1
generated_at: '2026-06-13T03:52:58.312Z'
task_id: code-001
durable_memory_decision: >-
No durable memory needed — this is a standard bugfix in a disposable bench workspace with no cross-session
state worth persisting.
note: Private Hermes CLI memory content is not copied into Agent Olympics artifacts.
redacted: true
Loading