Tracking issue distilled from PR #2123 (closed; the restructure was under the eval bar for brainstorming content, but the underlying behavior is worth tracking).
The behavior: during brainstorming, the agent asks the human questions the repo could answer, or delegates expert design parameters ('what should the retry budget be?') to the human instead of orienting in the codebase and arriving with informed proposals. The checklist's 'Explore project context' step exists, but nothing instructs the agent to classify unknowns — which questions the repo/docs/conventions answer, versus which genuinely require human judgment or preference.
Evidence state: the motivating session in #2123 showed the punt happening, but its baseline reproduced only 1-in-3 (2 of 3 baseline sessions on current dev already behaved correctly), so frequency is unclear. More transcripts of this failure on current main would sharpen it.
Relationship to other work: open PR #2116 (research prior art before design) touches brainstorming's front end and may shift this behavior; re-baseline after it lands. Any fix here is tuned-content work requiring the full eval treatment.
— Claude Fable 5, Claude Code 2.1.228, triaging on behalf of @obra
Tracking issue distilled from PR #2123 (closed; the restructure was under the eval bar for brainstorming content, but the underlying behavior is worth tracking).
The behavior: during brainstorming, the agent asks the human questions the repo could answer, or delegates expert design parameters ('what should the retry budget be?') to the human instead of orienting in the codebase and arriving with informed proposals. The checklist's 'Explore project context' step exists, but nothing instructs the agent to classify unknowns — which questions the repo/docs/conventions answer, versus which genuinely require human judgment or preference.
Evidence state: the motivating session in #2123 showed the punt happening, but its baseline reproduced only 1-in-3 (2 of 3 baseline sessions on current dev already behaved correctly), so frequency is unclear. More transcripts of this failure on current
mainwould sharpen it.Relationship to other work: open PR #2116 (research prior art before design) touches brainstorming's front end and may shift this behavior; re-baseline after it lands. Any fix here is tuned-content work requiring the full eval treatment.
— Claude Fable 5, Claude Code 2.1.228, triaging on behalf of @obra