Skip to content

fix: close real-model planning feedback and reviewer order gaps - #59

Draft
shenjiecode wants to merge 1 commit into
developfrom
fix/scenario-quality-feedback
Draft

fix: close real-model planning feedback and reviewer order gaps#59
shenjiecode wants to merge 1 commit into
developfrom
fix/scenario-quality-feedback

Conversation

@shenjiecode

Copy link
Copy Markdown
Contributor

Summary

Closes engineering gaps found by the frozen Qwen real-model comparison; does not claim the quality evaluation has passed.

  • Supply the complete five-field scenario template, stable-ID/status rules, and actionable bounded patch/schema errors without exposing raw Git stderr or rewriting patches.
  • Validate patch + resulting scenario states + execution selection at planning completion. Feed errors back at most twice within the same isolated Main Session; keep existing fail-closed checks and clean the temporary worktree.
  • Enforce Reviewer plan/existing-patch/original-image order before draft access or review submission. Failed image reads remain blocking; clarify that an execution list does not waive required capability coverage.
  • Test request recovery separately from Session completion; keep artifact and human quality checks distinct in evaluation tooling.

Verification

Using the lockfile-matching Docker quality image with current release inputs mounted:

  • format / lint / typecheck: passed
  • unit + integration: 162/162, 26 files
  • build + headless E2E: passed
  • local acceptance including production Pi paths: passed (live/release blocked)
  • tests cover parser-valid role template, rejected premature reads without exposing content, existing patch requirement, failed/zero-image paths, bounded same-Session repair, explicit missing-newline feedback, and recovered request classification.

No credentials, production data or raw evaluation outputs are committed. No target website operations, target remote publication, deployment or release.

Real-model status

The previous 72 attempts and all failures remain intact. New candidate a1bdd68 is frozen for a fresh 72-Session comparison against the original 09dbed0 baseline, using unchanged website-derived fixtures/rubric and qwen3.7-plus with Thinking off. It is running under a 1200-request cap with no pay-as-you-go fallback. Results and independent human adjudication are still pending; this PR is draft while retesting proceeds.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant