Skip to content

feat: production-hardening Track A/B (Phases 0–8, 9 except live spend) - #2

Merged
RickZee merged 17 commits into
mainfrom
feature/phase-harness-alignment
Sep 14, 2026
Merged

RickZee merged 17 commits into
mainfrom
feature/phase-harness-alignment

Conversation

@RickZee

@RickZee RickZee commented Sep 13, 2026

Copy link
Copy Markdown
Owner

Summary

  • Land production-hardening Tracks A and B (delete/truth/gates/control-plane/container/boundaries/complexity/guards) plus test hygiene.
  • Leave 9.1 (live benchmark, $25) human-triggered.
  • CI should run because this PR targets main.

Test plan

  • Lint / Test (3.11, 3.12) / Web UI E2E / Security green
  • pip-audit artifact downloadable from the Security job
  • No live model calls

RickZee and others added 17 commits September 13, 2026 09:19
Make the harness/decomposition thesis measurable (solo,
harnessed_solo, reference, ai_team) and record the suite
failures that would have gone red in CI.

Co-authored-by: Cursor <cursoragent@cursor.com>
The implementation landed without updating the checklist. Live
spends 4.3, 4.5, 5.4, and 11.4 stay open.

Co-authored-by: Cursor <cursoragent@cursor.com>
The scheduled run only skipped after finding no API keys. Keep
workflow_dispatch so it can be re-enabled when secrets exist.

Co-authored-by: Cursor <cursoragent@cursor.com>
Make design tokens the only literals, complete tab/dialog patterns, and
keep frozen testids except the documented nav-open-run removal.

Co-authored-by: Cursor <cursoragent@cursor.com>
Record the corpus audit campaign, journal notes, and related resource
links alongside the living checklists.

Co-authored-by: Cursor <cursoragent@cursor.com>
Reviewers need the 2.4/3.7/7.4 shots on the branch; the spec's "attach to the PR" only works if the PNGs are tracked next to BASELINE.md.

Co-authored-by: Cursor <cursoragent@cursor.com>
Wire the harness claims, lock the control plane behind a token, ship a
real UI container, and add repo guards so the docs cannot drift from the
code. Keep .archive/ with a retention README; defer Track B.

Co-authored-by: Cursor <cursoragent@cursor.com>
Move CrewAI agents/crews/tasks/flows under crewai_backend, break eval→backend
imports via core contracts, add deprecation shims, and dissolve utils/.

Co-authored-by: Cursor <cursoragent@cursor.com>
Extract RunOptions, MCP tool groups, evaluate_gate criteria, and web
routers so the three worst functions drop under 120 lines. Ratchet
58/7/15 → 55/4/14.

Co-authored-by: Cursor <cursoragent@cursor.com>
Skip-on-failure sites now fail; ToolBus and Settings reset after each test
so shuffled unit order is green. GitHub Release and the $25 benchmark still
need a human.

Co-authored-by: Cursor <cursoragent@cursor.com>
Feature-branch pushes do not run CI until a PR targets main. Do not hang
on gh auth or GitHub MCP.

Co-authored-by: Cursor <cursoragent@cursor.com>
PROMPTS.md pointed at a gitignored archive file that only existed locally.
Require git-tracked link targets. Stop RUN chown -R after COPY (2.55 GB).
Size ratchet 900→1800 MiB: GHA inspect Size is uncompressed layers.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Preserve the 1600s HITL receipt and the Husain debug loop so
open coding has a real run instead of another fixture dump.

Co-authored-by: Cursor <cursoragent@cursor.com>
Fixture-only suites can look alive while the real corpus is
all not_applicable. Measure that at $0 before claiming a rate.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Open coding never happened because the sitting was an editor
round-trip per trace. Ship a gated local page and a bundle
export so the human loop has a surface that matches the spec.

Co-authored-by: Cursor <cursoragent@cursor.com>
@RickZee
RickZee merged commit 9074ae0 into main Sep 14, 2026
10 checks passed
RickZee added a commit that referenced this pull request Sep 30, 2026
Tier 1 fixes #2 and #4 from the adversarial review.

Model confound (#2): the published n=5 table compared framework+model bundles
(deepseek for crewai/langgraph, Claude for the SDK), not frameworks. Added a
`smoke-claude` profile pinning every role to one Claude model — the same-model
control — and a `--team` flag on the batch runner that warns loudly when the
profile is mixed-model. Relabeled the README, COMPARISON_RESULTS, and
TEAM_PROFILES tables as confounded, added the Wilson CIs (5/5 -> 57-100%,
1/5 -> 4-62%, which overlap), and softened the "second-most-reliable" /
"consistency champion" prose that the intervals do not support. Documented the
caveat that the claude-agent-sdk backend ignores OpenRouter overrides, so the
control is exact only for crewai-vs-langgraph.

Trivial-task overreach (#4): the batch only ran add(a,b)+pytest, then
generalized to production verdicts. Added a `--demo` flag with per-tier
timeout scaling and a canary warning; verdicts are meant to come from
demos/02_todo_app, not the smoke canary. Batch output is now a provenance
bundle (team, same_model, demo, is_canary, timestamp) instead of a bare list.

CONTRIBUTING now tells contributors to run --team smoke-claude
--demo demos/02_todo_app and paste the interval-aware output.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@RickZee
RickZee deleted the feature/phase-harness-alignment branch September 30, 2026 21:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant