Repository navigation
feat: production-hardening Track A/B (Phases 0–8, 9 except live spend) - #2
Merged
Merged
Conversation
Make the harness/decomposition thesis measurable (solo, harnessed_solo, reference, ai_team) and record the suite failures that would have gone red in CI. Co-authored-by: Cursor <cursoragent@cursor.com>
The implementation landed without updating the checklist. Live spends 4.3, 4.5, 5.4, and 11.4 stay open. Co-authored-by: Cursor <cursoragent@cursor.com>
The scheduled run only skipped after finding no API keys. Keep workflow_dispatch so it can be re-enabled when secrets exist. Co-authored-by: Cursor <cursoragent@cursor.com>
Make design tokens the only literals, complete tab/dialog patterns, and keep frozen testids except the documented nav-open-run removal. Co-authored-by: Cursor <cursoragent@cursor.com>
Record the corpus audit campaign, journal notes, and related resource links alongside the living checklists. Co-authored-by: Cursor <cursoragent@cursor.com>
Reviewers need the 2.4/3.7/7.4 shots on the branch; the spec's "attach to the PR" only works if the PNGs are tracked next to BASELINE.md. Co-authored-by: Cursor <cursoragent@cursor.com>
Wire the harness claims, lock the control plane behind a token, ship a real UI container, and add repo guards so the docs cannot drift from the code. Keep .archive/ with a retention README; defer Track B. Co-authored-by: Cursor <cursoragent@cursor.com>
Move CrewAI agents/crews/tasks/flows under crewai_backend, break eval→backend imports via core contracts, add deprecation shims, and dissolve utils/. Co-authored-by: Cursor <cursoragent@cursor.com>
Extract RunOptions, MCP tool groups, evaluate_gate criteria, and web routers so the three worst functions drop under 120 lines. Ratchet 58/7/15 → 55/4/14. Co-authored-by: Cursor <cursoragent@cursor.com>
Skip-on-failure sites now fail; ToolBus and Settings reset after each test so shuffled unit order is green. GitHub Release and the $25 benchmark still need a human. Co-authored-by: Cursor <cursoragent@cursor.com>
Feature-branch pushes do not run CI until a PR targets main. Do not hang on gh auth or GitHub MCP. Co-authored-by: Cursor <cursoragent@cursor.com>
PROMPTS.md pointed at a gitignored archive file that only existed locally. Require git-tracked link targets. Stop RUN chown -R after COPY (2.55 GB). Size ratchet 900→1800 MiB: GHA inspect Size is uncompressed layers. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Preserve the 1600s HITL receipt and the Husain debug loop so open coding has a real run instead of another fixture dump. Co-authored-by: Cursor <cursoragent@cursor.com>
Fixture-only suites can look alive while the real corpus is all not_applicable. Measure that at $0 before claiming a rate. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Open coding never happened because the sitting was an editor round-trip per trace. Ship a gated local page and a bundle export so the human loop has a surface that matches the spec. Co-authored-by: Cursor <cursoragent@cursor.com>
RickZee
added a commit
that referenced
this pull request
Sep 30, 2026
Tier 1 fixes #2 and #4 from the adversarial review. Model confound (#2): the published n=5 table compared framework+model bundles (deepseek for crewai/langgraph, Claude for the SDK), not frameworks. Added a `smoke-claude` profile pinning every role to one Claude model — the same-model control — and a `--team` flag on the batch runner that warns loudly when the profile is mixed-model. Relabeled the README, COMPARISON_RESULTS, and TEAM_PROFILES tables as confounded, added the Wilson CIs (5/5 -> 57-100%, 1/5 -> 4-62%, which overlap), and softened the "second-most-reliable" / "consistency champion" prose that the intervals do not support. Documented the caveat that the claude-agent-sdk backend ignores OpenRouter overrides, so the control is exact only for crewai-vs-langgraph. Trivial-task overreach (#4): the batch only ran add(a,b)+pytest, then generalized to production verdicts. Added a `--demo` flag with per-tier timeout scaling and a canary warning; verdicts are meant to come from demos/02_todo_app, not the smoke canary. Batch output is now a provenance bundle (team, same_model, demo, is_canary, timestamp) instead of a bare list. CONTRIBUTING now tells contributors to run --team smoke-claude --demo demos/02_todo_app and paste the interval-aware output. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
main.Test plan