feat: deploy verified local AgentTeams research stack - #3
Merged
Conversation
Pin and bootstrap the official Controller and Manager, provision four Matrix-backed Workers, connect PostgreSQL/API/Bridge/Web, and fail closed on dirty sources or invalid health payloads.
Use the requested provider model by default and allocate larger bounded outputs to the matrix architect and independent reviewer.
Keep the GPU and cloud-memory boundaries honest while ensuring cockpit anchors land below sticky operator controls on desktop and mobile.
Document the one-command deployment, real Matrix collaboration receipt, browser QA, custom-input fail-closed result, and deferred GPU/Nexa claims. Co-Authored-By: OpenAI Codex <noreply@openai.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Official AgentTeams and local control plane
Live experts and Web truth
agnes-2.5-proby default with bounded per-role output budgets.Acceptance evidence
Test Coverage
Pre-Landing Review
Four findings were fixed before landing: malformed Bridge payload handling, dirty official checkout detection, first-run base-image availability, and atomic private-env persistence. Final review found no open issues.
Design Review
One browser-visible finding was fixed: in-page anchors were hidden behind the sticky task/operator controls. Real Chrome verification passed at 1440x900 and 390x844 with no horizontal overflow or console errors.
Eval Results
No prompt-template files matched the repository eval-suite trigger. A real custom-input call was preserved separately: PI and Scout completed, Architect violated the JSON contract twice, and the run failed closed before Matrix/GPU execution.
Plan Completion
No standalone plan file was detected. The user-requested pre-GPU deployment scope is implemented; physical GPU execution remains intentionally deferred.
Verification Results
Test plan
.venv/bin/pytest(683 passed, 1 skipped)npm test -- --run(26 passed)npm run buildpython3 scripts/deploy_local_live_stack.py verifyDocumentation
docs/architecture.md: documents the verified local Controller/Manager/Team/Worker/Matrix topology and preserves the GPU boundary.docs/claims-evidence.md: upgrades infrastructure collaboration toLIVE_LOCALwhile keeping full workflow and GPU claims unverified.docs/competition-mapping.md,docs/semifinal-scorecard.md, anddocs/semifinal-change-log.md: align the competition scorecard and remaining hard gates with the frozen acceptance receipt.docs/input-output-contract.md,docs/judge-feedback-implementation.md, anddocs/final-submission-20260902.md: separate live local collaboration infrastructure from the not-run experiment chain.apps/api/README.mdandintegrations/agentteams/README.md: explain the model-plane versus AgentTeams truth boundaries.docs/acceptance/2026-09-02-final-live-model.md: marks the older model-only record as a historical snapshot and repairs its input link.