Skip to content

Commit 3d115a3

Browse files
author
AI Assistant
committed
docs(memory-bank): close seven-task integration
Record completed local feature commits, final Windows and Linux integration evidence, browser checks, and remaining verification limits. Keep publication and remote mutation separately authorized. Co-authored-by: AI Assistant <ai@example.com>
1 parent 3b04d46 commit 3d115a3

4 files changed

Lines changed: 236 additions & 246 deletions

File tree

‎.memory-bank/activeContext.md‎

Lines changed: 108 additions & 111 deletions
Original file line numberDiff line numberDiff line change
@@ -9,119 +9,119 @@ source: current task evidence
99

1010
## Current focus
1111

12-
Client-specific adapters (task 06, correction round 2). Read-only
13-
`Get-CopilotAtelierClientAdapter` over private
14-
`Get-CopilotAtelierClientContract`, `ConvertFrom-`/`ConvertTo-`
15-
`CopilotAtelierAgentFrontmatter`, `ConvertTo-CopilotAtelierClientAgent`,
16-
`Export-CopilotAtelierClientAdapterArtifact`, plus the build task.
17-
18-
**Discovery is not parity, and the mismatch is specific.** The shipped profiles
19-
are authored in the VS Code shape. The published custom agents configuration the
20-
Copilot CLI follows documents one model string, a closed set of tool aliases,
21-
and no subagent allow-list, handoff, or argument hint — and it *ignores* an
22-
unrecognized tool name. `software-engineer` declares 45 tool identifiers; 5 map.
23-
Scope is VS Code and the CLI; no cloud client is claimed.
24-
25-
The VS Code files stay the only source; only frontmatter is rewritten. Four
26-
rules are tests, not prose: every mapping is explicit and an unmapped identifier
27-
is an error; frontmatter is a strict YAML subset that rejects an unknown or
28-
duplicate field rather than dropping it; **an `execute/` prefix is a namespace,
29-
not execution authority** — only `execute/runInTerminal` may reach the execute
30-
alias, so `execute/getTerminalOutput`, `runTests`, task runners, and `useMcp`
31-
stay unsupported; and a restriction that cannot be expressed removes what it
32-
guards, so the variant loses the `agent` tool. Every other mapping stays inside
33-
its capability class. A mandatory capability or workflow that cannot be provided
34-
fails the composition and emits nothing.
35-
36-
**`review: on` and `cycle: full` are refused, not degraded.** The composed file
37-
carries an additive limitation section, between explicit markers with the shared
38-
body's SHA-256 in the end marker, telling the agent to refuse those modes and
39-
return to VS Code. The body stays byte-identical and last.
40-
41-
Neither client is runtime verified: no receipt is bound to an artifact, so both
42-
are `StructurallyChecked`. The VS Code 1.136.1 observation is kept as historical
43-
source-profile evidence. The CLI is **not installed**;
44-
`docs/client-adapter-evals.md` carries the authored, unexecuted cases.
45-
46-
Variants land in `output/clientAdapters/<client>/` and are **never deployed**.
47-
The build task owns that directory through a manifest and removes only what it
48-
generated; an unowned directory, a reserved build name, a non-child path, or a
49-
reparse point is refused rather than deleted.
50-
51-
**Round 2 hardened the exporter, and only the exporter.** Every path component —
52-
output root, artifact directory, manifest, client directory, generated file,
53-
destination — goes through `Assert-CopilotAtelierRegularPath`, because a link
54-
*between* root and leaf redirects a delete just as well as one at either end.
55-
The whole operation is validated before the first mutation, so a late unsafe
56-
manifest entry no longer arrives after a valid earlier one was already deleted.
57-
Ownership is proved by content: schema 2 records a SHA-256 per generated file,
58-
so an edited generated file, an unowned destination collision — identical
59-
content included — and a names-only `schema 1` manifest are all refused.
60-
61-
Red 46/103 before round 1, then 103/103. Round 2 red 15/121, then 121/121 with
62-
no skips (elevated, so the link cases ran). Uncommitted; `review: off` —
63-
recommend `review: on` for the permission mapping and the ownership model.
12+
Completed the seven-task series on `main`. Tasks 01-06 are committed at
13+
`c0c7166`, `8e40815`, `09416a4`, `bb66d41`, `e04345e`, and `288a4ad`; the
14+
records commit is `555c260`. The earlier CI fix is `35fa926`. The user requested
15+
the dependency ignore rule, completion of task 07, and final integration.
16+
Remote mutation, paid model evaluations, and independent review remain off.
17+
18+
Task 07 (`tools/plan-review`) is committed in `3b04d46`. Integration fixed npm
19+
commands, ambiguous duplicate-heading anchors, linked feedback roots, missing
20+
launch-time input checks, invalid UTF-8, mobile stale-status overflow, document
21+
selection, unsent drafts, connection failures, and stale verdict-dialog identity.
22+
Final evidence: Windows combined gate 1,810 passed/67 skips, 90.72% coverage;
23+
Linux Unit/QA gate 1,707 passed/114 skips/56 tag exclusions, 90.42% coverage.
24+
Node tests: 133 passed on each OS; Edge desktop/mobile: 30 passed. All have
25+
zero failures. Focused repository checks passed 158; native AST, ScriptAnalyzer,
26+
syntax, and dependency audit checks passed. Memory budgets are within limits.
27+
Existing simulated-backend and near-budget warnings remain. macOS, live
28+
OneDrive, and paid model-backed evaluation are not verified by these runs.
29+
30+
**It is opt-in in the strong sense.** It is absent from `CustomizationDirectory`
31+
so the built module never carries it, no file under `source/` mentions it, and
32+
its dependencies are installed by hand in `tools/plan-review`. The repository
33+
gate runs `tests/PlanReview.Tests.ps1`, which asserts the packaging isolation
34+
and the trust-boundary invariants from the source and executes the
35+
dependency-free half of the Node suite when `node` is present — it never runs
36+
`npm install`, so `npm` is not imposed on an ordinary install.
37+
38+
**A browser verdict is feedback, not sign-off.** An HTTP request proves a
39+
content hash was posted, not who posted it. Verdicts are stored as
40+
`authority: "local-http-feedback"` under
41+
`approvalAuthority: "chat-sign-off-required"`, the header is repeated on the
42+
response and shown above the document, and there is no endpoint that writes a
43+
Decision record, fires a handoff, or runs a command. A `--state` root inside
44+
`.memory-bank/decisions` is refused at launch. The architect agent body points
45+
at the tool and says it does not replace the sign-off; its frontmatter is
46+
untouched, so the task-06 strict parser still accepts it.
47+
48+
Identity is content. A revision is the SHA-256 of the bytes; a section key is
49+
its heading slug plus an occurrence ordinal. A comment reloads as *current*,
50+
*revised*, or *orphaned*, and an orphaned comment is listed separately and never
51+
re-anchored by heading text. A verdict whose hash no longer matches is reported
52+
as invalidated, and a post against a stale hash is refused with `409`. Loopback
53+
is a reachability reduction, not authorization. A non-loopback bind is
54+
refused; every mutation needs Host, exact Origin, `Sec-Fetch-Site`, JSON content
55+
type, the per-launch session cookie, and a matching CSRF token, so a cookie from
56+
an earlier launch fails against the next one even on the same store. Documents
57+
are authorized at launch and addressed by an opaque id; paths are realpath
58+
checked with ancestor reparse rejection at read time, not only at launch. Raw
59+
HTML is off at `markdown-it` with DOMPurify after it, images are never fetched,
60+
Mermaid runs `securityLevel: 'strict'`, and the CSP is `default-src 'none'` with
61+
`script-src 'self'`. Earlier task-07 results were 124 Node tests and 22 browser
62+
checks; they are historical evidence, not the final integration results.
63+
Recommend `review: on` for the local HTTP, persistence, and approval boundaries.
64+
65+
## Previous focus: client-specific adapters
66+
67+
Task 06 is committed in `288a4ad`. VS Code profiles remain the source;
68+
the adapter rewrites strictly parsed frontmatter through explicit capability
69+
mappings, rejects unsupported grants, and preserves the body bytes. CLI
70+
`review: on` and `cycle: full` are refused in the generated variant. Neither
71+
client is runtime verified; authored cases remain in `docs/client-adapter-evals.md`.
72+
Generated variants are undeployed build artifacts under `output/clientAdapters`.
73+
The exporter preflights every path, rejects linked ancestors and unowned or
74+
modified collisions, and records SHA-256 ownership in schema 2. Round 2 passed
75+
121 focused regressions; the changelog retains earlier correction evidence.
6476

6577
## Previous focus: changed-file validation
6678

6779
Task 05 added the `changed-file-validation` Skill: an opt-in, bounded validation
6880
pass over one work batch, collected manually because no documented hook event
6981
reports an edit contract this implementation has verified. Validators read an
70-
isolated snapshot under a generated name, and a receipt binds to those bytes
71-
plus a plan identity hashing the linter entry point and the shipped checker
72-
code. Parse and PSScriptAnalyzer run in an owned child worker with a wall clock
73-
and inline settings. `Markdown.NativeStructure` is `coverage=partial` and never
74-
stands in for markdownlint. 12 regressions red, then 85/0/0.
82+
isolated snapshot, and a receipt binds to those bytes plus a plan identity
83+
hashing the linter entry point and the shipped checker code. Parse and
84+
PSScriptAnalyzer run in an owned child worker with a wall clock and inline
85+
settings. `Markdown.NativeStructure` is `coverage=partial` and never stands in
86+
for markdownlint. 12 regressions red, then 85/0/0.
7587

7688
## Previous focus: Skill health report
7789

7890
Task 04 added read-only `Get-CopilotAtelierSkillHealth` plus private
7991
`Import-CopilotAtelierSkillObservation`, `Measure-CopilotAtelierSkillHealth`,
80-
and the shared `Get-CopilotAtelierBoundedFile` enumerator.
81-
82-
The client-event contract is **verified, not asserted**: each of the eight
83-
documented hook events is checked against the deployed authoring Instruction,
84-
and the report publishes `VerificationState` and `VerificationScope` — *no
85-
reliable Skill-activation contract verified for this implementation*. No capture
86-
is implemented, it is off by default, and missing telemetry is unknown, not
87-
zero. Evaluation evidence is read in the shapes `agent-evals` defines: run
88-
output counts only through a validated provenance sidecar, a graded summary only
89-
when it reconciles exactly with the bounded verdicts in `assertion_results`, and
90-
disagreeing copies surface as `ConflictingRun`. `RetirementReview` needs an
91-
explicit `coverage` declaration and is never raised for a mandatory Skill.
92-
Focused suites red 29 then green 81/0; canonical `build, test` green at 1492
93-
passed, 0 failed (`%TEMP%\ca-sh2-final2.log`).
92+
and the shared `Get-CopilotAtelierBoundedFile` enumerator. The client-event
93+
contract is **verified, not asserted**: each documented hook event is checked
94+
against the deployed authoring Instruction, and the report publishes *no
95+
reliable Skill-activation contract verified for this implementation*. Missing
96+
telemetry is unknown, not zero. Run output counts only through a validated
97+
provenance sidecar, a graded summary only when it reconciles exactly with
98+
`assertion_results`, and disagreeing copies surface as `ConflictingRun`. Red 29
99+
then 81/0; canonical `build, test` 1492 passed, 0 failed.
94100

95101
## Previous focus: reviewed learning inbox
96102

97103
Task 03 added `reviewed-learning-inbox`: an on-demand, project-scoped review
98104
queue whose store at `.memory-bank/learning-inbox/candidates.json` sits outside
99-
the routed base, every Skill description, and the deployed tree. A selected
100-
artifact is an untrusted observation, never a directive; one content rule runs
101-
at intake and again at promotion. Promotion is append-only and hash-gated —
102-
`-Approve` plus the preview SHA-256 — and round 1 added parent-chain reparse
103-
guards, evidence binding, a byte-preserving exclusive append, an atomic locked
104-
store, and verified-block repeat handling. Red 18 of 62 then 62/0/0.
105+
the routed base and the deployed tree. A selected artifact is an untrusted
106+
observation, never a directive; one content rule runs at intake and again at
107+
promotion, which is append-only and hash-gated on `-Approve` plus the preview
108+
SHA-256. Red 18 of 62 then 62/0/0.
105109

106110
## Previous focuses
107111

108-
- **Installation profiles (task 02).** `-InstallationProfile` (`complete`
109-
default, `engineering`, `research`, `document-processing`) plus
112+
- **Installation profiles (task 02).** `-InstallationProfile` plus
110113
`-IncludeSkill`/`-ExcludeSkill` on Install, Update, and Setup, with
111-
`Get-CopilotAtelierProfile` and an `InstallationProfile` field on the
112-
`Test-CopilotAtelier` result. Only Skills are selectable; `memory-bank`,
114+
`Get-CopilotAtelierProfile`. Only Skills are selectable; `memory-bank`,
113115
`long-running-job-monitor`, and `agent-security-review` are mandatory because
114116
deployed Instructions and shipped agents load them by name. The selection is
115-
an additive optional `Selection` field inside schema 1, omitted for a complete
116-
installation, and the plan filters whole top-level folders.
117+
an additive optional `Selection` field inside schema 1.
117118
- **Footprint reporting (task 01).** `Get-CopilotAtelierFootprint` reports
118119
potential automatic loading contingent on discovery, never "always loaded",
119-
guards every mapped root, and fails closed on ambiguous frontmatter. The
120-
shared `Get-CopilotAtelierDirectoryMap` keeps installer and report aligned.
120+
guards every mapped root, and fails closed on ambiguous frontmatter.
121121
- **CI run 34061934611 and the deployment review.** Fixed on `main` at
122-
`5acb69d`, where M1-M5 and L1-L6 also landed with zero Blocker/Major findings;
123-
the per-ID ledger stays in `assessment-log.md`. Windows 1,266 passed, Linux
124-
1,172, 5.1 focused 120; macOS and live OneDrive remain unverified locally.
122+
`35fa926`; M1-M5 and L1-L6 landed at `5acb69d` with zero Blocker/Major findings;
123+
the per-ID ledger stays in `assessment-log.md`. macOS and live OneDrive remain
124+
unverified locally.
125125
- **Agent Plugins 1.0.** VS Code's documentation confirms all four Copilot-only
126126
component paths under `com.github.copilot/`, so Decision 0023's layout is
127127
documented upstream; only the accepted cross-type-link mismatch remains.
@@ -132,18 +132,17 @@ Two bulk PowerShell read-modify-write passes over this working tree replaced
132132
whole file contents with a monoalphabetic substitution cipher (`instructions`
133133
→ `nnkteuotnonk`, `applyTo` → `aeelyTo`), 129 files each time. Both were caught
134134
and restored from git; no corruption reached a commit. It is asynchronous: the
135-
script's own byte-exact read-back passed for all 175 files and `git diff` showed
136-
it afterwards, so a verify-after-write loop cannot detect it. Until the cause is
137-
found, edit through the editor tooling and treat any scripted bulk rewrite as
138-
unsafe; `git grep -l -e nnkteuotnon -e aeelyTo` detects it.
135+
script's own byte-exact read-back passed and `git diff` showed it afterwards, so
136+
a verify-after-write loop cannot detect it. Edit through the editor tooling and
137+
treat any scripted bulk rewrite as unsafe; `git grep -l -e nnkteuotnon -e
138+
aeelyTo` detects it.
139139

140140
## Blocked, not deferred
141141

142142
ShellPilot and `Invoke-ShpBatch` are absent here, so `-Mode Execute` is
143143
unavailable for both eval harnesses. The 75 prepared route-selection prompts
144-
cannot be answered and the authored trigger-query sets cannot be swept, so the
145-
baselined Skills stay unmeasured for discovery. Both need ShellPilot plus a paid
146-
model backend and an explicit go-ahead.
144+
cannot be answered and the authored trigger-query sets cannot be swept. Both
145+
need ShellPilot plus a paid model backend and an explicit go-ahead.
147146

148147
## Open findings
149148

@@ -152,8 +151,8 @@ model backend and an explicit go-ahead.
152151
and in a plain 5.1 child. The untouched `MemoryBankRoleMigration` suite fails
153152
identically, so it is a harness anomaly, not a code defect.
154153
- **Deployment review:** M1-M5 and L1-L6 are implemented and independently
155-
approved on the verified scope. macOS and live cloud-sync execution remain
156-
disclosed, not waived.
154+
approved on the verified scope; macOS and live cloud-sync execution remain
155+
disclosed, not waived.
157156
- **High:** `software-engineer-contoso` claims no egress while retaining an
158157
unrestricted terminal and mandating a generic `security-reviewer` delegate
159158
with web, GitHub, MCP, and terminal tools. Eleven older agents likewise
@@ -162,18 +161,17 @@ model backend and an explicit go-ahead.
162161
on native Windows without terminal sandboxing; replace copied omnibus tool
163162
lists with role-specific least-privilege surfaces.
164163
- **Medium:** Security Reviewer and Technical Writer delegate research to the
165-
full `research-analyst` profile (edit, terminal, browser, GitHub, MCP), and
166-
twelve agents expose `browser` while only Software Engineer carries an
167-
explicit ephemeral-loopback, shared-authentication, and user-confirmation
168-
contract. Both need narrower role-specific surfaces.
164+
full `research-analyst` profile, and twelve agents expose `browser` while
165+
only Software Engineer carries an explicit ephemeral-loopback,
166+
shared-authentication, and user-confirmation contract.
169167
- **Major:** `career-coach` (35,672 chars), `research-analyst` (43,376),
170168
`security-reviewer` (43,772), and `technical-writer` (35,018) exceed GitHub's
171169
30,000-character Custom agent prompt limit. The new test prevents growth; the
172170
separate Session handoff owns the refactor below the limit.
173171
- **Major:** every profile omits `target` but declares a VS Code model-priority
174172
array and mostly VS Code-qualified tool IDs. Copilot CLI documents one model
175-
string plus CLI tool names such as `view`, `edit`, `powershell`, `grep`, and
176-
`task`; product-specific profiles or a shared compatible subset remain open.
173+
string plus CLI tool names; product-specific profiles or a shared compatible
174+
subset remain open.
177175
- **Medium:** no executed agent behavioral eval set exists. The semantic tests
178176
catch structural regressions, but no live capability comparison was run.
179177
- **Low:** three test files parse agent frontmatter with independent regular
@@ -187,14 +185,13 @@ modes covering prompt isolation, label leakage, fallback, strict shape,
187185
reliability aggregation, and failure accounting. The first stage infers routes
188186
and fallback only; the resolver still receives human labels. Context cost,
189187
latency, and answer quality under routed versus full loading remain unmeasured,
190-
and no precision floor is set, so `Passed = True` at low precision is not a
191-
failing build. `WindowsAccessControl` slots 1 and 2 use the older ink-variant
192-
reading of dark/light, so two sets in one shared library disagree on "dark
193-
mode"; `brand-logo-system` integration was measured on one project only; and the
194-
`skill-creator` description edit remains unproven — train reached 100 % while
195-
validation fell, which is the overfitting signal.
188+
and no precision floor is set. `WindowsAccessControl` slots 1 and 2 use the
189+
older ink-variant reading of dark/light; `brand-logo-system` integration was
190+
measured on one project only; and the `skill-creator` description edit remains
191+
unproven — train reached 100 % while validation fell.
196192

197193
## Next step
198194

199-
Leave independent review, publication, and any remote mutation to an explicit
200-
user request.
195+
The seven-task implementation and integration are complete. Keep publication
196+
and any remote mutation behind a separate explicit request. Independent review
197+
of the new local HTTP, persistence, and approval boundaries remains recommended.

0 commit comments

Comments
 (0)