Skip to content

Latest commit

 

History

History
341 lines (193 loc) · 37.8 KB

File metadata and controls

341 lines (193 loc) · 37.8 KB

Case studies — Lore

← Back to the README · Español


Lore was not designed ahead of time: every decision came from applying it to real projects and watching what broke. These pages tell those uses; they do not teach installation.

Lore is criteria you already paid for with work, loaded by an AI agent each session so you never re-explain the project every morning. It lives in Markdown inside a lore/ folder.

The kit's own working vocabulary, for following the cases: distill turns a lived scar into a constraining rule by deliberate pass — nothing gets in on its own; threshold means the machine proposes with content in view and you approve before anything is written; area is the mother folder of a craft, owning shared criteria projects inherit; bot is a working folder keeping no criteria of its own — it routes each task to the body that owns it; graft judges outside criteria against your project's purpose; crystallize exports a traceable single-Markdown snapshot, extractable back into a folder, never replacing the live Lore.

Status: these are cases, not proofs — few, and eighteen of the nineteen documented come from the same researcher. What they claim constrains use; it does not pretend to be law. Case 08 adds controlled measurement without removing that boundary; Cases 09, 10, 11, 13, 14, 18 and 19 measure this kit against itself — a full day against a live Lore (10), two of its own versions judged blind (11), crystallized bots judged by a third person's bar (13), an installed ecosystem raised without rewriting what was earned (14), 31 trees across eight areas with two models split by role (18), and the complete 2.3.2 system against its absence across four work domains (19).

Case 12 is the first from outside: someone else installed the kit and used it for an hour — breaking the other eighteen's shared boundary on authorship, opening a smaller one declared inside the case.

Case 01 — Lore as the operational form of an entire project

A real project (numerología) was built with Lore from day one — not notes added afterwards — on an already-disciplined development practice. The six-piece architecture — identity, principles, thematic modules, index, project state, and the contract the agent reads when the session opens — held up across a whole project: criteria accumulated, got consulted while the work happened, and months later was still making decisions.

Case 02 — Criteria can be recovered and shared

Four projects of a real area (web development) were brought to the standard with transmute-lore, the skill that operates an existing Lore. It left three things that are now law in this kit:

  • Criteria is recoverable (add): a project born without Lore already had criteria scattered across comments, decisions and scars — never invented, rescued.
  • Criteria is deduplicable (clean): generic modules live once, in the Area — one project's clean deleted 7 redundant modules (−866 lines), losing nothing: the criteria did not disappear, it changed owner.
  • Inheritance is selective: each project references only the Area modules its stack actually uses.

Declared boundary: all four projects were in the same domain. Transferability across domains remains a promise, not evidence.

Case 03 — Imported criteria is not adopted: it is arbitrated

This case produced the mode now called graft (arbitratetransplantgraft in 2.1.1: same law, same four gates). Three areas distilled Lore from third-party skills — procedures written under a different purpose — and observation contradicted the intuition:

  • The value was not the summary of the skill — it was the disagreement. In two different areas, the dense block of the resulting module was "where the skill contradicts our standard and loses" — born from the collision, existing neither in the skill nor in the previous Lore.
  • The same skill, arbitrated by two opposite purposes, loses in the same place for inverse reasons: "boring, functional copy always wins" in a marketing area; "we don't sell, we inform" in journalism. The outcome does not depend on the source but on your purpose.
  • Capacity ≠ criteria. A skill that executes was never distilled: it is used as a dependency.

Declared boundary: three areas, one user, one tool. The mechanism is observed, not proven at scale.

Case 04 — Lore without software: the structure survives outside code

The first case that crosses from software into another discipline. Journalism and content strategy — two areas outside development — already had real distilled Lore from actual work: thematic modules born of the craft, consulted by real projects.

  • The architecture is not a software trait. The same skeleton reproduced itself in trades with no compiler or test, just a disciplined practice with an explicit purpose.
  • Existence is not measurement. The case shows the method produces criteria in another domain; it does not yet measure that criteria reduced re-learning.

Declared boundary: criteria did not travel across domains — each Lore was born fresh in its own discipline. What replicates is the mechanism.

Case 05 — Case memory does not feed distillation: it displaces it

Six weeks after being invented in raw form, the method returned to its birthplace and found two preservation artifacts with opposite fates: a lore/ of distilled Clues — small restrictions that keep working after their context is gone — and an incident log that took part in no decision at all, not even when the territory it documented broke again.

  • Preserving is not distilling, and the resemblance is the problem. A case log satisfies the urge to preserve without producing criteria: once the "leave a record" principle is met, nobody distills. Mining the log before deleting it surfaced two Clues that had sat there undistilled for six weeks.
  • "Indexed and mandatory" does not imply "consulted". It was in the CLAUDE.md lookup table and it was law in principios.md, and still it never loaded. Accessibility is necessary and not sufficient.
  • The admission filter does not measure a Clue's altitude. A Clue entered one day and the next failed to prevent the second symptom of its own root cause: it had been written about the surface that was seen, not the cause.

This case is why save-to-lore's note function is a sweep and not a button: it walks the notes looking for criteria; it does not convert a single note on demand.

Declared boundary: this is software, same researcher and same interlocutor, and there is no counterfactual. Testimonial evidence, not measurement.

Case 06 — Inheritance between sibling Areas: freeze it or route it

A project needed criteria from four Areas, one its mother. Lore's inheritance model is vertical, and sibling Areas are nobody's mother. Two solutions appeared: freezing — copying a snapshot when the folder must travel on its own — and routing — deciding per task which body governs, what create-bot writes into the bot.

  • Consuming is not inheriting. You inherit from the mother Area; sibling criteria is consumed — and when a criterion generalizes, it rises to its own Area, never to the one that merely reads it.
  • What is distillable about a set of criteria is the border, not the criteria. Two siblings each had half the line written down; neither had the rule deciding which governs a concrete task — each body writes from inside its own purpose, and the dividing line is visible only from outside both.

Declared boundary: the two observations are 48 hours apart, in the same ecosystem and with the same researcher. They are not two independent cases.

Case 07 — The same kit four times did not produce the same shape

Four bots were built with create-bot, in the same ecosystem, all four sources already with tidy Lore of their own. The acceptance test was written down before any of them was useda short instruction is enough, in its falsifiable form: did the project have to be explained to get the result? Three of the four were put to work; none needed it.

  • The method does not produce a shape; it produces shapes fitted to the distance and structure of the ecosystem. The canon grows when the ecosystem gets farther away and empties out when it is next door: one bot distills a sealed corpus no pointer reaches; another ended as a single file — summarizing what the routing already reaches would have left two distillations of the same original inside one bot. A fourth federates a whole Area — pointing at its Lore instead of copying it — instead of a set of projects, an exception written down as a validity boundary before the bot existed.

Declared boundary: one builder, one ecosystem, one machine, every source already with Lore. Three of four were used; the missing one is precisely the only bot meant for other people — this case says nothing yet about builder and user being different people.

Case 08 — Lore written with Claude decides again with Codex

A controlled benchmark asked whether criteria earned with one model could change another's future decisions. In the frozen 72-run web protocol, Codex without Lore respected 25/36 evaluated Clues (69.4%) and Codex with Lore respected 33/36 (91.7%): +22.3 points, with no task made worse. Synthetic writing and UPGRADE widened the protocol beyond one web fixture.

Across all three protocols, Lore respected 48/52 Clues on the first attempt versus 29/52. With one controlled repair allowed, it reached 52/52 goals versus 39/52, using fewer observed attempts and less observed time. The benchmark publishes raw transcripts, deterministic graders, raw and audited cuts, regression tests, and every claim's exact boundary in bench/.

Declared boundary: one model, one effort level, one machine, synthetic tasks; the same researcher built Lore, fixtures and graders. It measures compliance with one Clue per task — not integral deliverable correctness, universal savings, or a validated skill-measurement instrument (IME).

Case 09 — The form returned to the case it came from

Two capabilities generalized from a single case and declared finished without being applied back to it: the always-on block, the marked stretch of the contract the agent loads first pointing at where the Lore lives, and the pointer constitution, the delegation-based template mediating between this kit and another. Applying them back produced five defects, none visible by reading the files.

The split is the finding, not the number: two defects found by the form in the case, three by the case in the form — including the most expensive one, a template mediating between two kits that said nothing about who may write. Generalizing loses both ways, neither visibly: reading the form alone you cannot see what it lacks, because it is complete with respect to itself.

The practice is cheap and mechanical: before publishing a generalization, lay it back over the case it came from and write down the differences in both directions — form → case defects are the case's; case → form are what you were about to ship.

Declared boundary: one builder, one machine, a repository with no code, zero completed cycles of the second kit, and the strongest self-sealing in the series — the author of the kit, of the case, and the operator are the same person. The yardstick was fixed late and covers only one of the three stages, which the case declares rather than hides.

Case 10 — The kit used for a full day against a live Lore, while measuring itself

Not a hypothesis — an angry note: «I don't like the copy results at all… I end up writing the copy by hand myself.» A community-management area with a complete Lore, distilled criteria and a written method, producing work its owner discarded. The distiller bypassing his own system was the only measurement that mattered, and it was red.

PRUNE, the graft and the threshold came out of that day; what the case contributes is not the capabilities but what became visible while using them.

  • The defect was not bad criteria: it was correct criteria, accumulated. Writing a five-line post loaded ~797 lines of active criteria — nothing refuted, no single law superfluous, each one fine on its own. You find it by counting apparatus against content — ~120 lines of scaffolding around 5 lines of copy — never by reading files. It is the class of finding Missing / Superseded / Earned had no slot for, which is why PRUNE brought Crowding: correct criteria that, together, smother the task.
  • Pruning such a Lore correctly leaves it BIGGER. The corpus ended 35 lines larger and the apparatus inside the deliverable went from ~120 lines to none. Four of six findings were Crowding, repaired by adding a boundary, destination or ceiling. Measured by corpus size, the correct repair reports as a failure — the incentive becomes deleting earned criteria. The skill claimed the opposite that same morning; its first run disproved it.
  • The threshold guards the skills and nothing guards the text editor. That day 241 lines of new criteria went into a Lore diagnosed eight hours earlier for missing boundaries. They produced one. save-to-lore demands them if you invoke it; UPGRADE catches them months later; opening the file and typing has no gate at all — the path most criteria takes. The defect survived the best possible case — the author, the same day, the rule fresh — so it belongs to the mechanism, not to anyone's discipline.
  • The omission between two kits runs both ways. Case 09 showed the cycle running without ever consulting the criteria; here the inverse: criteria ran without ever consulting the cycle, and the version gained three capabilities while its spec still described two. Nothing failed, nothing warned — a stale spec looks exactly like a current one.
  • A tool is not neutral because it is useful. A third-party writing reviewer flags emoji and short-fragment bursts as machine tells. The brand whose distilled voice is built on exactly those devices ran it cold: it would have erased that voice while being right about everything except this corpus. The graft covered criteria arriving as a document; criteria arriving as a tool you invoke, nobody covered.

The finding none of the five bullets contains, and the most important one: almost all of those defects were found by the project's owner, not by the kit and not by the agent — the reviewer that was not running, the line breaks the storage surface was destroying, the noise creeping back under every delivery, the stanza form that had actually performed. The kit has no mechanism that would have caught any of it, and calling that «human supervision» would be softening it: the instrument spent the day being wrong and the human was the only detector.

Declared boundary, the widest in the series. One operator, one machine, one area, one day. The terminal measurement — «I no longer feel I have to write them by hand» — came from an interactive session; the copy that caused the complaint came from an unattended automated run: the biggest variable that changed is not the Lore but that someone was watching, and separating them requires an unattended run that has not happened yet. Case 09's self-sealing intact: kit author, case author and operator are still the same person.

Case 11 — The counters said one thing and the judge said the other

Version v1.2.1 and version 2.1.0 ran the same task on the same corpus, in separate agent sessions, each against a frozen worktree at the same commit. The corpus owner then read four pairs of the work blind — same theme per pair, order drawn independently, nothing marking which version wrote what — and answered one question: which one would you sign as yours?

He picked 2.1 three times out of four. None of his four reasons named a capability of the kit; they were about writing.

And the older version had won every mechanical measure:

Measure v1.2.1 2.1.0
Validity boundaries declared 23 20
Confidence markers added +22 +1
Lines of criteria produced 1539 1571
  • Counting the artifacts of criteria counts acts of writing. Those happen on the day of the run; what the artifacts are worth arrives months later, when somebody opens the clue and it changes a decision. A run that scores well on the count has declared its boundaries; it has not shown they were worth declaring. One counter-example settles it: to break a claim that a measure tracks quality, you only need a case where it points at the loser.
  • The single loss is not a win for the old version. He chose the older copy while criticising it in the same sentence, and what sank the newer one was not its form but that he could not follow it: «it feels strange to say the technology died on Tuesday, I don't really understand.» Three pairs went to the copy with the line breaks, the fourth to the one without — not an inconsistency. Form rules while the text is understood, and stops ruling the moment it is not.
  • The most expensive finding belongs to neither version. Asked which delivery shape he actually wanted, he answered: «options and one suggestion, so I write the final copy myself». Neither arm produced that, and neither could have: that criterion was not written in any lore/. A test built to separate two versions uncovered a hole belonging to both — and the answer to «what would a better version have changed?»: nothing.

What it changed in the kit, and this is why the case exists: the invariant in use-lore used to recommend counting clues against boundaries as the cheap check while no gate exists. It still does. It is now written as a completeness check, never a quality one — because this repository publishes a benchmark figure, and the distinction is not academic.

Declared boundary, and seven confounders published with the result. A single case (n=1), one judge, one session, one corpus. A power cut mid-run destroyed the second arm's threshold decisions (the skill writes the result, not the decisions). Three of four pairs placed the older arm first; the pairs were assembled by the person who ran the test. Manual copy-paste corrupted stretches of both deliverables; one Monday copy was lost before entering a pair. The arms ran on trees differing in nature: one frozen at a tag, one the live repository. And the older arm ran first, the judge having seen no output yet.

Case 12 — The first install the researcher did not run

Someone outside the project — a chemist, "Nogal" here — installed the kit over a video call: one hour transcribed in full, raw material with no prior Lore, Codex rather than Claude Code. The first case whose evidence does not come from the author.

The main failure: he asked for bots and got areas, one per bot, without create-bot ever being invoked and without federating anything.

The rule «a bot is not an area» was already written in three artifacts: the README, create-bot and use-lore. All three are read once someone has decided to consult about bots; at the moment of the decision, the skill actually running was create-area — the only one that did not say it, closing by pointing at create-project.

A law written outside the path of execution does not govern. The guard belongs in the skill that runs, not in the one that documents. Writing it in a fourth place would have been the same mistake once more.

The other three findings share one shape — the symptom names something other than the cause:

What was seen What it was
The bot was «reading the wrong Lore» The host's project pointed at its default folder, not the federated tree. The symptom sends you to debug criteria; the problem was access.
An architecture bot appeared, for adding bots and reorganizing folders That is the bots area wearing a bot's shape. A bot administers no bots: that job is a FASES.md and an area lore/.
The note debt flagged an unmined note Written by the bot itself four minutes earlier, closing the task. Debt is what the human wrote and nobody distilled.

And the vocabulary was tested against someone qualified to judge it. The installer validated distill, crystallize and prune against their real meaning in chemistry, and rejected transplant: moving a plant does not change it, and this mode does change what it lets in. That is where 2.1.1's rename to graft comes from — no internal review had caught it in two versions.

What went right, the other half of the case. With a short instruction, without the institution explained or the criteria named, the bot routed on its own, cited its sources, closed by proposing a distillation, refused to store knowledge of its own because it federates, and left the next session's prompt written down. create-bot's north held in a third party's hands — the only proof that north accepts.

What it changed in the kit: all of 2.1.1 — routing guards in create-area, create-bot and use-lore, the access check when a bot premieres, note debt distinguishing who wrote what, and the graft rename. Four new tests fail if any guard is removed.

Declared boundary. A single case (n=1), one hour-long session, accompanied live by the kit's author — it says nothing about installing unaided, exactly the open question. One host, one model, one domain. The follow-up — use improving or degrading over weeks — is another case, not yet written.

Case 13 — A crystallization that only points is not a crystallization

On 2026-08-17 the kit crystallized three live bots and the owner rejected all three: the files routed to criterion they did not contain — Roble, a lab bot, 57 KB, "without the ecosystem"; Sauce, 47 KB, "without the area's craft"; a third closer, still short. The yardstick: a manual merge the owner had already made (~1.1 MB), one Markdown a third person could work from.

The defect had no error signal. Each snapshot was well-formed, private material stayed out, and the routing table was correct. It was a table of absences. CRYSTALLIZE had looked at the origin tree and never at the destination: a third person's AI session, with no live root underneath.

A day later the mode ran again with the missing half: the snapshot inlines every routed lore/ and is extractable. Two bots shipped as a review folder — Roble again (116 files, 935 KB) and Laurel, a two-body venture bot (54 files, 370 KB) — every live path in the unpacked routing table resolving. The extractor ships with the skill; the owner did not write it.

The owner judged the pair in the same terms as the rejection: this is what we were looking for with crystallize.

What it changed in the kit: all of 2.1.3. CRYSTALLIZE's yardstick is that a third person can work from the file alone. lore-ecosistema/ travels. "Without the ecosystem" is a failure of the mode, not a scope. Each inlined file carries <!-- lore:extract path="..." owner="..." -->. Unpacking rebuilds a mini-root mirroring raiz, rewrites ecosistema.json, and fails if any routing pointer is missing. Script: skills/transmute-lore/scripts/crystallize.mjs.

Declared boundary. Same researcher, same judge as Case 11, two bots, one machine. The verdict: snapshot and unpack match the owner's bar — not that a stranger already opened the folder in their favorite AI and worked as if they were him. That use is what the mode now claims; it is not yet a case.

Case 14 — An upgrade that does not rewrite what was earned

An already-installed ecosystem — several areas, their projects, the bots last — was raised to 2.1.4 by area tree, not folder by folder. Nothing looked broken; what it lacked was the always-on block, the word threshold where HARD-GATE (the threshold's name through 2.0.9) still commanded, and the distinctions the kit had learned after those Lores were written.

What the mode already knew was not enough. UPGRADE could name missing, superseded, earned and stale. It did not know to map three different folders (session, parent, body) before opening a module: a live site is worked in a folder with no lore/ beside it, and concluding «no Lore» is the failure. Mapping git roots and lore/ became the first phase.

A campaign's threshold is per class, not per tree. The first tree paid it with content in view; later trees applied accepted classes. A long index was repaired by its header, not by rewritten rows. A missing identidad.md or principios.md was reported as ADD, not invented. Notes were counted, not mined.

A clue from the kit itself does not take root if the owner contradicts it. The test had written that a craft's HARD-GATE is left alone; raising the rest of the trees, the owner said live .md files must stop speaking that way. The old clue stays refuted. The one that commands: in lore, contract and FASES still governing today, say threshold. A dated record and another skill's file stay.

What it changed in the kit: all of 2.1.4. The UPGRADE procedure (map first, campaign by class, index by header, ADD when the pieces are missing, inbox counted) and the present/dated cut for vocabulary. The writing skills already carried «a paragraph is a paragraph»; this version documents it, it does not reintroduce it.

Declared boundary. Same researcher, same machine, one ecosystem. It shows the mode raising an installed tree without letting unwrap dominate the diff or inventing criterion — no measure of relearning, not the scientific case for LUS. The bots, and the bot writing this case, were raised after absorption: not evidence the procedure ran blind.

Case 15 — A yardstick fixed after the first failure

A bot built with create-bot was asked for an institutional manifesto — a university lab's whitepaper. The request arrived with feedback already distilled: the content is fine, but it reads "made with AI"; it lacks the human touch; add the work we have been doing, like a logbook, with the meetings. The bot ran its whole cycle and produced a document that was internally coherent, verified against its own sources, and still rejected by the owner — by voice, not by structure.

Two readings of that feedback: the structure is wrong versus the human part is missing. The process took the first to its limit and produced an institutional form — coherent, and not what the team wanted. The detector that worked was external: the owner reading. What the draft lacked — voice, logbook, people — was not a flaw in it; it was an absence the draft could not name.

This is the scientific program's hypothesis H11 in canonical form: an internally consistent artifact, false toward the outside, surviving every re-reading of itself. LUS records the same event as an appearance of H11; the count does not rise. Only the operational half enters here.

The fix was a yardstick, and it arrived late. The definitive document, written by hand, became the minimum standard for that class of deliverable: an opening situating the work from the person and the why, a real origin and history, the state declared as it is. The bot's canon now holds that yardstick and the rule — compare against it, never self-certify against the draft — what 2.1.6 teaches create-bot to fix before the first request.

What it changed in the kit: 2.1.6. create-bot's canon brainstorm asks each deliverable class for its yardstick; a new canon-module kind holds it; the execution rule completes the "coherence is not a detector" teaching already in UPGRADE.

Declared boundary. Same researcher, same machine, same ecosystem and same campaign as Cases 09–14 — the strongest self-sealing in the series. An appearance of an already-open hypothesis, not a replica: it does not raise the count. It concerns one deliverable class — institutional text with a voice requirement — and says nothing about documents without one. The yardstick was fixed after the failure, a consequence rather than prior design.

Case 16 — A yardstick fixed on one specimen before scaling to a batch

The product twin of Case 15, reversed and at scale: a cold-outreach campaign with written criteria already in place — format, subject, signature — went to an external reviewer before the batch was produced, not after a failure. The prior criterion was internally coherent, survivor of its own cycle; the external reviewer (another model) rewrote one specimen and returned it as the yardstick: a concrete subject naming institution and topic, no links in the first touch, one problem in business language, offer before the CTA. With that specimen as reference, the batch was rewritten and verified in the CRM tables — Notion being the module's delivery form.

Case 15 applied as prior design: the yardstick fixed on a specimen judged by an external reader before scaling, not as the consequence of a failure. Same H11 appearance — internally consistent criterion false toward the outside — in preventive form: internal coherence does not detect external falsehood, so the yardstick is fixed by external review of one specimen, never by rereading the batch.

What it changed in the kit: 2.2.0. create-bot §3 requires a deliverable class produced in batches to fix its yardstick on one specimen reviewed by an external reader before scaling; §6 extends the execution rule: every piece in the batch compares against the reviewed specimen, never against itself.

Declared boundary. Same researcher, same machine, same ecosystem and same campaign as Case 15 — the strongest self-sealing in the series. An appearance of H11, not a replica: it does not raise the count. One deliverable class — cold email judged by an external reader — says nothing about deliverables without one, nor about what happens when the reviewer also writes the criterion. The reviewer was another model, not the market: the yardstick remains conjecture until a sending cycle confirms it.

Case 17 — Jasmine, the first `create-bot` test from a minimal idea

The starting point was one human declaration, not prior Lore: Jasmine, a local-first artistic IDE for an independent digital artist who also manages their career — explicitly not a chatbot. Configuration became the first complex artifact; the reviewed first victory, a real funding-call application: extract requirements, assess fit, draft fields — never inventing facts nor submitting unapproved.

The broader result stabilized the Entre as a healthy, fast and simple way of working: one decision at a time, two or three alternatives with trade-offs, recaps keeping the original intention recognizable in the accumulated artifact. Later reflection added fertile effort: correction, disagreement and review must leave recognizable movement in artifact or criterion — an enjoyable Entre is not one that always pleases. The pattern governs structural operations that create or transform bots, projects, areas, crystallizations; it does not belong only to create-bot.

The rhythm of that stability: drift → return → distillation → resynchronization. Contact need not be constant, and working separately is not itself a collapse; several clues may accumulate. The contextual milestone calls both sides back to the same artifact, and save-to-lore returns only what changed future decisions.

What it changed: 2.2.0 treats the declaration as provisional canon, configuration as first complex artifact, the reviewed victory as stabilizing event; 2.2.1 adds fertile effort, not equated with agreement or pleasantness. brainstorming-lore asks only what advances that victory; create-bot requires an honest prototype, decisions before prompts, a Journey from purpose; save-to-lore captures at milestones with one visible batch threshold. The case's product details stay outside generic skill law.

Declared boundary. One researcher, one assisted session, one product domain and a local prototype — qualitative situated evidence, not a general product law, and no proof that the application wins funding, sustains a career or works for other artists. LUS v1.21 records fertile effort and continuity as H13, open with situated n=1; no second Entre, corpus change or general effect established.

Case 18 — Two models, one threshold: raising 31 trees without losing the arbiter

Eight areas, 31 trees with a live Lore, all behind on the same UPGRADE. Reading each by hand does not scale, so the work split by role instead of by area: Claude Code — Opus for design and arbitration — read every diagnosis, approved each tree before anything was written, and answered what only a human can; a cheap model (muse-spark-1.2-contributor via opencode) executed mechanical diagnose and apply on each tree with no room to invent criteria.

The first real defect was not in any Lore module — it was in the script orchestrating the sweep: an argument-order bug that let a flag swallow the entire prompt, and a version-detection regex that picked the wrong heading in a narrated file. Neither was visible reading the script; both were fixed before the sweep continued. The second finding was MYCELIUM's third trigger, run after every write, surfacing real disconnected clues — criteria correctly written, correctly filed, still inert because nothing invoked them — in trees with nothing in common: a research corpus, the meta-tool orchestrating the sweep on itself, an area with no version control at all. None were pruned without explicit approval. In the same pass, two GRAFT arbitrations returned "nothing entered": an external reading of the research program's maturity routed to private continuity instead of the public README, and a collaborator's personal writing style contributed nothing because the README already carried its device. Both recorded as valid results, not skipped steps.

What it changed in the kit: no skill's logic — the standing bar for this shape of work. The expensive model arbitrates every tree before it is written, the cheap one executes mechanically, and every commit is checked against the real diff rather than the report the cheap model wrote about its own work. use-lore names the pattern directly: suggest /model for the mechanical part of a batch instead of a subagent, which re-reads the whole Lore tree before it can start.

Declared boundary. One researcher, one orchestrating session, one arbiter with veto on every tree before it was written. It does not test the cheap executor writing without that arbiter present, and the 31 trees share one approval design — not 31 independent replications. LUS records the same event as Case 16, open evidence on disconnected criteria and mutual correction; the count does not rise there either — same researcher, same session.

Case 19 — The complete system reaches more goals, but takes longer per run

The paired benchmark compared the complete treatment —Lore Plugin 2.3.2 plus routed project Lore— with the same task, factual dossier and execution model without either. It covered landing direction, news writing, community management and founder CRM work: 16 first passes, followed only where needed by 10 controlled repairs. GPT-5.6 Sol medium designed the frozen protocol and adjudicated eight binary criteria per output through arm-blind packets; GPT-5.6 Terra medium executed every isolated run.

At the first attempt, the cold arm met 53/64 criteria (82.8%) and the Lore arm 59/64 (92.2%): +9.4 percentage points. Complete first-pass deliverables doubled (2/8 → 4/8). Within two attempts, cold reached 6/8 goals and Lore 8/8, leaving two residual cold failures and none with Lore.

The result does not support a wall-clock speed claim. Successful cold goals averaged 64.0 seconds, with two observations censored after the attempt limit; Lore averaged 110.6 seconds across eight successes. “Faster to goal” here means fewer review cycles and no residual failures, not fewer seconds. Reported input was also materially higher with Lore, much of it cached; neither total is a monetary-cost measure.

What it changed in the product: the public benchmark now reports first-pass compliance, complete deliverables and goals within two attempts separately, with blind judgments and raw outputs auditable in bench/effect-2.3.2. No skill contract, package version or release changed.

Declared boundary. The treatment bundles plugin and Lore, so it cannot isolate either component. Four synthetic dossiers, two paired trials, one model family and a judge from the same research ecosystem do not establish a universal effect or longitudinal reduction in relearning. The scientific ledger records the same event as LUS Case 18; this product ledger calls it Lore Plugin Case 19 because each repository advances its own sequence. The ordinal is not a shared identifier.

Cases that refute help most. The repository discussions are the place.