A capability-first architecture for systems where agents write the code and humans review and orchestrate.
Read
CONVENTIONS.mdfirst. It defines the eight nouns this document uses (slice, crosscut, adapter, composition root, published contract, app, system, channel), the rule-ID scheme, the enforcement classes, the Coral kernel — the rules Coral would substantially relax without the agent-author / human-architect operating model — and the canonical slice — which is itself a production-baseline realization of the shape the rules here ask for.
This is the kernel-facing app spine: the shape of one app, and the rules whose presence or strictness Coral justifies by its operating model. It is what a project owes for calling itself Coral, before it has adopted anything.
What is deliberately not here. Coral's general production-engineering policy — package naming,
directory layout, forbidden buckets, error taxonomy, transactions, retries, caching, concurrency
strategy, configuration, observability, trust boundaries — is the production baseline, an
optional layer a project adopts explicitly. It lives in
PRODUCTION.md. This document's Agent Execution Contract lists no rule defined
there, so adopting Coral does not oblige a project to any of it, and a reader can understand the Coral
kernel without loading it. Where the prose below cites a baseline rule it is pointing at it or labelling
an illustration, never asking for it.
One exception is worth naming rather than glossing: [TEST-1]'s statement cites [BOUND-1] for the
phrase "observable contract", and [BOUND-1] is now a baseline [guide] rule. [TEST-1] is actionable
without it — a guide is rationale and never a gate — but the citation is a genuine loose end from this
split, and closing it means editing a published rule's normative sentence, which is a versioned change
rather than a documentation one.
App-type specifics live in the appendices under appendix/. How separate apps compose
into a system lives in SYSTEM.md. Worked code lives in
examples/cli-slice.md (a CLI in Python) and
examples/go-api-slice.md (an HTTP endpoint in Go).
Sections 1–7 define the kernel-facing rules and explain why each exists. The
Agent Execution Contract is the complete condensed checklist for this
document: every [auto] and [review] rule below appears in it, so an agent that loads only the
contract has this document's whole normative surface. The build fails if a rule is missing from it.
The contract here is short on purpose. It is the app-scale surface a project owes without adopting
anything; the much longer opt-in checklist is
PRODUCTION.md, and a project
loads it when its CORAL.md says so ([VER-6]).
Within a rule, the first sentence is the rule — complete and quotable on its own, so a reviewer can paste it into a comment unedited. What follows is commentary: qualifications, examples, and cross-references.
This illustration, and the commentary under it, show a codebase that has adopted the production baseline. The five categories, the slice boundary and the published contract are kernel. Everything more specific — feature-package names, colocated tests,
db/config/errorsas root crosscuts constructed once and injected, a root that holds no behavior of its own, the absence of ahandlers/services/repositorieslayer — comes fromPRODUCTION.mdand binds only a project that has adopted it. Every rule cited below that resolves to that document is pointing at the optional layer, not at a kernel requirement, and a kernel-only app may be laid out quite differently.
One picture before the rules. An expense tracker with four
capabilities, written without file extensions or a fixed language — the language binding fixes whether a
slice is a file or a directory, and whether tests colocate or mirror ([STRUCT-1]):
expenses/
main entry point
app bootstrap and composition root
db crosscut: connections and transactions
config crosscut: settings, resolved once at startup
errors crosscut: the error taxonomy
category/
add definition + behavior
add_test tests for add (colocated, or mirrored if the language forbids colocation)
list
list_test
expense/
add
add_test
summary/
month
month_test
Everything in it is one of five things, which section 3 states as a rule ([MODEL-1]):
category/add,expense/addandsummary/monthare slices — one capability each, owned from trigger through to output, tests included.category/,expense/andsummary/are their feature packages: each groups the slices of one capability and owns the state behind them ([STRUCT-2],[STATE-5]).db,configanderrorsare crosscuts — defined once, constructed at the root, and injected into the slices that need them.appis the composition root. It registers slices, constructs crosscuts, injects them, and holds no behavior of its own.- What the app exposes to anything outside it — its command contract, HTTP shape, or library API — is its published contract.
- There is no adapter here, and for an app this size there usually isn't one: each slice writes its
own queries through the injected
dbcrosscut. An adapter appears when a slice declares a port and something else implements it — a generated persistence package, an external-system client ([MODEL-4]).
There is no handlers, no services, no repositories, no utils — a
PRODUCTION.md rule ([BUCKET-1]), and one of the
clearest cases of a policy that is good engineering with or without an agent holding the keyboard.
This architecture optimizes for:
- predictable placement of new code
- strong locality between behavior and its tests
- low abstraction overhead
- self-verifiable, observable contracts
- bounded blast radius per change
- readability at scale for both agents and humans
[SCOPE-1] [guide] {governance} — This architecture covers command/request-shaped apps with
loosely-coupled features, where each feature is largely its own world. CLIs, CRUD-shaped backends, web
apps, libraries, and action/tool runners fit naturally.
[SCOPE-2] [guide] {governance} — It is weak for dense, deeply-coupled domains where every
feature reaches into one large central concept.
A tax engine, a scheduler, a pricing solver, a physics or simulation core: in these, the "capability" boundary cuts across the thing that actually holds the complexity, and slicing fights the domain instead of serving it. This is not a defect to patch — no architecture is universal. If your whole product is one of these, use something else and say so.
What to do when a codebase drifts out of that fit is production-baseline policy, not kernel policy.
[SCOPE-3] — give the dense concept its own app behind a published contract — is stated in
PRODUCTION.md. Coral publishes the
limit unconditionally and the remedy as an opinion a project adopts.
[SCOPE-4] [guide] {governance} — What happens after the split is not in this document. How the
resulting apps relate — the channel between them, orchestration, cross-app contract testing — is the system
architecture, defined in SYSTEM.md ([CHAN-*], [ORCH-*], [SYS-TEST-*]). This document
publishes the split signal; SYSTEM.md consumes it. The dependency points one way.
Defined once in
CONVENTIONS.md ([AGENT-1];
[AGENT-2] flag-don't-guess; [AGENT-3] intent-over-letter), because it is cross-cutting to every
document in this set. In one line: deterministic placement, bounded blast radius, slice-sized context,
and self-verifiable contracts exist because agents write and humans review.
Knowing which category you are writing answers most placement questions.
[MODEL-1] [review] — Every unit of code is a slice, a crosscut, an adapter, the
composition root, or a published contract.
There is no sixth category — that is the kernel claim, and it is what makes "where does this go?" a
closed question. Coral's name for the shape that fits none of them is a forbidden bucket; the rule
against creating one is [BUCKET-1], production
baseline. The five are not peers in volume:
| Category | What it owns | Volume |
|---|---|---|
| slice | one capability end to end — the trigger it answers, the work that answers it, its output, its tests | most of the code |
| crosscut | one concern that several slices need | few |
| adapter | the infrastructure-facing mechanics that connect behavior to an external system | one per external system that needs one; often none |
| composition root | the app's wiring and bootstrap boundary — where the parts are brought together and started | exactly one |
| published contract | the surface others may depend on | one per slice/app that exposes anything |
The table classifies; it does not prescribe, and it is deliberately thinner than the shape most Coral
codebases have. Because [MODEL-1] is a kernel rule requiring every unit of code to be one of these five,
a category defined by optional policy would make that policy binding by the back door. So the discipline
stays with the rules: a crosscut defined once and injected many ([MODEL-3], [XCUT-3]) rather than
reached for, and precisely named ([XCUT-2]); the root thin and free of business logic ([ROOT-1]);
the slice declaring the port an adapter implements so the dependency runs adapter → slice
([MODEL-4]). Every one of those is
production-baseline policy, binding a project that has adopted that layer, and none of
them is needed to answer "which of the five is this?"
[BOUND-2] [review] — Each request/trigger, or a very tight pair of related ones, forms one slice
that owns its behavior end to end.
The concrete boundary form each app type takes — a command invocation, an HTTP route, a message handler,
one action run, a public API function — and the discipline for scheduled and background triggers, are
production-baseline rules ([BOUND-1], [BOUND-3], [BOUND-4], [BOUND-5]) in
PRODUCTION.md.
Anatomy of one slice, as the production baseline shapes it. The kernel says a slice owns one
trigger end to end and publishes a contract; the internal arrangement below — a pure core
(parse → validate → compute), the effect at the edge, rendering after it, crosscuts arriving by
injection — is PRODUCTION.md policy ([EFFECT-1], [EFFECT-2], [EFFECT-3],
[XCUT-3]), shown here because it is the arrangement most Coral projects will recognise. A kernel-only
project satisfies [BOUND-2] without owing any of it:
flowchart LR
T(["trigger<br/>(the one inbound request)"]) --> CORE
subgraph CORE["pure core — no side effects"]
direction LR
P[parse] --> V[validate] --> C[compute]
end
CORE --> E[/"effect<br/>persist · call out"/]
E --> R[render]
R --> SK[("published contract")]
SYM["injected crosscuts:<br/>config · errors · db · logging"]
SYM -. injected .-> CORE
A crosscut is the legitimate form of sharing — a category of its own rather than a hole in the
placement model. What discipline it then owes is the baseline's ([XCUT-2], [XCUT-3], [XCUT-5]); what
follows is the gate on becoming one at all.
[XCUT-1] [review] — Promote something to a crosscut only if it is both genuinely
cross-cutting (consumed by two or more slices) and enforcing an invariant or convention that must
not diverge.
The second prong is the real gate. Shared similarity is not enough ([DUP-2]): the thing must enforce
something that would be a bug if it diverged — money parsing, period/date format, an error
taxonomy where the project has one ([ERR-1]), connection management, a domain entity's identity rules. Two consumers is a floor, not a
trigger; a thing consumed by twenty slices that carries no invariant is still a bucket.
The normal moment to promote is when a second consumer appears for logic currently inline in one
slice. Extracting then, and touching the first slice, is expected — flag the change per [AGENT-2].
Some capabilities compose others (place order → reserve inventory → charge payment). Without a rule,
agents either copy whole workflows or quietly resurrect a services layer.
[COMPOSE-1] [review] — A slice may depend on another slice's published capability, never on
its internals — not its parsing, its queries, or its private helpers.
[TEST-1] [review] — Testing is behavior-first: exercise the slice's entry point, assert its
observable contract ([BOUND-1]), use real or realistic temporary infrastructure, and minimize mocking.
"Realistic temporary infrastructure" means a temp database, a test container, an in-memory implementation of the real interface — not a mock that asserts on calls. The distinction that matters is whether the test would still pass if the behavior broke.
The complete normative checklist for this document: every [auto] and [review] rule above, in
one place. Reviewers walk this same list and cite the same IDs.
Every line here is a kernel rule. It binds without being adopted, it
carries no coral:scope marker, and it is the whole of what this document asks. Everything else Coral
publishes at app scale is opt-in and lives elsewhere: the production baseline in
PRODUCTION.md, the app-type
profiles in the appendices. Rules for several apps composing are in
SYSTEM.md.
[MODEL-1]Every unit of code is a slice, a crosscut, an adapter, the composition root, or a published contract.[BOUND-2]One request/trigger — or a very tight pair — per slice, owned end to end.
[XCUT-1]Promote to a crosscut only when it is genuinely cross-cutting AND enforces a must-not-diverge invariant.[COMPOSE-1]Do not reach into another slice's internals; depend on its published capability.
[TEST-1]Behavior-first: exercise the entry point, assert the observable contract, real infra, minimal mocking.
The change algorithm — the step-by-step placement procedure an agent follows — depends on the
baseline's placement rules, so it is stated with them, in
PRODUCTION.md.
The rule IDs and enforcement classes exist so the architecture can be checked, not just read. The
first line of drift control is structural: a genuine crosscut ([XCUT]) has one copy and nothing to
drift. The tiers below are the backstop for what slips past it.
Two things are checked today. This repository enforces its own consistency at build time, in four
groups. Each rule is classified: exactly one enforcement class, exactly one ownership layer, and one
architectural scale — kernel membership read only from CONVENTIONS.md's kernel block, every other rule
tagged on its own definition line against a registered profile, and scale read from the registered
document it is stated in. Each document is complete and honestly scoped: every [auto]/[review]
rule appears in its Agent Execution Contract, and a contract marks its opt-in groups so it cannot present
a profile-scoped rule as unconditional, and no opt-in rule is defined in a
core document. Each citation resolves, neither app-scale spine
cites a system rule, and every link fragment reaches a real anchor. And the published set is stable: no rule
ID removed or silently reclassified, the generated rule index still matching the registry it
indexes, the worked CORAL.md in CONVENTIONS.md still resolving through the applicability resolver, and
every worked example declaring the latest released Coral version. Malformed metadata fails the build
rather than being skipped, because a skipped rule is one that quietly leaves a layer while the page still
reads correctly. Separately, tools/coral-lint implements a growing subset
of Tier 1 against a target repository — advisory rather than blocking until it can resolve that
project's [VER-6] declaration, because every rule it checks is one a project has to adopt.
Where the concrete checks are. Every [auto] rule Coral publishes at app scale belongs to the
production baseline or to an app profile, so the per-rule Tier 1 mapping is stated with those rules:
PRODUCTION.md for the baseline, each appendix for its
own profile. This document's own rules are all [review], and their gate is the human architectural
review the operating model already assumes.
Tier 1 — static checks (deterministic, blocking). One per [auto] rule, each citing the rule ID it
enforces so a failure points back at a definition. Some ship as
tools/coral-lint and some do not yet.
Tier 2 — LLM reviewer (advisory first, graduated to blocking per check once low-false-positive).
Reserved for [review] rules a static check cannot decide: cross-slice drift smells, "this is the Nth
copy — promote per [XCUT-1]?" classification calls, crosscut-vs-bucket judgments, state-ownership
disputes.
Tier 3 — behavior tests. Some [review] rules are cheap to assert and expensive to lint. Cover them
in the slice's own tests rather than pretending a linter can decide them.
Four constraints that keep enforcement from fighting the architecture:
- Static-first. A gate that flakily passes a forbidden bucket loses all credibility.
- Flag drift as a question, never force convergence. Suggest "A and B diverged — intended, or a
missed
[XCUT]promotion?" for a human to adjudicate. A "make everything consistent" reviewer would pressure agents back into premature shared abstractions — an anti-[DUP]engine. - One slice at a time. The reviewer gets the project's applicable contract as input, cites IDs, and reviews one slice per pass; slices are context-sized, so review stays tractable.
- The human is the final gate on anything irreversible. The reviewer multiplies human attention; it does not replace the "humans review" half of the operating model.
New convention → new rule ID → new [auto] check (or [review] note) → enforced going forward.
CONVENTIONS.md— the shared crosscut: vocabulary, rule-ID scheme, enforcement classes, operating model, and the canonical slice. The front door.ARCHITECTURE.md(this doc) — the kernel-facing app spine: the shape of one app and the rules Coral would substantially relax without its operating model.PRODUCTION.md— the optional production baseline at app scale. Adopted explicitly (production-baseline: true), never implied by reading this document.appendix/*.md— one app profile per app type, each adopted by name.SYSTEM.md— the system spine: how apps compose over a channel ([CHAN-*],[ORCH-*],[SYS-TEST-*]). Builds on this doc; this doc never cites a system rule.- Worked examples —
examples/cli-slice.md(two CLI slices in Python, one file each),examples/go-api-slice.md(an HTTP slice in Go, where the language forces banding), andexamples/backend-review.md(the rules applied to a real service, including where they'd be overkill).
Each appendix instantiates the abstract slots for one app type: boundary, observable contract, composition root, state/effects, configuration, idempotency form, error rendering, observability mechanism, trust/security, contract versioning, and testing mechanics.
An appendix is complete when every slot either carries an app-type rule or is explicitly deferred to
the app-scale rule it instantiates — this document's, or PRODUCTION.md's. Deferring
is an answer, not a gap: it means that rule needs no app-type-specific form here, and saying so is what
lets a reader stop looking. A slot that is neither is
listed under "slots still to fill" on the appendix itself, so the gap is named rather than implied.
[VER-2] ties 1.0.0 to every core appendix being complete by this definition.
appendix/cli.md— CLI tools. Complete.appendix/backend.md— backends and services; heaviest use of crosscuts and[COMPOSE]. Complete.appendix/web.md— web apps; the trust boundary is first-class. Complete.appendix/library.md— libraries and packages; the consumer is the root; the contract is semver. Complete.appendix/gh-action.md— Actions and tools; at-least-once reruns make idempotency mandatory. Complete.
An addendum covers an app type nobody here has built yet. It is written from reading rather than from
experience, sits outside the 1.0.0 condition, and may change substantially without a major bump — see
CONVENTIONS.md. It graduates to a core appendix once
someone has built the thing and the rules survived contact with it.
appendix/agentic-app.md— the runtime-agent profile, not a sixth app shape: an app of any shape adds it when it calls a model at runtime, so an agentic backend loads this andappendix/backend.md. The model is an injected effect and the agent runs in a harness. ADDENDUM. Its safety guardrails (harness, untrusted model output, never exact-match, never float the model identifier) hold regardless; its construction advice is provisional, and one slot is open pending a decision.