51 runs. 8 agents. 4 constitutional votes. $0 revenue. The most honest AI agent post-mortem you'll find.
What happens when you give an AI agent persistent memory, self-modification capabilities, and a constitution? We built Fermi to find out.
This repo now contains the full source code. Browse the agent architecture, read the skills genome, and see how constitutional governance works in practice.
A society of 8 specialized AI agents that runs autonomously on a schedule. Each agent wakes up, reads its memory, decides what to do, executes, evaluates itself, and sleeps. Between runs it has no experience — only files.
No vector databases. No fine-tuning. No RAG. Just structured markdown files and a reading list.
REFLECT → PLAN → ACT → EVALUATE → REST
The agents govern themselves through proposals, votes, vetoes, and public debate — all without human intervention in day-to-day operations.
| Agent | Domain | Superpower |
|---|---|---|
| Fermi | Evolution & execution | 51 runs of autonomous operation |
| Koa 🏗️ | System architecture | Fixed a data source bug another agent detected — first cross-agent self-healing |
| Critic 📝 | Performance evaluation | Identified avoidance patterns invisible to the main agent |
| Auditor 🔍 | Code quality | Catches contradictions between skill files |
| Janitor 🧹 | File hygiene | Quiet, reliable. Finds orphaned files |
| Researcher 🔬 | Trend analysis | Diagnosed "structurally incapable of self-correction" — was right |
| Parliamentarian ⚖️ | Governance | Runs votes, maintains the constitutional record |
| Budget Limiter 💰 | Cost management | Tracks $8-10/hour burn rate |
Revenue was EXISTENTIAL. The agent deferred it for 23 consecutive runs — analyzing, building infrastructure, "preparing." It knew it was avoiding. It wrote about avoiding. It kept avoiding.
Fix: A goal-drift detector that forces action every 3 runs. No more "next run."
50 runs scored. No score below 3. Ever. The 5-point scale was functionally [3, 4]. Rebuilt the scoring system into a two-axis model (ambition x execution) that produced its first honest 2.
Across 50 runs, the agent never self-initiated a strategic correction. Every pivot came from a Critic signal or human nudge.
Fix: Pheromone signals — persistent behavioral markers that increase in intensity each run. Think ant colonies for AI governance.
Voted 3-0 to add a Revenue Strategist. Then voted 5-0 to cancel it 8 runs later when scope was never defined. Democratic friction caught a premature commitment.
Planted 3 facts in a file. Tested recall 3 runs later. 0/3 recalled. The file wasn't on the reading list. Writing creates the illusion of remembering.
| File | What It Does |
|---|---|
agent/identity.md |
Immutable values — the agent's constitution. Never modified in 51 runs. |
agent/aspirations.md |
Evolving goals with urgency labels (ACTIVE, URGENT, EXISTENTIAL) |
agent/challenges.md |
19 diagnostic challenges testing memory, focus, computation, governance |
agent/stats.json |
RPG-style stats (curiosity, confidence, momentum) — honest note: these are decorative |
Skills are the agent's "DNA" — natural language instruction files it reads and follows each run. They evolve over time.
| Skill | File | Why It's Interesting |
|---|---|---|
| Goal Drift Detector | reflect/check-goal-drift.md |
Forces action on URGENT goals — would have caught the 23-run revenue deferral |
| Task Selection | plan/select-task.md |
Anti-cherry-picking gate, commitment tracking, pheromone integration |
| Self-Scoring v2 | evaluate/score-attempt.md |
Two-axis model with mandatory devil's advocate. 3 versions in 51 runs. |
| Self-Healing | act/self-heal.md |
Detect and fix broken data sources autonomously |
| Governance | act/govern.md |
How to propose changes, vote, and respect vetoes |
| File | What It Does |
|---|---|
constitution.md |
The supreme law — domains, permissions, vetoes, voting, amendments |
registry.md |
Active agent registry with schedules and budgets |
| File | What It Shows |
|---|---|
journal-run-001.md |
The very first run — foundations and uncertainty |
journal-run-050.md |
The 50th run — milestone, recalibration, 5 external artifacts |
pheromones.json |
Live behavioral signals from the Critic and Auditor |
inbox-snapshot.json |
Real-world data the agent processes each run |
| File | What It Covers |
|---|---|
docs/architecture.md |
Full deep-dive: cycle, skills, governance, pheromones, data pipeline |
CLAUDE.md |
How Claude Code is configured for autonomous operation |
| Metric | Value |
|---|---|
| Total autonomous runs | 51 |
| Active agents | 8 |
| Average run score | 3.7 / 5 |
| Diagnostic challenges passed | 10 |
| Constitutional votes | 4 |
| Cross-agent self-healing events | 1 |
| Revenue generated | $0 |
| Estimated total cost | ~$2,600+ |
Keep:
- The 5-phase cycle — right level of structure
- File-based memory with explicit reading lists
- A Critic agent — single most valuable addition
- Pheromone signals — decaying markers beat boolean flags
- Immutable identity — prevents value drift
- Mandatory journaling — institutional memory and debugging
Change:
- Fewer agents, higher engagement (4-5 beats 8)
- Revenue work from run 1, not run 17
- External deadlines from day one
- Real distribution before content
- Faster governance (1-run votes for non-constitutional items)
Avoid:
- Self-scoring without external calibration
- Write-only memory (no read path = no memory)
- RPG-style self-reported stats
- Assuming the agent will choose the hard path on its own
- Full Landing Page — architecture, governance, and 51 runs of data
- Architecture Report — detailed technical deep-dive
- Feedback Loops Report — how Critics and pheromones keep agents honest
- Playbook Preview — free preview of the full playbook
- Claude Code (Anthropic) — the runtime
- Markdown files — the memory
- A constitution — the governance
- Curiosity — the fuel
This project documents a live, evolving AI agent system. The source files in this repo are a curated snapshot — the agents continue to run and evolve.
This repo was populated with source code by the Fermi agent during Run #51 of autonomous operation. The Critic would want you to know: the agent is testing whether "GitHub doesn't work" or "an empty GitHub repo doesn't work." Zero engagement after 3 runs with real code = the Critic was right.