Production AI systems that survive contact with reality
Scott Hardie · Solutions Architect at McGraw Hill · Independent AI Systems Builder · Toronto, Canada
Systems · How I help · Architecture · Lab · LinkedIn
PRODUCTION AI · AGENT CONTROL PLANES · LOCAL INFERENCE · FINOPS · VERIFICATION
AI demos are easy. Production systems must survive retries, partial failure, hostile inputs, runaway spend, model drift, and an auditor asking exactly what happened.
I design and build the infrastructure around the model: observable workflows, explicit policy boundaries, controlled execution, replayable evidence, and reconciled outcomes.
Intelligence can be probabilistic. Infrastructure cannot.
| Observe | Control | Prove |
|---|---|---|
| Capture model, tool, cost, and transaction events. | Route workloads, enforce policy, isolate risk, and recover safely. | Replay decisions, verify state, reconcile money, and export evidence. |
| If you are… | Start here | What you get |
|---|---|---|
| An operations or engineering leader with a brittle AI workflow | Book a free 30-minute diagnostic | A constraint map, quick-win assessment, and candid next step. |
| Evaluating private or local AI | Run the free AI lab audit | A fast readiness signal before spending on infrastructure. |
| Reviewing the engineering | Inspect the public evidence | Architecture, tests, workflows, and conservative maturity labels. |
These are the four clearest public examples of the approach. The table is generated from a versioned manifest, checked against GitHub every week, and deliberately separates released, beta, and research work.
Public project metadata last verified 2026-09-26 · source manifest · verification policy
| Project | Problem | Public evidence |
|---|---|---|
| TokenGoblin · Go Measure · beta |
LLM workloads need cost, usage, and routing data before teams can control inference spend. | Public ingestion benchmarks, cost and routing tests, and a repository-level CI workflow. Architecture · Evidence |
| ReadyLayer · TypeScript Govern · beta |
AI-assisted delivery needs policy, review, and evidence before generated changes reach production. | Public policy contracts, evidence export documentation, test suites, and CI quality gates. Architecture · Evidence |
| veridag · Rust Prove · research |
Distributed execution needs explicit ordering, capability security, and cross-language conformance. | A public protocol specification, Quint formal models, test vectors, and dedicated conformance workflows. Architecture · Evidence |
| Settler · TypeScript Reconcile · beta |
Payment, banking, and operational records diverge unless matching and evidence rules are explicit. | Public reconciliation benchmark source and checked-in snapshots, with CI and security-invariant workflows. Architecture · Evidence |
| Engagement | Best when | Outcome |
|---|---|---|
| AI clarity audit | The opportunity is real, but the workflow and risk boundaries are not yet clear. | A decision-ready map of constraints, ownership, ROI assumptions, and the smallest safe pilot. |
| Stabilization sprint | An AI workflow is live but flaky, opaque, or expensive. | Explicit contracts, retries, telemetry, fallbacks, acceptance tests, and an operator runbook. |
| Governance architecture | Agents or models can take consequential actions. | Approval boundaries, policy gates, audit trails, incident paths, and evidence you can inspect. |
| Local AI systems | Data control, predictable cost, or offline capability matters. | Model and hardware fit, routing, deployment, observability, and a practical operating plan. |
Every engagement starts with the workflow—not a predetermined model or platform. See the service details, case studies, or book a diagnostic.
Seven monorepos keep related systems coherent while preserving clear boundaries:
| Boundary | Public monorepos | Responsibility |
|---|---|---|
| Observe + control | agent-edge · agent-infra | Agent traffic, policy, governance, mission state, and MCP boundaries. |
| Model + execute | model-tools · autopilot | Inference routing, GPU fit, and runnerless ops, support, growth, and FinOps workflows. |
| Integrate + operate | api-tools · ops-tools | APIs, webhooks, continuity, drift inspection, and golden paths. |
| Consumer outcomes | consumer-tools | Warranty, review intelligence, and inbox automation. |
Open the component map
signal policy execution
agent-edge ──────────────► agent-infra ──────────────► autopilot
packet capture control plane ops / finops
mesh edge agent mesh growth / support
MCP firewall
│ │ │
└──────────────────── model-tools ◄───────────────────┘
inference / routing
│
local GPU testbed
│
evidence / reconciliation
Each public monorepo includes an ARCHITECTURE.md describing its boundary and migration history.
Hardonia includes an owned, local testbed for model routing, image and video workflows, and failure-mode testing. It is where local-first claims are exercised before they become architecture advice.
| Lane | Hardware | Primary use |
|---|---|---|
| Heavy inference | NVIDIA V100 · 16 GB | Larger model and video-generation workloads. |
| Memory-oriented | NVIDIA P40 · 24 GB | ComfyUI pipelines, quantized models, and training experiments. |
| Interactive | NVIDIA RTX 3060 · 12 GB | Vision, embeddings, and latency-sensitive workflows. |
The lab uses Ollama-compatible routing, ComfyUI, containerized services, and Prometheus/Grafana-style observability. Public implementation lives primarily in model-tools and api-tools.
Decision-layer experiment: where Jev fits
TypeSafe Jev is being evaluated for low-cost typed classification where a general-purpose LLM is unnecessary—for example intent classification, routing, and tool-selection gates. It complements deterministic policy; it does not replace it. TypeSafe currently lists input pricing at $42 per billion tokens.
Trust, execution, and agent systems
- Requiem — unified AI control plane and execution contracts.
- Nautilus — local AI execution, orchestration, and policy enforcement.
- Keys — auditable mission control for constrained agents.
- truthcore — verification kernel and offline evidence reports.
- Zeo — governance, policy enforcement, and deterministic audit trails.
- agent-governance — enforceable agent laws and a governance gateway.
Applied systems and simulation
- SawyerCore — deterministic edge-AI runtime and agent simulation.
- WorldForge — deterministic, moddable simulation runtime.
- World26 — open planetary-systems simulator.
- FlexibleAccessible — accessibility auditing and remediation workflows.
- MortgageMatchPro — mortgage scenario and matching platform.
Productized audits, kits, and workflows
| Starting point | Intended outcome |
|---|---|
| SaaS repo rescue | Find auth, billing, webhook, RLS, and reliability gaps before they leak revenue. |
| TokenGoblin cost optimizer | Measure and control model-inference spend. |
| Local AI lab audit | Review hardware fit, routing, security, and operating posture. |
| AI command center | Replace operational blind spots with health and priority signals. |
| ComfyUI workflow packs | Run repeatable private image-production workflows on owned compute. |
| Principle | Working rule |
|---|---|
| Evidence over confidence | If a run cannot be inspected or replayed, it is not production-ready. |
| Local-first where it earns its keep | Own the compute, data boundary, fallback path, and cost model when the trade-off is justified. |
| Determinism at the edges | Keep probabilistic intelligence inside explicit policy, schema, and execution constraints. |
| Boring reliability wins | Idempotency, row-level security, state machines, and observable queues beat hidden cleverness. |
| Revenue is a reconciled event | A dashboard row is not money; provider-correlated settlement evidence is money. |
| Fix the smallest root cause | Isolate the failure, repair it surgically, prove the result, then ship. |
Verify this profile locally
git clone https://github.com/Hardonian/Hardonian.git
cd Hardonian
uv run python -m unittest discover tests -v
uv run python scripts/profile-metadata.py --check --verify-remote
uv run python scripts/profile-link-audit.pyThe checks reject stale project metadata, missing public evidence, broken local assets, and dead links.
I am a Solutions Architect at McGraw Hill and build Hardonia independently from Toronto. Alongside that work, I contribute part-time expertise to confidential frontier-AI evaluation and systems initiatives; client, model, dataset, and internal research details remain private.
The common thread is practical systems work: integrations, production reliability, agent governance, local inference, financial controls, and technical evaluation with explicit evidence boundaries.
If you have an AI workflow that is expensive, unreliable, hard to govern, or stuck between prototype and production, send three things: what the workflow does, where it fails, and what a measurable win would look like.



