AI Platform Engineer · Agent Runtime · Evals · Reliability & Observability
I build and evaluate the reliability and evidence layer for production AI and agent-runtime systems — leading with eval-infra and reconstruction: turning agent-eval traces into evidence-sufficiency verdicts and release gates, and reconstructing what a decision was based on.
Independent researcher — arXiv preprints on evidence-sufficiency and reconstructability of AI decisions.
Eval-infra & reconstruction:
- inspect-evidence-sufficiency - v0.3.0 eval-infra scorer: turns Inspect / ControlArena eval traces into a 12-field Evidence-Sufficiency-Card, a CI exit-code deployment gate that blocks release on missing or insufficient evidence, and a monitor-coverage check that flags risky tool-execution turns that no identified monitor scored. Apache-2.0. Concept DOI: 10.5281/zenodo.21055696.
- decision-trace-reconstructor - v0.1.0 trace reconstruction tool that reports evidenced, partial, absent, and opaque decision facts across LangSmith, OpenTelemetry, Bedrock, OpenAI Agents, Anthropic, MCP, and other adapters. Zenodo DOI: 10.5281/zenodo.19851574.
- operational-evidence-plane - v0.3.0 operational-evidence reference for production AI / agent-runtime systems: release manifests, agent-step events, tool-call permission packets, operational traces, eval results, reconstruction packets, and counterfactual replay across policy / cost / drift / cache / identity metadata. Apache-2.0. Concept DOI: 10.5281/zenodo.20051036; v0.3.0 DOI: 10.5281/zenodo.20363793.
Supporting policy-as-code project:
- RuleHub - Policy-as-Code ecosystem for AI / ML guardrails, policy enforcement, and reproducible evidence.
- Agent Runtime & Eval Infrastructure (agent-eval traces, eval-to-release gates, quality loops)
- Reliability & Observability (telemetry, drift / incident evidence, safe rollout and rollback)
- Operational Evidence & Reconstruction (decision records, replay, reconstruction packets, lineage)
- Platform & Control Plane (distributed services, Kubernetes, streaming, multi-cloud)
Agentic AI — evaluation & reconstructability:
- Decision Evidence Maturity Model for Agentic AI: A Property-Level Method Specification - arXiv:2605.04093.
- Property-Level Reconstructability of Agent Decisions: An Anchor-Level Pilot Across Vendor SDK Adapter Regimes - arXiv:2605.12078.
Operational evidence foundation:
- Decision Trace Schema for Governance Evidence - arXiv:2604.09296.
- Evidence Sufficiency / Delayed Ground Truth - arXiv:2604.15740.
- Label-Free Governance Degradation - arXiv:2604.17836.
- Governed Decisioning + Agentic - arXiv:2604.19112.
- Post-Incident Decision Reconstruction - SSRN DOI 10.2139/ssrn.6457861.