Skip to content
View dev404ai's full-sized avatar

Highlights

  • Pro

Block or report dev404ai

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
dev404ai/README.md

Oleg Solozobov

AI Platform Engineer · Agent Runtime · Evals · Reliability & Observability

I build and evaluate the reliability and evidence layer for production AI and agent-runtime systems — leading with eval-infra and reconstruction: turning agent-eval traces into evidence-sufficiency verdicts and release gates, and reconstructing what a decision was based on.

Independent researcher — arXiv preprints on evidence-sufficiency and reconstructability of AI decisions.


Selected Open-Source Work

Eval-infra & reconstruction:

  • inspect-evidence-sufficiency - v0.3.0 eval-infra scorer: turns Inspect / ControlArena eval traces into a 12-field Evidence-Sufficiency-Card, a CI exit-code deployment gate that blocks release on missing or insufficient evidence, and a monitor-coverage check that flags risky tool-execution turns that no identified monitor scored. Apache-2.0. Concept DOI: 10.5281/zenodo.21055696.
  • decision-trace-reconstructor - v0.1.0 trace reconstruction tool that reports evidenced, partial, absent, and opaque decision facts across LangSmith, OpenTelemetry, Bedrock, OpenAI Agents, Anthropic, MCP, and other adapters. Zenodo DOI: 10.5281/zenodo.19851574.
  • operational-evidence-plane - v0.3.0 operational-evidence reference for production AI / agent-runtime systems: release manifests, agent-step events, tool-call permission packets, operational traces, eval results, reconstruction packets, and counterfactual replay across policy / cost / drift / cache / identity metadata. Apache-2.0. Concept DOI: 10.5281/zenodo.20051036; v0.3.0 DOI: 10.5281/zenodo.20363793.

Supporting policy-as-code project:

  • RuleHub - Policy-as-Code ecosystem for AI / ML guardrails, policy enforcement, and reproducible evidence.

Focus Areas

  • Agent Runtime & Eval Infrastructure (agent-eval traces, eval-to-release gates, quality loops)
  • Reliability & Observability (telemetry, drift / incident evidence, safe rollout and rollback)
  • Operational Evidence & Reconstruction (decision records, replay, reconstruction packets, lineage)
  • Platform & Control Plane (distributed services, Kubernetes, streaming, multi-cloud)

Research Preprints

Agentic AI — evaluation & reconstructability:

Operational evidence foundation:


Agent Runtime & Evals

Python Go OPA/Rego MCP OpenTelemetry Evals Replay


Platform, Cloud & Data

Kubernetes Terraform AWS GCP Kafka Flink ClickHouse PostgreSQL


Observability & Controls

Prometheus Grafana Vault OAuth2/OIDC

Pinned Loading

  1. rulehub/rulehub rulehub/rulehub Public

    Policy-as-Code guardrails for ML and LLM systems: OPA/Kyverno policies, compliance mappings, signed bundles, evidence trails, and plugin index.

    Python 5 1

  2. costscope/costscope costscope/costscope Public

    Open FinOps and governance data plane for FOCUS 1.2 cost normalization across cloud, on-prem, GPU, and AI/LLM workloads.

    Go 1 1

  3. governance-evidence/decision-event-schema governance-evidence/decision-event-schema Public

    JSON Schema for decision events as governance evidence units in automated decision and real-time risk systems. MIT.

    Python 1 1