This skill builds typed Jev decision workflows and evaluates them over
reviewed golden sets. It supports TypeSafe Jev, the community OpenJev native
server, the open-weight OpenJev-HF checkpoint through SGLang, the LocalJev Bun
bridge, local Ollama, and OpenAI-compatible runtimes such as vLLM, LM Studio,
and llama.cpp server.
It also supports native /v1/systemone adapters for Von, LitJev, and
Simple-JEV, plus dependency-free calibration fitting and matched-report
comparison. It documents the optional official LangChain integration for using Jev as
an agent routing and pre-tool decision layer, and includes provider-neutral
control-plane policies for routing, fail-closed tool-risk assessment, and
rubric judging.
The included harness validates Noul, Choice, and Score answers; can aggregate
independent samples; can run a typed verifier; and marks uncertain cases for
review instead of silently forcing a decision.
- Typed routing, ranking, extraction, verification, and bounded scoring.
- Comparing Jev with local models on the same golden, adversarial, and metamorphic cases.
- Measuring accuracy, calibration, consensus, latency, retries, token usage, and optional cost.
- Building a safe local-model decision layer with schema-constrained output and explicit abstention.
- Adding Jev decisions to an agent control plane without granting model output authority over permissions or side effects.
- Claiming that a local checkpoint has Jev's weights or accuracy without a matched benchmark.
- Replacing deterministic arithmetic, date handling, permissions, or policy enforcement with model output.
- Treating confidence, consensus, or a verifier score as proof of truth.
- Executing a tool directly from a Jev answer without deterministic policy and authorization checks.
- API keys come from environment variables and are never written to reports.
- Input state is not printed or stored in reports; use appropriate privacy controls for golden sets and model endpoints.
- Local output is strict JSON-schema validated by default; JSON repair is an explicit opt-in for diagnostics.
- Low consensus or verifier support becomes
decision: "review"; the harness does not silently override the proposed answer.