Skip to content

Latest commit

 

History

History
50 lines (42 loc) · 2.26 KB

File metadata and controls

50 lines (42 loc) · 2.26 KB

Jev Skill

What It Does

This skill builds typed Jev decision workflows and evaluates them over reviewed golden sets. It supports TypeSafe Jev, the community OpenJev native server, the open-weight OpenJev-HF checkpoint through SGLang, the LocalJev Bun bridge, local Ollama, and OpenAI-compatible runtimes such as vLLM, LM Studio, and llama.cpp server. It also supports native /v1/systemone adapters for Von, LitJev, and Simple-JEV, plus dependency-free calibration fitting and matched-report comparison. It documents the optional official LangChain integration for using Jev as an agent routing and pre-tool decision layer, and includes provider-neutral control-plane policies for routing, fail-closed tool-risk assessment, and rubric judging. The included harness validates Noul, Choice, and Score answers; can aggregate independent samples; can run a typed verifier; and marks uncertain cases for review instead of silently forcing a decision.

Best For

  • Typed routing, ranking, extraction, verification, and bounded scoring.
  • Comparing Jev with local models on the same golden, adversarial, and metamorphic cases.
  • Measuring accuracy, calibration, consensus, latency, retries, token usage, and optional cost.
  • Building a safe local-model decision layer with schema-constrained output and explicit abstention.
  • Adding Jev decisions to an agent control plane without granting model output authority over permissions or side effects.

Not For

  • Claiming that a local checkpoint has Jev's weights or accuracy without a matched benchmark.
  • Replacing deterministic arithmetic, date handling, permissions, or policy enforcement with model output.
  • Treating confidence, consensus, or a verifier score as proof of truth.
  • Executing a tool directly from a Jev answer without deterministic policy and authorization checks.

Safety Notes

  • API keys come from environment variables and are never written to reports.
  • Input state is not printed or stored in reports; use appropriate privacy controls for golden sets and model endpoints.
  • Local output is strict JSON-schema validated by default; JSON repair is an explicit opt-in for diagnostics.
  • Low consensus or verifier support becomes decision: "review"; the harness does not silently override the proposed answer.