Detect behavioural drift between LLM versions before you upgrade. Compare model responses, classify regressions, and generate migration reports with validated prompt patches.
-
Updated
Jul 18, 2026 - Rust
Detect behavioural drift between LLM versions before you upgrade. Compare model responses, classify regressions, and generate migration reports with validated prompt patches.
git diff for prompt engineers
Hey LLM, you okay? — pyramid-ordered LLM testing CLI for CI/CD. One YAML for every layer, LLM-as-a-judge gates, and A/B triage that tells prompt regressions from model drift.
AI/LLM test strategy for an e-commerce product recommendation engine — prompt regression, hallucination detection, toxicity safety gate, latency SLOs, and contract testing
Catch prompt regressions from model drift — on a schedule, not just on PRs.
PromptOps — Evaluate, improve, test, and run your prompts in Claude Code. Score prompts 1–5, auto-improve with guardrails, regression test against golden datasets, benchmark costs across models, and execute — all in one session.
Audit log for AI storyboard prompt regressions, frame continuity, camera changes, scene coverage, and fixes.
Portfolio-grade AI quality evaluation lab with golden datasets, prompt regression, groundedness checks, hallucination tests and CI thresholds.
CI-ready quality gates for LLM/RAG systems: response quality, hallucination risk, retrieval metrics, latency SLAs, and prompt regression testing.
Offline prompt regression CI checks for OpenAI-compatible gateways, model routes, JSON output, and tool-call readiness.
AI agent prompt regression test template mirror for Codex, Claude Code, Cursor teams. Routes to $203 Agent Ops team license.
Add a description, image, and links to the prompt-regression topic page so that developers can more easily learn about it.
To associate your repository with the prompt-regression topic, visit your repo's landing page and select "manage topics."