RAG evaluation toolkit. Full spec in tapat-prd-v0.md — read it before implementing anything. The PRD is the source of truth; if my instruction conflicts with it, flag the conflict instead of guessing.
- Python 3.12, Typer (CLI), Pydantic v2 (all schemas), httpx (async calls)
- pytest for tests, ruff for lint, pyright for types
- No other dependencies without asking me first
- Type hints on everything. Pydantic models for any data crossing a boundary.
- Every metric function ships with unit tests in the same PR. No exceptions.
- Tests NEVER call real APIs. Use MockAdapter and fixtures only.
- No API keys in code, ever. Config via tapat.yaml + env vars.
- Small, reviewable changes. One module per task. Don't refactor code I didn't ask about.
- v0 scope only: no web UI, no dashboard, no database. Section 3 non-goals are binding.
- src layout: package code in tapat/, tests mirror the package structure
- Judge prompts live in judges/*.md, versioned, never inline in Python