Compile real-world Claude Code and Codex trajectories into verified, tradable post-training assets.
-
Updated
Aug 21, 2026 - Python
Compile real-world Claude Code and Codex trajectories into verified, tradable post-training assets.
Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.
Context & Guide For Reinforcement Learning with Verifiable Rewards with Large Language Models
Score the trustworthiness of outputs from any LLM in real-time
🐀 Fuzz your verifier before an RL agent does. Static + dynamic LLM security auditor to detect reward-hacking in RL post-training environments (OpenEnv, verifiers-spec, Gymnasium).
A curated list of rubrics, checklists, criteria sets, principles, and scoring guides used to score, rank, verify, filter, or train modern generative models.
Write loops, not prompts. The loop is easy — the verifier is the whole game. VCN night one at Network School.
Open Arena: SLM verifiers across observability platforms, dataset and model catalogs, and value scenarios for LLM, agentic and harness evals.
An RL Enviorment for AES Inversion
A verifiers RLM environment for testing whether adaptive recursive search outperforms brittle manual RAG choreography on long synthetic corpora.
Adversarial QA for LLM-RL environments: find out what reward an empty answer earns. Model-free, zero API cost.
Verifiable RL environments for corporate law & governance — deterministic reward, no LLM judge, fully synthetic worlds.
An open reinforcement-learning (RL) environment that trains LLM agents to use the current fact, not the stale one — verifiable reward for temporal fact-currency, built on verifiers / prime-rl (GRPO, LoRA).
Reproducible verifier audits, datasheets, agreement metrics, and release gates
Typed asset shapes + visual + headless views for AI agents. One asset definition. Three rendering targets (HTML / Markdown / Text).
RLVR coding environment for training and evaluating LLM agents on automated code-fixing tasks.
Research proposal for verifier-gated on-policy distillation with explicit evaluation and claim boundaries
A verifiable RL environment for TRP ion-channel ligand pharmacology, built on Prime Intellect's Verifiers
Deterministic SRT captioning environment for LLM evaluation and reinforcement learning.
Running, reproducing, and testing AI agent environments to understand how they work.
To associate your repository with the verifiers topic, visit your repo's landing page and select "manage topics."