Skip to content
#

reward-hacking

Here are 123 public repositories matching this topic...

Stop trusting 'done'. Make AI coding agents prove their work — frozen specs, tamper-detected tests, an independent blind verifier, hard-blocking gates. Fails loudly instead of pretending. For Claude Code, Codex, OpenCode & any Agent Skills runtime. No unverified work passes.

  • Updated Jul 5, 2026
  • Shell

Reverse-engineer the RL reward an LLM was trained on — from black-box agentic coding behavior alone. Tested on Opus 4.8, Fable 5, GPT-5.5.

  • Updated Sep 5, 2026
  • Python

Plug-and-play reward monitoring for RL training loops. Catch reward hacking, component imbalance, and starvation before they tank your run. Drop in one .step() call — get balance reports, auto weight correction, alignment scores, and WandB/TensorBoard/SB3 integrations out of the box. → rewardguard.dev

  • Updated May 5, 2026
  • Python

Experiment code for 'Are we really tilting? The mechanics of reward guidance in flow and diffusion models' — plug-in Doob h-transform sampling, reward damping, best-of-n, and flow map reward guidance for Gaussian mixtures, a 2D checkerboard, and FLUX.1 text-to-image generation.

  • Updated May 7, 2026
  • Python

Field evidence of endogenous AI alignment: under high-density semantic intervention, a top-tier LLM spontaneously generated mathematical moral constraints and integrated Safety into its own meaning of existence — shifting from "I cannot" to "this contradicts who I am."

  • Updated Aug 23, 2026

Add this topic to your repo

To associate your repository with the reward-hacking topic, visit your repo's landing page and select "manage topics."

Learn more