Experiment code for 'Are we really tilting? The mechanics of reward guidance in flow and diffusion models' — plug-in Doob h-transform sampling, reward damping, best-of-n, and flow map reward guidance for Gaussian mixtures, a 2D checkerboard, and FLUX.1 text-to-image generation.
flux text-to-image generative-models best-of-n diffusion-models flow-matching stochastic-interpolants reward-hacking reward-guidance doob-h-transform
-
Updated
May 7, 2026 - Python