Penalize the Path, Reward the Outcome — verifiable per-action penalties as a dense channel for deployable, sample-efficient agentic RL (GRPO). Paper: arXiv:2607.07435
-
Updated
Aug 12, 2026 - Python
Penalize the Path, Reward the Outcome — verifiable per-action penalties as a dense channel for deployable, sample-efficient agentic RL (GRPO). Paper: arXiv:2607.07435
Gymnasium-style RL framework for LLM agent training — MDP environments, three-layer process reward & SFT/DPO/GRPO policy optimization. CLI + MCP ready.
Official implementation of "Advancing Reasoning in Diffusion Language Models with Denoising Process Rewards" (ACL 2026).
Compiler feedback as process reward for coding agent RL training (Junhao Fu, 2025)
LLM/Agent 发布评测决策:数据治理、Judge 校准、安全回归、漂移与门禁
To associate your repository with the process-reward topic, visit your repo's landing page and select "manage topics."