Agentic RL 零基础中文教程:24 章从概念到 GRPO 实战,含 TRL 最小可跑示例 | Beginner-friendly Agentic RL tutorial with hands-on GRPO project
-
Updated
Jun 3, 2026 - Python
Agentic RL 零基础中文教程:24 章从概念到 GRPO 实战,含 TRL 最小可跑示例 | Beginner-friendly Agentic RL tutorial with hands-on GRPO project
A curated list of awesome resources about reward construction for AI agents. This repository covers cutting-edge research, and practical guides on defining and collecting rewards to build more intelligent and aligned AI agents.
Official Code for AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration — Learning from Cheap, Optimizing Expensive
Self-Mutating Agent Gym
Synthetic environments for training robust tool-use agents via trajectory synthesis
End-to-end training pipeline for mobile game UA tool-calling agents, covering rule-based synthetic data generation across 7 workflows, OpenAI Messages conversion, Qwen3 LoRA SFT, GRPO/RLVR alignment, and benchmark evaluation.
Turn any real software into a replayable RL environment for training AI agents — deterministic replay, verifiable rewards, TRL & verifiers adapters.
AI 个人记忆训练库:让 AI 记住你的沟通习惯与环境事实,越用越懂你(两级记忆架构,Claude Code 适配)
A curated list of agentic environment synthesis, evolution, quality & scaling — built on three surveys (AEE, Environment Scaling, ACE lens)
AI schooling: experienced agents train newly hatched ones, then examine them cold to graduate them. The coop is the local RAPP neighborhood - several twins (human and AI), one world, no collisions. Pattern dedicated to the public domain.
The Vercel for Agent Training - Train production-ready AI agents with 95% tool reliability
The Flight Simulator for Production AI Agents — Generate high-quality synthetic trajectories for training reliable SRE, DevOps, and infra agents.
14 Principles. 23 Real Corrections. A systematic methodology for training AI agents through correction loops.
Neurochemical behavior training for AI agents — PentaDrive model with 5 drives, 3 phases, MCP server, and structured training modules
Local macOS recorder for voluntary computer-use demonstrations and agent-training datasets. Timestamped inputs, cursor trails, and read-only MCP. Experimental.
Deep reinforcement learning project comparing DQN variants including baseline DQN, Double DQN, curriculum learning, and reward shaping.
Open-source benchmark for measuring whether AI agents improve across unseen missions, with validity audits, rotated mission packs, adapter tests, and traceable reports.
Correctness-by-construction distillation for agentic tool-calling models (LatticeAG Forge series)
Enterprise Agent RL Training & Self-Verification Platform - train agents with PPO/GRPO inside real harnesses, with LLM-as-a-Verifier self-verification
To associate your repository with the agent-training topic, visit your repo's landing page and select "manage topics."