Dataflow-Oriented Reinforcement Learning for (Multi-)Agentic LLMs
-
Updated
Jul 26, 2026 - Python
Dataflow-Oriented Reinforcement Learning for (Multi-)Agentic LLMs
A repo for survey paper "The Evolving Landscape of LLM- and VLM-Integrated Reinforcement Learning" and a collection of AWESOME papers focused on using LLMs, VLMs for improving RL.
Official Implementation of "J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data" (https://arxiv.org/abs/2608.26582)
LMGym: Distributed Orchestration for LLM Reinforcement Learning(面向 LLM 强化学习的分布式编排方案)
LMCode: Continual Training and Serving for LLM Coding Agent(面向 LLM Coding Agent 的继续训练与服务方案)
LMChat: Continual Fine-Tuning and Serving for LLM Chatbot(面向 LLM Chatbot 的继续微调与服务方案)
To associate your repository with the llm-rl topic, visit your repo's landing page and select "manage topics."