🎯
Focusing
working on post-training, reinforcement learning, and agents.
Pinned Loading
-
tinyvlm-implementation
tinyvlm-implementation PublicImplementation of tiny modular VLM with scaling study using FSDP
Python 1
-
agent-credit-bench
agent-credit-bench PublicExact-oracle conformance tests for turn-level credit assignment in agentic RL (GRPO, RLOO, GAE, GiGPO; verl, TRL, OpenRLHF).
Python 6
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.

