Thanks a lot for your awesome work on Verl!
Just wondering—are there any plans to add multi-turn RL training?
A couple of quick ideas:
- Try multi-turn tasks and compare with non-quantized runs
- Make agent_loop work with importance sampling in Verl
Would love to hear your thoughts. Thanks again!
Thanks a lot for your awesome work on Verl!
Just wondering—are there any plans to add multi-turn RL training?
A couple of quick ideas:
Would love to hear your thoughts. Thanks again!