Hi, thanks for your great work on FastRL!
I found that Eagle1 speculative decoding successfully accelerates inference benchmarking, but does not bring any speedup in end-to-end GRPO training.
Running bash examples/bench_sd.sh:

Running bash examples/grpo_7B.sh:
Autoregressive:

+Eagle1 without draft training

I was using 8 × H800 80GB, Qwen2.5-7B-instruct+Qwen2.5-7B-Eagle-RL
Hi, thanks for your great work on FastRL!
I found that Eagle1 speculative decoding successfully accelerates inference benchmarking, but does not bring any speedup in end-to-end GRPO training.
Running bash examples/bench_sd.sh:

Running bash examples/grpo_7B.sh:


Autoregressive:
+Eagle1 without draft training
I was using 8 × H800 80GB, Qwen2.5-7B-instruct+Qwen2.5-7B-Eagle-RL