Skip to content

Evaluation Time Issue #23

Description

@Lorsense

Hello authors,

Thank you very much for open-sourcing this great work.

I am currently reproducing your work using 8 × NVIDIA A800 GPUs, and I have encountered some issues regarding the evaluation that is performed every 20 training steps.

During evaluation, I found that the evaluation process takes an extremely long time (more than 36 hours), and it still does not produce the final evaluation results. I also tried evaluating intermediate checkpoints using run_eval_aime.sh, but encountered the same issue.

It is worth noting that during evaluation, I frequently see errors such as:

API Request Error: Gateway Timeout
Sandbox API call failed. Last error: ...

Similar errors also occur during training. However, during training, each step takes approximately 30 minutes on average, whereas the evaluation process seems to run indefinitely without completing.

I would like to ask what might be causing this issue. Is such a long evaluation time expected, or could it be related to the sandbox service, API gateway timeouts, or some configuration issue?

I would greatly appreciate any advice or suggestions on how to diagnose and resolve this problem.

Thank you very much for your time and help!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions