Skip to content

Source of Tools and Tool-Calling Judgment in Multi-Step Reasoning #21

Description

@Lorsense

Hello, the scenario discussed in the paper is tool-integrated reasoning. However, during the model distillation stage, I did not see a clear description of the source of the tools or the tool-calling process. In addition, I do not fully understand how the paper determines whether a tool call is incorrect at a certain step. I would be very grateful if you could kindly clarify these points despite your busy schedule. Thank you.

Activity

  1. YoungZ365 commented on Jul 11, 2026

    @YoungZ365
    Owner

    Thank you for your interest in our work and for your thoughtful question.

    Regarding the first question, in our work the tools are provided through SandboxFusion, which offers an isolated Python execution environment. During rollout, the student model interacts with the environment in the standard agent manner: it generates Python code, the sandbox executes the code, and the execution result is returned to the model as part of the subsequent context. The model then continues reasoning based on this observation. Therefore, the tool itself is not part of the model, but part of the external environment that both training and evaluation interact with.

    For the second question, SOD does not explicitly determine whether a tool call is "correct" or "incorrect" at any particular step. This is a common misunderstanding.

    Instead, SOD measures the divergence between the teacher and student policies at each reasoning step. Specifically, after each tool interaction, we compute the average token-level log-probability difference between the teacher and student over the student's generated tokens, which is denoted as $d_k$ in the paper. The step-wise distillation weight is then adjusted according to these divergences.

    The intuition is that an incorrect tool call often produces an incorrect execution result, which changes the subsequent reasoning context. As the student continues reasoning on this altered context, the divergence between the teacher and student naturally increases. Consequently, SOD automatically reduces the distillation weight for those later steps.

    Therefore, SOD never needs an explicit binary signal such as "this tool call is correct" or "this tool call is wrong." The effect is captured implicitly through the increasing teacher–student divergence after the environment state has drifted.

    In other words, the method is divergence-driven rather than error-label-driven. This design avoids the need to annotate or detect incorrect tool calls explicitly while still reducing the influence of unreliable supervision in trajectories where the student has significantly deviated from the teacher.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions