Skip to content

Per-sample Gradient Scaling #219

Description

@sleepymalc

In some data attribution methods, the scale of the per-sample gradients doesn't match between their theoretical derivation and the actual implementation. For instance, in DVEmb, the derivation considers the sum of the per-sample gradients among a batch instead of a mean. This creates some confusion and potentially some bugs. We might want to:

  1. Check all the per-sample gradient calculation and their scale in the implementation. For instance, when using gradient hooks, the scale of the per-sample gradients might not be correctly scaled.
  2. Check those data attribution methods where the derivation involves training dynamics.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions