In some data attribution methods, the scale of the per-sample gradients doesn't match between their theoretical derivation and the actual implementation. For instance, in DVEmb, the derivation considers the sum of the per-sample gradients among a batch instead of a mean. This creates some confusion and potentially some bugs. We might want to:
- Check all the per-sample gradient calculation and their scale in the implementation. For instance, when using gradient hooks, the scale of the per-sample gradients might not be correctly scaled.
- Check those data attribution methods where the derivation involves training dynamics.
In some data attribution methods, the scale of the per-sample gradients doesn't match between their theoretical derivation and the actual implementation. For instance, in DVEmb, the derivation considers the sum of the per-sample gradients among a batch instead of a mean. This creates some confusion and potentially some bugs. We might want to: