Thanks for your amazing work!
I've been trying to replace CUT3R backbone and found out that Human3R's scene prediction aligns well with the GT depth while other metric-depth estimation like Depth-Anything 3 fails to do so. I think this alignment is critical to train SMPL's translation parameter. Is there any idea how you managed to align the scene points to the GT?
Fig. 1. Initial sample from Human3R training (Grey: prediction, Green: GT)
Fig. 2. DA3-metric result (Blue: prediction, Green: GT)
Thanks for your amazing work!
I've been trying to replace CUT3R backbone and found out that Human3R's scene prediction aligns well with the GT depth while other metric-depth estimation like Depth-Anything 3 fails to do so. I think this alignment is critical to train SMPL's translation parameter. Is there any idea how you managed to align the scene points to the GT?
Fig. 1. Initial sample from Human3R training (Grey: prediction, Green: GT)
Fig. 2. DA3-metric result (Blue: prediction, Green: GT)