Hello authors,
Thank you for releasing this amazing work! I really appreciate the approach of fusing VGGT-like features into video generation to ensure geometric consistency in video-based world generation. This seems like it would bring a significant improvement in quality when lifted to the 3D space via methods like 3DGS.
I have a couple of questions regarding the capabilities and extension of your pretrained models:
Full View Coverage (360° Orbit Videos): Is it currently possible to generate videos with full 360-degree orbit trajectories using your pretrained models? I am working on generating scenes that require full view coverage.
Extending Frame Count for Longer Trajectories: Is there a recommended way to extend the frame count for longer camera trajectories? I noticed that trying to squeeze a longer movement into the standard video length introduces severe motion blur (though it remains geometrically superior to other video-based methods).
This is an example generated using the default frame count and a supposed orbit trajectory, with the Wan2.2 model and identical images for the start and end images:
https://github.com/user-attachments/assets/9a4b14da-2dda-4c4b-b1c7-719e070a80d9
Thanks in advance for your help and insights!
Hello authors,
Thank you for releasing this amazing work! I really appreciate the approach of fusing VGGT-like features into video generation to ensure geometric consistency in video-based world generation. This seems like it would bring a significant improvement in quality when lifted to the 3D space via methods like 3DGS.
I have a couple of questions regarding the capabilities and extension of your pretrained models:
Full View Coverage (360° Orbit Videos): Is it currently possible to generate videos with full 360-degree orbit trajectories using your pretrained models? I am working on generating scenes that require full view coverage.
Extending Frame Count for Longer Trajectories: Is there a recommended way to extend the frame count for longer camera trajectories? I noticed that trying to squeeze a longer movement into the standard video length introduces severe motion blur (though it remains geometrically superior to other video-based methods).
This is an example generated using the default frame count and a supposed orbit trajectory, with the Wan2.2 model and identical images for the start and end images:
https://github.com/user-attachments/assets/9a4b14da-2dda-4c4b-b1c7-719e070a80d9
Thanks in advance for your help and insights!