I tried generating videos using around 10 different input images with varying lighting conditions and backgrounds. All the images were in 1920×1080 resolution. However, about 5 out of the 10 generated videos showed significant facial distortions or diffusion. In 2 cases, the lip movements were not generated properly at all.
I wanted to understand if there are any specific image requirements that the model works best with. For example:
- Does the model prefer a particular image resolution or aspect ratio?
- Are there any recommendations around face focus, framing, or camera angle?
- What kind of lighting conditions work best (bright, neutral, studio lighting, etc.)?
- Are there any constraints on background complexity?
- Is there a minimum face size or face-to-frame ratio that should be maintained?
What confuses me is that some images with darker backgrounds generated excellent results, while other images with brighter and seemingly better lighting produced poor outputs.
Are there any specific quality metrics, preprocessing steps, or parameter values that we should look at to achieve more consistent and reliable video generation across different input images?
I tried generating videos using around 10 different input images with varying lighting conditions and backgrounds. All the images were in 1920×1080 resolution. However, about 5 out of the 10 generated videos showed significant facial distortions or diffusion. In 2 cases, the lip movements were not generated properly at all.
I wanted to understand if there are any specific image requirements that the model works best with. For example:
What confuses me is that some images with darker backgrounds generated excellent results, while other images with brighter and seemingly better lighting produced poor outputs.
Are there any specific quality metrics, preprocessing steps, or parameter values that we should look at to achieve more consistent and reliable video generation across different input images?