May I ask which image the caption corresponds to? During training, the resolution is high (GT of the target resolution, like 1024 * 1024), while during testing, it is low resolution (for example, the input is a downsampled 4x low resolution image 256 * 256 to generate caption)
May I ask which image the caption corresponds to? During training, the resolution is high (GT of the target resolution, like 1024 * 1024), while during testing, it is low resolution (for example, the input is a downsampled 4x low resolution image 256 * 256 to generate caption)