比较好的是里面介绍了三个 objectives:Image-Text matching, Image-Grounded Text Generation, Image-Text Constrastive Learning