Hello,
Thank you for the awesome work!
We are trying to reproduce the data processing pipeline and we cannot seem to find anywhere a description of the exact process of ROI-aware captioning, VQA generation and the filtering of QA pairs. We assume LLMs were used for all three. Could you kindly provide the LLM prompts used to generate these?
Thanks in advance!
Hello,
Thank you for the awesome work!
We are trying to reproduce the data processing pipeline and we cannot seem to find anywhere a description of the exact process of ROI-aware captioning, VQA generation and the filtering of QA pairs. We assume LLMs were used for all three. Could you kindly provide the LLM prompts used to generate these?
Thanks in advance!