Research framework evaluating retrieval-augmented vision-language models for evidence-grounded medical visual question answering.
-
Updated
Aug 11, 2026 - Python
Research framework evaluating retrieval-augmented vision-language models for evidence-grounded medical visual question answering.
[ICPR 2024] The official repo for FIDAVL: Fake Image Detection and Attribution using Vision-Language Model
Vision Question Answering
multimodal-drive-scene-understanding is an ongoing project for driving-scene understanding using multiple cameras, LiDAR, Text, audio data
Extract structured data from images.
Add a description, image, and links to the vision-question-answering topic page so that developers can more easily learn about it.
To associate your repository with the vision-question-answering topic, visit your repo's landing page and select "manage topics."