Collaborative research exploring multimodal question answering using OCR, RAG, and document/image understanding techniques.
-
Updated
Mar 5, 2026
Collaborative research exploring multimodal question answering using OCR, RAG, and document/image understanding techniques.
For the work on "MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing"
Quasar PoC, Multitenant PoC.
(IJCNLP-AACL 2025) MAMMQA: Rethinking Information Synthesis in Multimodal Question Answering from a Multi-Agent Perspective
MMTabReal is a benchmark suite for multimodal table reasoning on real-world tables with mixed text, images, charts, maps, and visual encodings. It includes curated QA pairs, baseline implementations, and evaluation scripts for multimodal table understanding research.
To associate your repository with the multimodal-question-answering topic, visit your repo's landing page and select "manage topics."