Multi-Modal learning toolkit based on PaddlePaddle and PyTorch, supporting multiple applications such as multi-modal classification, cross-modal retrieval and image caption.
-
Updated
May 7, 2023 - Python
Multi-Modal learning toolkit based on PaddlePaddle and PyTorch, supporting multiple applications such as multi-modal classification, cross-modal retrieval and image caption.
The source code of AMFMN and the dataset RSITMD
[IJCAI2022] Unsupervised Voice-Face Representation Learning by Cross-Modal Prototype Contrast
Next-gen Cross-Modal Graph RAG engine bridging video keyframes, Whisper speech transcripts, architecture diagrams, and technical PDFs into an interconnected knowledge graph with pgvector and Gemini.
To associate your repository with the crossmodal-retrieval topic, visit your repo's landing page and select "manage topics."