Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
-
Updated
Oct 30, 2025 - Python
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
Reproducible scaling laws for contrastive language-image learning (https://arxiv.org/abs/2212.07143)
Ultralytics fork of Apple MobileCLIP for fast image-text inference, training, evaluation, and an iOS demo.
Using Segment-Anything and CLIP to generate pixel-aligned semantic features.
Model merging, task-vector rebasin, and fine-tuning for vision and LLM models.
[Official] [IROS 2024] A goal-oriented planning to lift VLN performance for Closed-Loop Navigation: Simple, Yet Effective
Clipora is a powerful toolkit for fine-tuning OpenCLIP models using Low Rank Adapters (LoRA).
I switched phones and found thousands of WhatsApp photos waiting for me, mostly memes and screenshots. This finds the ones worth keeping: search your pictures by describing them, and bin the rest into a quarantine you can undo.
Turn any YouTube video into viral clips for free!
A performant cross-platform image viewer and editor built with Rust and Slint, with an extensible plugin system.
A simple open-sourced SigLIP model finetuned on Genshin Impact's image-text pairs.
Text-to-image search with OpenCLIP, Docker, Flask, Faiss, etc. and a basic front-end.
A doctor-assistive AI system that interprets medical knowledge and patient images simultaneously. It utilizes a Dual-Encoder architecture to cross-reference textbook theory with visual pathology, generating clinically grounded diagnoses.
Mori_Cloud is a web platform that allows users to store, manage, and share memorable moments through images and text. Powered by artificial intelligence for smart image search, Mori_Cloud delivers a personalized, secure, and modern user experience.
use SAM and OpenCLIP to perform zero-shot object detection using COCO 2017 val split.
Multimodal RAG chatbot with voice input for fashion product search 🤖🎙️👟⚡
CLIP based Zero Shot Instance Segmentation
An AI-powered computer vision system that automatically selects the best wedding photos from thousands of images.
Dynamic cluster-based data sampling for efficient and long-tail-aware vision-language model pre-training.
Local-first semantic image explorer powered by OpenCLIP embeddings and Qdrant vector search, enabling natural-language retrieval across screenshots and photos.
To associate your repository with the openclip topic, visit your repo's landing page and select "manage topics."