Ovis-Omni-Embedding is an omni-modal embedding model developed by the Alibaba ATH-MaaS team. It maps heterogeneous modalities — including text, image, video, and audio — into a unified representation space, enabling comprehensive cross-modal retrieval and understanding within a single model.
The first release, Ovis-Omni-Embedding-v0.1-3B, achieves leading performance on the Massive Multimodal Embedding Benchmark (MMEB).
The technical report will be released in the near future. Stay tuned!
- [26/07/30] 🔥 Announcing Ovis-Omni-Embedding, an omni-modal embedding model for text, image, video, and audio. The technical report of Ovis-Omni-Embedding-v0.1-3B are coming soon.
| Model | Parameters | Supported Modalities | Tech Report |
|---|---|---|---|
| Ovis-Omni-Embedding-v0.1-3B | 3B | Text / Image / Video / Audio | Coming soon |
The technical report is forthcoming. Citation information will be provided upon its release.
This project is licensed under the Apache License, Version 2.0 (SPDX-License-Identifier: Apache-2.0).
We used compliance-checking algorithms during the training process, to ensure the compliance of the trained model to the best of our ability. Due to the complexity of the data and the diversity of language model usage scenarios, we cannot guarantee that the model is completely free of copyright issues or improper content. If you believe anything infringes on your rights or generates improper content, please contact us, and we will promptly address the matter.
