Omni-Embed-Mini: 0.9B model embeds text and five more modalities

A new paper presents Omni-Embed-Mini, a 0.9B-parameter embedding model that maps text, speech, audio, images, video and visually-rich documents into a single shared cosine space. It does this without updating any text-side parameter.
The authors start from a familiar problem. Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. Omni-Embed-Mini is built to avoid both the quality loss and the size.
The key insight, in the authors' words, is that the teacher signal requires no separate embedding model. Each media sample is paired with a dense cascaded caption, and the teacher target is simply the frozen backbone's own embedding of that caption. Because teacher and student share the same backbone weights, they inhabit byte-identical geometry. That means lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment.
Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner, whose negatives sharpen as the encoder improves.
The 0.9B model keeps its text weights bit-identical to the backbone, so training cannot regress text retrieval. Its text score is 49.57 nDCG@10 on MTEB-v2 BEIR-8. At the same time it extends the model to five additional modalities, and it is about 2.7x to 9.5x smaller than every open omni embedder the authors compare against.
The recipe carries over to a 2.3B variant, made by swapping in a native vision-language backbone. The authors report that this variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average. Models, code, data and the evaluation harness are on the project page at omniembed.cvmbzuai.com.
Key facts
- Omni-Embed-Mini is a 0.9B-parameter model that maps text, speech, audio, images, video and visually-rich documents into one shared cosine space without updating any text-side parameter.
- The teacher target is the frozen backbone's own embedding of a dense cascaded caption for each media sample, so no separate embedding model is needed; lightweight projectors and phased LoRA adapters handle alignment.
- Text weights stay bit-identical to the backbone, so text retrieval cannot regress; the reported score is 49.57 nDCG@10 on MTEB-v2 BEIR-8.
- The 0.9B model is about 2.7x to 9.5x smaller than every open omni embedder the authors compare against.
- A 2.3B variant, built by swapping in a native vision-language backbone, is reported as competitive with the closed gemini-embedding-2 and edging ahead on the overall-modality average.
Why it matters
Adding new modalities to a text embedder usually costs text retrieval quality, and the existing omni-modal embedders cope by growing to multi-billion parameters. Omni-Embed-Mini takes a different route: it freezes the text side entirely, so text retrieval cannot get worse by construction, and it aligns the other modalities to the backbone's own embeddings of captions. The authors present this as a way to get one shared space for six kinds of content at under a billion parameters.
Who it affects
Teams building retrieval or search over mixed content will care most: the model covers text, speech, audio, images, video and visually-rich documents in one cosine space. People who already depend on a text embedding backbone and do not want its text behaviour to shift are the natural audience, since the text weights are unchanged. Researchers working on multimodal embedding can also use the released data and evaluation harness.
How to use it
According to the authors, models, code, data and the evaluation harness are available on the project page at omniembed.cvmbzuai.com. The 0.9B model is the small option; the 2.3B variant uses a native vision-language backbone. The licence of the released material is not stated in the source, so check the project page before building on it.
How solid is it
This is the authors' own report in a paper abstract, and the headline comparisons are their claims. The 49.57 nDCG@10 figure on MTEB-v2 BEIR-8 is stated for text retrieval; the claim that text does not regress rests on the weights being bit-identical to the backbone. The size advantage of about 2.7x to 9.5x is measured against the open omni embedders the authors chose to compare with. The statement that the 2.3B variant edges ahead of gemini-embedding-2 refers to the overall-modality average, and the source gives no scores for it or for the non-text modalities. Models, code and data being released means others can check the results.
Risks and caveats
The text score is preserved by design, but the source gives no per-modality numbers for speech, audio, images, video or documents, so how well the new modalities perform is not visible here. The 2.3B result is described as competitive, with only a slim edge on one average, against a closed model. The names and sizes of the compared open embedders, the backbone model, and the training data and compute are not given in the source. Licence terms for the released models and data are also not stated.
“Because teacher and student share the same backbone weights, they inhabit byte-identical geometry, and lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment.”
— Omni-Embed-Mini paper abstract