Jina's jina-embeddings-v5-omni-small Searches Images and Audio Without Re-Embedding Text

Jina ships a 1.74B multimodal embedding model that encodes text, images, video, and audio into one shared vector space without reindexing.

·
·
·
Jina's jina-embeddings-v5-omni-small Searches Images and Audio Without Re-Embedding TextPRO
Read2 min
TypeModel
  • Jina released jina-embeddings-v5-omni-small, a 1.74B multimodal embedder for text, image, video, audio
  • Text embeddings are bit-identical to v5-text-small, so no index rebuild is needed
  • Uses GELATO frozen-tower design, training only 0.35% of weights on 4 H100s
  • 1024-dim output with Matryoshka truncation to 32-768 and 32k context length
  • Ships four task adapters: retrieval, classification, clustering, text-matching, plus vLLM support
  • License is CC BY-NC 4.0; commercial use requires contacting Jina

Jina Adds Multimodal Search Without Re-Embedding Text

Jina has released jina-embeddings-v5-omni-small, a 1.74-billion-parameter model that maps text, images, video, and audio into one vector space. Items with related meaning land near one another, enabling cross-modal nearest-neighbor search. Its text embeddings are bit-for-bit identical to those from the corresponding v5 Text model when configuration and task settings match, so existing text vectors can remain in place while teams add media queries and documents.

Compatibility depends on using the corresponding v5 Text variant, the same task adapter, the same query or document role, and the same output dimension. Under those conditions, a RAG system built on v5 Text can search its current index with an image, audio clip, or video without re-embedding the stored text corpus.

Four media types, one geometry

Specification Value
Parameters 1.74 billion
Modalities Text, images, video frames, and audio
Native vector size 1,024 dimensions
Text context 32,768 tokens
Truncated sizes 32, 64, 128, 256, 512, or 768 dimensions
License CC BY-NC 4.0

The model accepts formats including .mp4, .wav, .mp3, .pdf, .jpg, and .png. Matryoshka training allows applications to retain only the leading dimensions of each vector, reducing storage and search costs at the expense of some retrieval quality.

Choose the objective

  • retrieval: asymmetric query-to-document search and RAG

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads