English

jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval

Artificial Intelligence 2025-07-08 v3 Computation and Language Information Retrieval

Abstract

We introduce jina-embeddings-v4, a 3.8 billion parameter multimodal embedding model that unifies text and image representations through a novel architecture supporting both single-vector and multi-vector embeddings in the late interaction style. The model incorporates task-specific Low-Rank Adaptation (LoRA) adapters to optimize performance across diverse retrieval scenarios, including query-document retrieval, semantic text similarity, and code search. Comprehensive evaluations demonstrate that jina-embeddings-v4 achieves state-of-the-art performance on both single-modal and cross-modal retrieval tasks, with particular strength in processing visually rich content such as tables, charts, diagrams, and mixed-media formats. To facilitate evaluation of this capability, we also introduce Jina-VDR, a novel benchmark specifically designed for visually rich image retrieval.

Keywords

Cite

@article{arxiv.2506.18902,
  title  = {jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval},
  author = {Michael Günther and Saba Sturua and Mohammad Kalim Akram and Isabelle Mohr and Andrei Ungureanu and Bo Wang and Sedigheh Eslami and Scott Martens and Maximilian Werk and Nan Wang and Han Xiao},
  journal= {arXiv preprint arXiv:2506.18902},
  year   = {2025}
}

Comments

22 pages, 1-10 main, 14-22 experimental results, benchmark tables

R2 v1 2026-07-01T03:29:57.144Z