English
Related papers

Related papers: Gemini Embedding 2: A Native Multimodal Embedding …

200 papers

Graph embedding techniques, which learn low-dimensional representations of a graph, are achieving state-of-the-art performance in many graph mining tasks. Most existing embedding algorithms assign a single vector to each node, implicitly…

Social and Information Networks · Computer Science 2020-10-22 Jisung Yoon , Kai-Cheng Yang , Woo-Sung Jung , Yong-Yeol Ahn

We introduce the Granite Embedding R2 models, a comprehensive family of high-performance English encoder-based embedding models engineered for enterprise-scale dense retrieval applications. Building upon our first-generation release, these…

Contrastive language-image pre-training aligns the features of text-image pairs in a common latent space via distinct encoders for each modality. While this approach achieves impressive performance in several zero-shot tasks, it cannot…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Christian Schlarmann , Francesco Croce , Nicolas Flammarion , Matthias Hein

Multimodal tasks, such as image-text retrieval and generation, require embedding data from diverse modalities into a shared representation space. Aligning embeddings from heterogeneous sources while preserving shared and modality-specific…

Machine Learning · Computer Science 2024-12-03 Dongfang Zhao

Recent efforts on training visual navigation agents conditioned on language using deep reinforcement learning have been successful in learning policies for different multimodal tasks, such as semantic goal navigation and embodied question…

Machine Learning · Computer Science 2019-02-05 Devendra Singh Chaplot , Lisa Lee , Ruslan Salakhutdinov , Devi Parikh , Dhruv Batra

Massively multilingual sentence representation models, e.g., LASER, SBERT-distill, and LaBSE, help significantly improve cross-lingual downstream tasks. However, the use of a large amount of data or inefficient model architectures results…

Computation and Language · Computer Science 2024-05-31 Zhuoyuan Mao , Chenhui Chu , Sadao Kurohashi

Designing powerful tools that support cooking activities has rapidly gained popularity due to the massive amounts of available data, as well as recent advances in machine learning that are capable of analyzing them. In this paper, we…

Computation and Language · Computer Science 2018-05-01 Micael Carvalho , Rémi Cadène , David Picard , Laure Soulier , Nicolas Thome , Matthieu Cord

Embeddings mapping high-dimensional discrete input to lower-dimensional continuous vector spaces have been widely adopted in machine learning applications as a way to capture domain semantics. Interviewing 13 embedding users across…

Human-Computer Interaction · Computer Science 2022-03-07 Angie Boggust , Brandon Carter , Arvind Satyanarayan

We present Liquid, an auto-regressive generation paradigm that seamlessly integrates visual comprehension and generation by tokenizing images into discrete codes and learning these code embeddings alongside text tokens within a shared…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Junfeng Wu , Yi Jiang , Chuofan Ma , Yuliang Liu , Hengshuang Zhao , Zehuan Yuan , Song Bai , Xiang Bai

The realization of Artificial General Intelligence (AGI) necessitates Embodied AI agents capable of robust spatial perception, effective task planning, and adaptive execution in physical environments. However, current large language models…

Biological multimodal large language models (MLLMs) have emerged as powerful foundation models for scientific discovery. However, existing models are specialized to a single modality, limiting their ability to solve inherently cross-modal…

Machine Learning · Computer Science 2026-03-17 Wonbin Lee , Dongki Kim , Sung Ju Hwang

We propose a novel probabilistic model for visual question answering (Visual QA). The key idea is to infer two sets of embeddings: one for the image and the question jointly and the other for the answers. The learning objective is to learn…

Computer Vision and Pattern Recognition · Computer Science 2018-06-12 Hexiang Hu , Wei-Lun Chao , Fei Sha

Recent breakthroughs in large multimodal models (LMMs), such as the impressive GPT-4o-Native, have demonstrated remarkable proficiency in following general-purpose instructions for image generation. However, current benchmarks often lack…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Jiayu Wang , Yang Jiao , Yue Yu , Tianwen Qian , Shaoxiang Chen , Jingjing Chen , Yu-Gang Jiang

General-purpose robots need a deep understanding of the physical world, advanced reasoning, and general and dexterous control. This report introduces the latest generation of the Gemini Robotics model family: Gemini Robotics 1.5, a…

Robotics · Computer Science 2025-12-02 Gemini Robotics Team , Abbas Abdolmaleki , Saminda Abeyruwan , Joshua Ainslie , Jean-Baptiste Alayrac , Montserrat Gonzalez Arenas , Ashwin Balakrishna , Nathan Batchelor , Alex Bewley , Jeff Bingham , Michael Bloesch , Konstantinos Bousmalis , Philemon Brakel , Anthony Brohan , Thomas Buschmann , Arunkumar Byravan , Serkan Cabi , Ken Caluwaerts , Federico Casarini , Christine Chan , Oscar Chang , London Chappellet-Volpini , Jose Enrique Chen , Xi Chen , Hao-Tien Lewis Chiang , Krzysztof Choromanski , Adrian Collister , David B. D'Ambrosio , Sudeep Dasari , Todor Davchev , Meet Kirankumar Dave , Coline Devin , Norman Di Palo , Tianli Ding , Carl Doersch , Adil Dostmohamed , Yilun Du , Debidatta Dwibedi , Sathish Thoppay Egambaram , Michael Elabd , Tom Erez , Xiaolin Fang , Claudio Fantacci , Cody Fong , Erik Frey , Chuyuan Fu , Ruiqi Gao , Marissa Giustina , Keerthana Gopalakrishnan , Laura Graesser , Oliver Groth , Agrim Gupta , Roland Hafner , Steven Hansen , Leonard Hasenclever , Sam Haves , Nicolas Heess , Brandon Hernaez , Alex Hofer , Jasmine Hsu , Lu Huang , Sandy H. Huang , Atil Iscen , Mithun George Jacob , Deepali Jain , Sally Jesmonth , Abhishek Jindal , Ryan Julian , Dmitry Kalashnikov , M. Emre Karagozler , Stefani Karp , Matija Kecman , J. Chase Kew , Donnie Kim , Frank Kim , Junkyung Kim , Thomas Kipf , Sean Kirmani , Ksenia Konyushkova , Li Yang Ku , Yuheng Kuang , Thomas Lampe , Antoine Laurens , Tuan Anh Le , Isabel Leal , Alex X. Lee , Tsang-Wei Edward Lee , Guy Lever , Jacky Liang , Li-Heng Lin , Fangchen Liu , Shangbang Long , Caden Lu , Sharath Maddineni , Anirudha Majumdar , Kevis-Kokitsi Maninis , Andrew Marmon , Sergio Martinez , Assaf Hurwitz Michaely , Niko Milonopoulos , Joss Moore , Robert Moreno , Michael Neunert , Francesco Nori , Joy Ortiz , Kenneth Oslund , Carolina Parada , Emilio Parisotto , Amaris Paryag , Acorn Pooley , Thomas Power , Alessio Quaglino , Haroon Qureshi , Rajkumar Vasudeva Raju , Helen Ran , Dushyant Rao , Kanishka Rao , Isaac Reid , David Rendleman , Krista Reymann , Miguel Rivas , Francesco Romano , Yulia Rubanova , Peter Pastor Sampedro , Pannag R Sanketi , Dhruv Shah , Mohit Sharma , Kathryn Shea , Mohit Shridhar , Charles Shu , Vikas Sindhwani , Sumeet Singh , Radu Soricut , Rachel Sterneck , Ian Storz , Razvan Surdulescu , Jie Tan , Jonathan Tompson , Saran Tunyasuvunakool , Jake Varley , Grace Vesom , Giulia Vezzani , Maria Bauza Villalonga , Oriol Vinyals , René Wagner , Ayzaan Wahid , Stefan Welker , Paul Wohlhart , Chengda Wu , Markus Wulfmeier , Fei Xia , Ted Xiao , Annie Xie , Jinyu Xie , Peng Xu , Sichun Xu , Ying Xu , Zhuo Xu , Jimmy Yan , Sherry Yang , Skye Yang , Yuxiang Yang , Hiu Hong Yu , Wenhao Yu , Wentao Yuan , Yuan Yuan , Jingwei Zhang , Tingnan Zhang , Zhiyuan Zhang , Allan Zhou , Guangyao Zhou , Yuxiang Zhou

Recent advancements in contrastive learning have revolutionized self-supervised representation learning and achieved state-of-the-art performance on benchmark tasks. While most existing methods focus on applying contrastive learning to…

Machine Learning · Computer Science 2024-04-16 Lihui Liu , Jinha Kim , Vidit Bansal

Multimodal foundation models have shown compelling but conflicting performance in medical image interpretation. However, the mechanisms by which these models integrate and prioritize different data modalities, including images and text,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Thomas Buckley , James A. Diao , Pranav Rajpurkar , Adam Rodman , Arjun K. Manrai

We introduce jina-embeddings-v3, a novel text embedding model with 570 million parameters, achieves state-of-the-art performance on multilingual data and long-context retrieval tasks, supporting context lengths of up to 8192 tokens. The…

Large Language models (LLMs) have demonstrated impressive performance on a wide range of tasks, including in multimodal settings such as speech. However, their evaluation is often limited to English and a few high-resource languages. For…

Recent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for multi-modal pretraining, which can benefit from modality…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Haiyang Xu , Qinghao Ye , Ming Yan , Yaya Shi , Jiabo Ye , Yuanhong Xu , Chenliang Li , Bin Bi , Qi Qian , Wei Wang , Guohai Xu , Ji Zhang , Songfang Huang , Fei Huang , Jingren Zhou

We develop an approach to learning visual representations that embraces multimodal data, driven by a combination of intra- and inter-modal similarity preservation objectives. Unlike existing visual pre-training methods, which solve a proxy…

Computer Vision and Pattern Recognition · Computer Science 2021-04-28 Xin Yuan , Zhe Lin , Jason Kuen , Jianming Zhang , Yilin Wang , Michael Maire , Ajinkya Kale , Baldo Faieta