English
Related papers

Related papers: OmniBind: Large-scale Omni Multimodal Representati…

200 papers

Any-to-any generation seeks to translate between arbitrary subsets of modalities, enabling flexible cross-modal synthesis. Despite recent success, existing flow-based approaches are challenged by their inefficiency, as they require…

Machine Learning · Computer Science 2026-04-14 Yeonwoo Cha , Semin Kim , Jinhyeon Kwon , Seunghoon Hong

Lossless compression is essential for efficient data storage and transmission. Although learning-based lossless compressors achieve strong results, most of them are designed for a single modality, leading to redundant compressor deployments…

Machine Learning · Computer Science 2026-03-03 Yan Zhao , Zhengxue Cheng , Junxuan Zhang , Dajiang Zhou , Qunshan Gu , Qi Wang , Li Song

Omni-modal Large Language Models (Omni-LLMs) have demonstrated strong capabilities in audio-video understanding tasks. However, their reliance on long multimodal token sequences leads to substantial computational overhead. Despite this…

Computation and Language · Computer Science 2026-05-14 Yue Ding , Yiyan Ji , Jungang Li , Xuyang Liu , Xinlong Chen , Junfei Wu , Bozhou Li , Bohan Zeng , Yang Shi , Yushuo Guan , Yuanxing Zhang , Jiaheng Liu , Qiang Liu , Pengfei Wan , Liang Wang

Multi-modal learning combines various modalities to provide a comprehensive understanding of real-world problems. A common strategy is to directly bind different modalities together in a specific joint embedding space. However, the…

Machine Learning · Computer Science 2026-02-09 Zhuo Huang , Runnan Chen , Bo Han , Gang Niu , Masashi Sugiyama , Tongliang Liu

The deployment of humanoid robots for dexterous manipulation in unstructured environments remains challenging due to perceptual limitations that constrain the effective workspace. In scenarios where physical constraints prevent the robot…

Robotics · Computer Science 2026-03-09 Pei Qu , Zheng Li , Yufei Jia , Ziyun Liu , Liang Zhu , Haoang Li , Jinni Zhou , Jun Ma

Multimodal models have been proven to outperform text-based models on learning semantic word representations. Almost all previous multimodal models typically treat the representations from different modalities equally. However, it is…

Computation and Language · Computer Science 2018-01-03 Shaonan Wang , Jiajun Zhang , Chengqing Zong

We present TerraBind, a foundation model for protein-ligand structure and binding affinity prediction that achieves 26-fold faster inference than state-of-the-art methods while improving affinity prediction accuracy by $\sim$20\%. Current…

Accurate identification of protein nucleic-acid-binding residues poses a significant challenge with important implications for various biological processes and drug design. Many typical computational methods for protein analysis rely on a…

Biomolecules · Quantitative Biology 2023-12-21 Linglin Jing , Sheng Xu , Yifan Wang , Yuzhe Zhou , Tao Shen , Zhigang Ji , Hui Fang , Zhen Li , Siqi Sun

Representation learning of networks has witnessed significant progress in recent times. Such representations have been effectively used for classic network-based machine learning tasks like node classification, link prediction, and network…

Social and Information Networks · Computer Science 2018-12-07 Arunkumar Bagavathi , Siddharth Krishnan

Aligning objects with corresponding textual descriptions is a fundamental challenge and a realistic requirement in vision-language understanding. While recent multimodal embedding models excel at global image-text alignment, they often…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Shenghao Fu , Yukun Su , Fengyun Rao , Jing Lyu , Xiaohua Xie , Wei-Shi Zheng

We introduce Gemini Embedding 2, a native multimodal embedding model that allows embedding video, audio, image, and text modalities in a unified representation space. We leverage the multimodal capabilities of Gemini to produce embeddings…

In this report, we present OpenUni, a simple, lightweight, and fully open-source baseline for unifying multimodal understanding and generation. Inspired by prevailing practices in unified model learning, we adopt an efficient training…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Size Wu , Zhonghua Wu , Zerui Gong , Qingyi Tao , Sheng Jin , Qinyue Li , Wei Li , Chen Change Loy

The exploration of multimodal language models integrates multiple data types, such as images, text, language, audio, and other heterogeneity. While the latest large language models excel in text-based tasks, they often struggle to…

Artificial Intelligence · Computer Science 2023-11-23 Jiayang Wu , Wensheng Gan , Zefeng Chen , Shicheng Wan , Philip S. Yu

In human-centric scenes, the ability to simultaneously understand visual and auditory information is crucial. While recent omni models can process multiple modalities, they generally lack effectiveness in human-centric scenes due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-28 Jiaxing Zhao , Qize Yang , Yixing Peng , Detao Bai , Shimin Yao , Boyuan Sun , Xiang Chen , Shenghao Fu , Weixuan chen , Xihan Wei , Liefeng Bo

Multimodal large models have been recognized for their advantages in various performance and downstream tasks. The development of these models is crucial towards achieving general artificial intelligence in the future. In this paper, we…

Sound · Computer Science 2023-09-12 Sen Fang , Bowen Gao , Yangjian Wu , Teik Toe Teoh

This paper presents OmniDataComposer, an innovative approach for multimodal data fusion and unlimited data generation with an intent to refine and uncomplicate interplay among diverse data modalities. Coming to the core breakthrough, it…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Dongyang Yu , Shihao Wang , Yuan Fang , Wangpeng An

Multimodal foundation models that can holistically process text alongside images, video, audio, and other sensory modalities are increasingly used in a variety of real-world applications. However, it is challenging to characterize and study…

Heterogeneous information networks(HINs) become popular in recent years for its strong capability of modelling objects with abundant information using explicit network structure. Network embedding has been proved as an effective method to…

Machine Learning · Computer Science 2021-04-12 Xinyi Zhang , Lihui Chen

For natural language understanding and generation, embedding concepts using an order-based representation is an essential task. Unlike traditional point vector based representation, an order-based representation imposes geometric…

Computation and Language · Computer Science 2024-04-18 Croix Gyurek , Niloy Talukder , Mohammad Al Hasan

The scarcity of high-quality multimodal biomedical data limits the ability to effectively fine-tune pretrained Large Language Models (LLMs) for specialized biomedical tasks. To address this challenge, we introduce MINT (Multimodal…

Quantitative Methods · Quantitative Biology 2026-02-18 Zhanliang Wang , Da Wu , Quan Nguyen , Zhuoran Xu , Kai Wang