中文
相关论文

相关论文: PianoBind: A Multimodal Joint Embedding Model for …

200 篇论文

Multimodal representation alignment is pivotal for large language models and robotics. Traditional methods are often hindered by cross-modal information discrepancies and data scarcity, leading to suboptimal alignment spaces that overlook…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Zeyu Chen , Jie Li , Kai Han

Understanding neural activity and information representation is crucial for advancing knowledge of brain function and cognition. Neural activity, measured through techniques like electrophysiology and neuroimaging, reflects various aspects…

神经元与认知 · 定量生物学 2024-07-22 Fengyu Yang , Chao Feng , Daniel Wang , Tianye Wang , Ziyao Zeng , Zhiyang Xu , Hyoungseob Park , Pengliang Ji , Hanbin Zhao , Yuanning Li , Alex Wong

We present ImageBind, an approach to learn a joint embedding across six different modalities - images, text, audio, depth, thermal, and IMU data. We show that all combinations of paired data are not necessary to train such a joint…

计算机视觉与模式识别 · 计算机科学 2023-06-01 Rohit Girdhar , Alaaeldin El-Nouby , Zhuang Liu , Mannat Singh , Kalyan Vasudev Alwala , Armand Joulin , Ishan Misra

We present UniBind, a flexible and efficient approach that learns a unified representation space for seven diverse modalities -- images, text, audio, point cloud, thermal, video, and event data. Existing works, eg., ImageBind, treat the…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Yuanhuiyi Lyu , Xu Zheng , Jiazhou Zhou , Lin Wang

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Given the foundational role of multimodal joint representation in understanding and generation pipelines, high-quality…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Zehan Wang , Ziang Zhang , Hang Zhang , Luping Liu , Rongjie Huang , Xize Cheng , Hengshuang Zhao , Zhou Zhao

Unified multi-model representation spaces are the foundation of multimodal understanding and generation. However, the billions of model parameters and catastrophic forgetting problems make it challenging to further enhance pre-trained…

计算机视觉与模式识别 · 计算机科学 2024-05-13 Zehan Wang , Ziang Zhang , Xize Cheng , Rongjie Huang , Luping Liu , Zhenhui Ye , Haifeng Huang , Yang Zhao , Tao Jin , Peng Gao , Zhou Zhao

Modeling various aspects that make a music piece unique is a challenging task, requiring the combination of multiple sources of information. Deep learning is commonly used to obtain representations using various sources of information, such…

声音 · 计算机科学 2021-04-05 Andres Ferraro , Xavier Favory , Konstantinos Drossos , Yuntae Kim , Dmitry Bogdanov

We present TaxaBind, a unified embedding space for characterizing any species of interest. TaxaBind is a multimodal embedding space across six modalities: ground-level images of species, geographic location, satellite image, text, audio,…

计算机视觉与模式识别 · 计算机科学 2024-11-04 Srikumar Sastry , Subash Khanal , Aayush Dhakal , Adeel Ahmad , Nathan Jacobs

A flexible recommendation and retrieval system requires music similarity in terms of multiple partial elements of musical pieces to allow users to select the element they want to focus on. A method for music similarity learning using…

声音 · 计算机科学 2025-07-18 Yuka Hashizume , Li Li , Atsushi Miyashita , Tomoki Toda

Multimodal sensing systems are increasingly prevalent in various real-world applications. Most existing multimodal learning approaches heavily rely on training with a large amount of synchronized, complete multimodal data. However, such a…

机器学习 · 计算机科学 2025-03-06 Xiaomin Ouyang , Jason Wu , Tomoyoshi Kimura , Yihan Lin , Gunjan Verma , Tarek Abdelzaher , Mani Srivastava

Style transfer of polyphonic music recordings is a challenging task when considering the modeling of diverse, imaginative, and reasonable music pieces in the style different from their original one. To achieve this, learning stable…

声音 · 计算机科学 2018-11-30 Chien-Yu Lu , Min-Xin Xue , Chia-Che Chang , Che-Rung Lee , Li Su

Learning musical structures and composition patterns is necessary for both music generation and understanding, but current methods do not make uniform use of learned features to generate and comprehend music simultaneously. In this paper,…

声音 · 计算机科学 2024-12-10 Xiao Liang , Zijian Zhao , Weichao Zeng , Yutong He , Fupeng He , Yiyi Wang , Chengying Gao

A unified representation space in multi-modal learning is essential for effectively integrating diverse data sources, such as text, images, and audio, to enhance efficiency and performance across various downstream tasks. Recent binding…

机器学习 · 计算机科学 2025-10-08 Minoh Jeong , Zae Myung Kim , Min Namgung , Dongyeop Kang , Yao-Yi Chiang , Alfred Hero

Research on multi-modal learning dominantly aligns the modalities in a unified space at training, and only a single one is taken for prediction at inference. However, for a real machine, e.g., a robot, sensors could be added or removed at…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Yuanhuiyi Lyu , Xu Zheng , Dahun Kim , Lin Wang

We simplify space binding by focusing on two core components, a single encoder per modality and high-quality data; enabling training state-of-the-art models on a single GPU in a few hours as opposed to multiple days. We present EBind, an…

机器学习 · 计算机科学 2025-11-19 Jim Broadbent , Felix Cohen , Frederik Hvilshøj , Eric Landau , Eren Sasoglu

We have seen remarkable success in representation learning and language models (LMs) using deep neural networks. Many studies aim to build the underlying connections among different modalities via the alignment and mappings at the token or…

声音 · 计算机科学 2025-03-04 Daniel Chin , Gus Xia

Self-supervised representation learning maps high-dimensional data into a meaningful embedding space, where samples of similar semantic contents are close to each other. Most of the recent representation learning methods maximize cosine…

计算机视觉与模式识别 · 计算机科学 2022-06-15 Chuang Niu , Ge Wang

Semantic embeddings play a crucial role in natural language-based information retrieval. Embedding models represent words and contexts as vectors whose spatial configuration is derived from the distribution of words in large text corpora.…

计算与语言 · 计算机科学 2024-01-09 Silvan David Peter , Shreyan Chowdhury , Carlos Eduardo Cancino-Chacón , Gerhard Widmer

Piano cover generation aims to create a piano cover from a pop song. Existing approaches mainly employ supervised learning and the training demands strongly-aligned and paired song-to-piano data, which is built by remapping piano notes to…

声音 · 计算机科学 2024-08-06 Chih-Pin Tan , Hsin Ai , Yi-Hsin Chang , Shuen-Huei Guan , Yi-Hsuan Yang

Definitive embeddings remain a fundamental challenge of computational musicology for symbolic music in deep learning today. Analogous to natural language, music can be modeled as a sequence of tokens. This motivates the majority of existing…

声音 · 计算机科学 2020-10-19 Hongru Liang , Wenqiang Lei , Paul Yaozhu Chan , Zhenglu Yang , Maosong Sun , Tat-Seng Chua
‹ 上一页 1 2 3 10 下一页 ›