中文
相关论文

相关论文: Sound Source Localization is All about Cross-Modal…

200 篇论文

Multi-modal semantic understanding requires integrating information from different modalities to extract users' real intention behind words. Most previous work applies a dual-encoder structure to separately encode image and text, but fails…

计算与语言 · 计算机科学 2024-03-12 Ming Zhang , Ke Chang , Yunfang Wu

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Hyeonggon Ryu , Seongyu Kim , Joon Son Chung , Arda Senocak

An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, and where someone is…

计算机视觉与模式识别 · 计算机科学 2022-12-26 Rahul Sharma , Krishna Somandepalli , Shrikanth Narayanan

Cross-modal retrieval is generally performed by projecting and aligning the data from two different modalities onto a shared representation space. This shared space often also acts as a bridge for translating the modalities. We address the…

计算机视觉与模式识别 · 计算机科学 2022-03-22 Kranti Kumar Parida , Gaurav Sharma

Sound sources localization using multichannel signal processing has been a subject of active research for decades. In recent years, the use of deep learning in audio signal processing has allowed to drastically improve performances for…

音频与语音处理 · 电气工程与系统科学 2021-06-16 Hadrien Pujol , Éric Bavu , Alexandre Garcia

Video action recognition is a challenging but important task for understanding and discovering what the video does. However, acquiring annotations for a video is costly, and semi-supervised learning (SSL) has been studied to improve…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Seokun Kang , Taehwan Kim

Visual and audio events simultaneously occur and both attract attention. However, most existing saliency prediction works ignore the influence of audio and only consider vision modality. In this paper, we propose a multitask learning method…

计算机视觉与模式识别 · 计算机科学 2021-11-17 Minglang Qiao , Yufan Liu , Mai Xu , Xin Deng , Bing Li , Weiming Hu , Ali Borji

There is extensive interest in metric learning methods for image retrieval. Many metric learning loss functions focus on learning a correct ranking of training samples, but strongly overfit semantically inconsistent labels and require a…

机器学习 · 计算机科学 2023-06-05 Christopher Liao , Theodoros Tsiligkaridis , Brian Kulis

The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data. Recent advances in self-supervised learning have shown that effective representations can be learnt from natural cross-modal…

声音 · 计算机科学 2020-11-05 Soo-Whan Chung , Hong Goo Kang , Joon Son Chung

Imagine hearing a dog bark and turning toward the sound only to see a parked car, while the real, silent dog sits elsewhere. Such sensory conflicts test perception, yet humans reliably resolve them by prioritizing sound over misleading…

声音 · 计算机科学 2025-10-27 Yanhao Jia , Ji Xie , S Jivaganesh , Hao Li , Xu Wu , Mengmi Zhang

Music similarity search is useful for a variety of creative tasks such as replacing one music recording with another recording with a similar "feel", a common task in video editing. For this task, it is typically necessary to define a…

音频与语音处理 · 电气工程与系统科学 2020-08-14 Jongpil Lee , Nicholas J. Bryan , Justin Salamon , Zeyu Jin , Juhan Nam

Music captioning has gained significant attention in the wake of the rising prominence of streaming media platforms. Traditional approaches often prioritize either the audio or lyrics aspect of the music, inadvertently ignoring the…

声音 · 计算机科学 2023-10-24 Zihao He , Weituo Hao , Wei-Tsung Lu , Changyou Chen , Kristina Lerman , Xuchen Song

Cross-modal similarity search is a problem about designing a search system supporting querying across content modalities, e.g., using an image to search for texts or using a text to search for images. This paper presents a compact coding…

计算机视觉与模式识别 · 计算机科学 2019-02-05 Ting Zhang , Jingdong Wang

In the task of audio-visual sound source separation, which leverages visual information for sound source separation, identifying objects in an image is a crucial step prior to separating the sound source. However, existing methods that…

计算机视觉与模式识别 · 计算机科学 2022-03-31 Takashi Oya , Shohei Iwase , Shigeo Morishima

Music exists in various modalities, such as score images, symbolic scores, MIDI, and audio. Translations between each modality are established as core tasks of music information retrieval, such as automatic music transcription…

声音 · 计算机科学 2026-04-08 Jongmin Jung , Dongmin Kim , Sihun Lee , Seola Cho , Hyungjoon Soh , Irmak Bukey , Chris Donahue , Dasaem Jeong

Speaker diarization, the process of segmenting an audio stream or transcribed speech content into homogenous partitions based on speaker identity, plays a crucial role in the interpretation and analysis of human speech. Most existing…

机器学习 · 计算机科学 2024-08-23 Luyao Cheng , Hui Wang , Siqi Zheng , Yafeng Chen , Rongjie Huang , Qinglin Zhang , Qian Chen , Xihao Li

With the exponential surge in diverse multi-modal data, traditional uni-modal retrieval methods struggle to meet the needs of users seeking access to data across various modalities. To address this, cross-modal retrieval has emerged,…

信息检索 · 计算机科学 2024-10-01 Tianshi Wang , Fengling Li , Lei Zhu , Jingjing Li , Zheng Zhang , Heng Tao Shen

Audio-Visual Source Localization (AVSL) aims to localize the source of sound within a video. In this paper, we identify a significant issue in existing benchmarks: the sounding objects are often easily recognized based solely on visual…

多媒体 · 计算机科学 2024-09-12 Liangyu Chen , Zihao Yue , Boshen Xu , Qin Jin

Audio-visual sound source localization (AV-SSL) estimates the position of sound sources by fusing auditory and visual cues. Current AV-SSL methodologies typically require spatially-paired audio-visual data and cannot selectively localize…

声音 · 计算机科学 2025-08-07 Yu Chen , Hongxu Zhu , Jiadong Wang , Kainan Chen , Xinyuan Qian

Music information is often conveyed or recorded across multiple data modalities including but not limited to audio, images, text and scores. However, music information retrieval research has almost exclusively focused on single modality…

声音 · 计算机科学 2021-06-03 Ho-Hsiang Wu , Magdalena Fuentes , Juan P. Bello