中文
相关论文

相关论文: Learning Audio-Visual Dynamics Using Scene Graphs …

200 篇论文

In audio-visual navigation, an agent intelligently travels through a complex, unmapped 3D environment using both sights and sounds to find a sound source (e.g., a phone ringing in another room). Existing models learn to act at a fixed…

计算机视觉与模式识别 · 计算机科学 2021-02-12 Changan Chen , Sagnik Majumder , Ziad Al-Halah , Ruohan Gao , Santhosh Kumar Ramakrishnan , Kristen Grauman

We propose a self-supervised approach for learning to perform audio source separation in videos based on natural language queries, using only unlabeled video and audio pairs as training data. A key challenge in this task is learning to…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Reuben Tan , Arijit Ray , Andrea Burns , Bryan A. Plummer , Justin Salamon , Oriol Nieto , Bryan Russell , Kate Saenko

Dynamic objects in the environment, such as people and other agents, lead to challenges for existing simultaneous localization and mapping (SLAM) approaches. To deal with dynamic environments, computer vision researchers usually apply some…

机器人学 · 计算机科学 2021-08-04 Tianwei Zhang , Huayan Zhang , Xiaofei Li , Junfeng Chen , Tin Lun Lam , Sethu Vijayakumar

While neural network approaches have made significant strides in resolving classical signal processing problems, it is often the case that hybrid approaches that draw insight from both signal processing and neural networks produce more…

音频与语音处理 · 电气工程与系统科学 2024-02-13 Karim Helwani , Masahito Togami , Paris Smaragdis , Michael M. Goodwin

Few-shot audio-visual acoustics modeling seeks to synthesize the room impulse response in arbitrary locations with few-shot observations. To sufficiently exploit the provided few-shot data for accurate acoustic modeling, we present a…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Diwei Huang , Kunyang Lin , Peihao Chen , Qing Du , Mingkui Tan

We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise. Solving this task using only audio as input is extremely challenging and does not provide an…

Source separation is the process of isolating individual sounds in an auditory mixture of multiple sounds [1], and has a variety of applications ranging from speech enhancement and lyric transcription [2] to digital audio production for…

声音 · 计算机科学 2024-12-10 Bradford Derby , Lucas Dunker , Samarth Galchar , Shashank Jarmale , Akash Setti

Modern smart glasses leverage advanced audio sensing and machine learning technologies to offer real-time transcribing and captioning services, considerably enriching human experiences in daily communications. However, such systems…

Never having seen an object and heard its sound simultaneously, can the model still accurately localize its visual position from the input audio? In this work, we concentrate on the Audio-Visual Localization and Segmentation tasks but under…

计算机视觉与模式识别 · 计算机科学 2024-02-05 Yaoting Wang , Weisong Liu , Guangyao Li , Jian Ding , Di Hu , Xi Li

Acoustic Scene Classification (ASC) and Sound Event Detection (SED) are two separate tasks in the field of computational sound scene analysis. In this work, we present a new dataset with both sound scene and sound event labels and use this…

音频与语音处理 · 电气工程与系统科学 2019-07-02 Helen L. Bear , Ines Nolasco , Emmanouil Benetos

In this paper, we investigate how to learn rich and robust feature representations for audio classification from visual data and acoustic images, a novel audio data modality. Former models learn audio representations from raw signals or…

计算机视觉与模式识别 · 计算机科学 2020-02-12 Andrés F. Pérez , Valentina Sanguineti , Pietro Morerio , Vittorio Murino

Multimodal dialogue emotion recognition captures emotional cues by fusing text, visual, and audio modalities. However, existing approaches still suffer from notable limitations in modeling emotional dependencies and learning multimodal…

多媒体 · 计算机科学 2026-03-12 Yunsheng Wang , Yuntao Shou , Yilong Tan , Wei Ai , Tao Meng , Keqin Li

The thud of a bouncing ball, the onset of speech as lips open -- when visual and audio events occur together, it suggests that there might be a common, underlying event that produced both signals. In this paper, we argue that the visual and…

计算机视觉与模式识别 · 计算机科学 2018-10-10 Andrew Owens , Alexei A. Efros

We introduce a state-of-the-art audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify limitations of previous…

声音 · 计算机科学 2021-10-15 Efthymios Tzinis , Scott Wisdom , Tal Remez , John R. Hershey

We present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual speech separation is a…

Propelled by the breakthrough in deep generative models, audio-to-image generation has emerged as a pivotal cross-modal task that converts complex auditory signals into rich visual representations. However, previous works only focus on…

声音 · 计算机科学 2025-12-11 Hao Zhou , Xiaobao Guo , Yuzhe Zhu , Adams Wai-Kin Kong

Can machines recording an audio-visual scene produce realistic, matching audio-visual experiences at novel positions and novel view directions? We answer it by studying a new task -- real-world audio-visual scene synthesis -- and a…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Susan Liang , Chao Huang , Yapeng Tian , Anurag Kumar , Chenliang Xu

Automatic speech recognition (ASR) in multimedia content is one of the promising applications, but speech data in this kind of content are frequently mixed with background music, which is harmful for the performance of ASR. In this study,…

音频与语音处理 · 电气工程与系统科学 2020-08-28 Jeongwoo Woo , Masato Mimura , Kazuyoshi Yoshii , Tatsuya Kawahara

3D scene graphs have empowered robots with semantic understanding for navigation and planning. However, current functional scene graphs primarily focus on static element detection, lacking the actionable kinematic information required for…

机器人学 · 计算机科学 2026-03-24 Qiuyi Gu , Yuze Sheng , Jincheng Yu , Jiahao Tang , Xiaolong Shan , Zhaoyang Shen , Tinghao Yi , Xiaodan Liang , Xinlei Chen , Yu Wang

In this paper we propose a method for separation of moving sound sources. The method is based on first tracking the sources and then estimation of source spectrograms using multichannel non-negative matrix factorization (NMF) and extracting…

声音 · 计算机科学 2017-10-30 Joonas Nikunen , Aleksandr Diment , Tuomas Virtanen