中文
相关论文

相关论文: SoundSculpt: Direction and Semantics Driven Ambiso…

200 篇论文

We present a single-stage casual waveform-to-waveform multichannel model that can separate moving sound sources based on their broad spatial locations in a dynamic acoustic scene. We divide the scene into two spatial regions containing,…

声音 · 计算机科学 2022-07-01 Dejan Markovic , Alexandre Defossez , Alexander Richard

Accurately estimating and simulating the physical properties of objects from real-world sound recordings is of great practical importance in the fields of vision, graphics, and robotics. However, the progress in these directions has been…

声音 · 计算机科学 2024-09-23 Xutong Jin , Chenxi Xu , Ruohan Gao , Jiajun Wu , Guoping Wang , Sheng Li

The audio-visual segmentation (AVS) task aims to segment sounding objects from a given video. Existing works mainly focus on fusing audio and visual features of a given video to achieve sounding object masks. However, we observed that prior…

声音 · 计算机科学 2023-08-02 Chen Liu , Peike Li , Xingqun Qi , Hu Zhang , Lincheng Li , Dadong Wang , Xin Yu

Cinematic audio source separation is a relatively new subtask of audio source separation, with the aim of extracting the dialogue, music, and effects stems from their mixture. In this work, we developed a model generalizing the Bandsplit…

音频与语音处理 · 电气工程与系统科学 2024-08-27 Karn N. Watcharasupat , Chih-Wei Wu , Yiwei Ding , Iroro Orife , Aaron J. Hipple , Phillip A. Williams , Scott Kramer , Alexander Lerch , William Wolcott

Target audio source separation with natural language queries presents a promising paradigm for extracting arbitrary audio events through arbitrary text descriptions. Existing methods mainly face two challenges, the difficulty in jointly…

声音 · 计算机科学 2025-12-03 Xinlei Yin , Xiulian Peng , Xue Jiang , Zhiwei Xiong , Yan Lu

Topic modeling is a powerful technique to discover hidden topics and patterns within a collection of documents without prior knowledge. Traditional topic modeling and clustering-based techniques encounter challenges in capturing contextual…

计算与语言 · 计算机科学 2024-10-04 Melkamu Abay Mersha , Mesay Gemeda yigezu , Jugal Kalita

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work…

Acoustic source localization has been applied in different fields, such as aeronautics and ocean science, generally using multiple microphones array data to reconstruct the source location. However, the model-based beamforming methods fail…

声音 · 计算机科学 2022-04-01 Guanxing Zhou , Hao Liang , Xinghao Ding , Yue Huang , Xiaotong Tu , Saqlain Abbas

Numerous embedding models have been recently explored to incorporate semantic knowledge into visual recognition. Existing methods typically focus on minimizing the distance between the corresponding images and texts in the embedding space…

计算机视觉与模式识别 · 计算机科学 2017-06-06 Dong Li , Hsin-Ying Lee , Jia-Bin Huang , Shengjin Wang , Ming-Hsuan Yang

Recent advances in active noise control have enabled the development of hearables with spatial selectivity, which actively suppress undesired noise while preserving desired sound from specific directions. In this work, we propose an…

音频与语音处理 · 电气工程与系统科学 2025-05-16 Tong Xiao , Simon Doclo

Spatial audio reasoning enables machines to interpret auditory scenes by understanding events and their spatial attributes. In this work, we focus on spatial audio understanding with an emphasis on reasoning about moving sources. First, we…

声音 · 计算机科学 2025-09-19 Arvind Krishna Sridhar , Yinyi Guo , Erik Visser

Large-scale pre-trained image-text models exhibit robust multimodal representations, yet applying the Contrastive Language-Image Pre-training (CLIP) model to audio-visual localization remains challenging. Replacing the classification token…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Khanh Binh Nguyen , Chae Jung Park

Intent classification is a task in spoken language understanding. An intent classification system is usually implemented as a pipeline process, with a speech recognition module followed by text processing that classifies the intents. There…

计算与语言 · 计算机科学 2021-02-16 Bidisha Sharma , Maulik Madhavi , Haizhou Li

We propose a novel deep neural network architecture for speech recognition that explicitly employs knowledge of the background environmental noise within a deep neural network acoustic model. A deep neural network is used to predict the…

计算与语言 · 计算机科学 2016-10-03 Suyoun Kim , Bhiksha Raj , Ian Lane

In conversational speech, the acoustic signal provides cues that help listeners disambiguate difficult parses. For automatically parsing spoken utterances, we introduce a model that integrates transcribed text and acoustic-prosodic features…

计算与语言 · 计算机科学 2018-04-17 Trang Tran , Shubham Toshniwal , Mohit Bansal , Kevin Gimpel , Karen Livescu , Mari Ostendorf

Employing pre-trained language models (LM) to extract contextualized word representations has achieved state-of-the-art performance on various NLP tasks. However, applying this technique to noisy transcripts generated by automatic speech…

计算与语言 · 计算机科学 2020-11-03 Chao-Wei Huang , Yun-Nung Chen

Our objective is to transform a video into a set of discrete audio-visual objects using self-supervised learning. To this end, we introduce a model that uses attention to localize and group sound sources, and optical flow to aggregate…

计算机视觉与模式识别 · 计算机科学 2020-08-11 Triantafyllos Afouras , Andrew Owens , Joon Son Chung , Andrew Zisserman

Reasoning about spatial audio with large language models requires a spatial audio encoder as an acoustic front-end to obtain audio embeddings for further processing. Such an encoder needs to capture all information required to detect the…

音频与语音处理 · 电气工程与系统科学 2025-11-04 Kevin Wilkinghoff , Zheng-Hua Tan

The rapid advances in audio analysis underscore its vast potential for humancomputer interaction, environmental monitoring, and public safety; yet, existing audioonly datasets often lack spatial context. To address this gap, we present two…

声音 · 计算机科学 2025-12-10 Shuaihang Yuan , Congcong Wen , Muhammad Shafique , Anthony Tzes , Yi Fang

Target sound extraction consists of extracting the sound of a target acoustic event (AE) class from a mixture of AE sounds. It can be realized using a neural network that extracts the target sound conditioned on a 1-hot vector that…

音频与语音处理 · 电气工程与系统科学 2021-06-15 Marc Delcroix , Jorge Bennasar Vázquez , Tsubasa Ochiai , Keisuke Kinoshita , Shoko Araki