中文
相关论文

相关论文: Self-Supervised Audio-Visual Soundscape Stylizatio…

200 篇论文

We demonstrate how conditional generation from diffusion models can be used to tackle a variety of realistic tasks in the production of music in 44.1kHz stereo audio with sampling-time guidance. The scenarios we consider include…

声音 · 计算机科学 2023-12-06 Mark Levy , Bruno Di Giorgi , Floris Weers , Angelos Katharopoulos , Tom Nickson

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the…

音频与语音处理 · 电气工程与系统科学 2022-05-13 Otavio Braga , Olivier Siohan

In this paper, we teach machines to understand visuals and natural language by learning the mapping between sentences and noisy video snippets without explicit annotations. Firstly, we define a self-supervised learning framework that…

计算机视觉与模式识别 · 计算机科学 2021-01-12 Yujie Zhong , Linhai Xie , Sen Wang , Lucia Specia , Yishu Miao

The large amount of audiovisual content being shared online today has drawn substantial attention to the prospect of audiovisual self-supervised learning. Recent works have focused on each of these modalities separately, while others have…

机器学习 · 计算机科学 2021-06-18 Pingchuan Ma , Rodrigo Mira , Stavros Petridis , Björn W. Schuller , Maja Pantic

We present an unsupervised approach that converts the input speech of any individual into audiovisual streams of potentially-infinitely many output speakers. Our approach builds on simple autoencoders that project out-of-sample data onto…

计算机视觉与模式识别 · 计算机科学 2021-07-06 Kangle Deng , Aayush Bansal , Deva Ramanan

In speech recognition, it is essential to model the phonetic content of the input signal while discarding irrelevant factors such as speaker variations and noise, which is challenging in low-resource settings. Self-supervised pre-training…

计算与语言 · 计算机科学 2023-01-04 Sreepratha Ram , Hanan Aldarmaki

Audio-driven visual scene editing endeavors to manipulate the visual background while leaving the foreground content unchanged, according to the given audio signals. Unlike current efforts focusing primarily on image editing, audio-driven…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Kaixin Shen , Ruijie Quan , Linchao Zhu , Jun Xiao , Yi Yang

Speech editing systems aim to naturally modify speech content while preserving acoustic consistency and speaker identity. However, previous studies often struggle to adapt to unseen and diverse acoustic conditions, resulting in degraded…

音频与语音处理 · 电气工程与系统科学 2025-11-04 Taewoo Kim , Uijong Lee , Hayoung Park , Choongsang Cho , Nam In Park , Young Han Lee

With the advance in self-supervised learning for audio and visual modalities, it has become possible to learn a robust audio-visual speech representation. This would be beneficial for improving the audio-visual speech recognition (AVSR)…

图像与视频处理 · 电气工程与系统科学 2022-07-12 Zi-Qiang Zhang , Jie Zhang , Jian-Shu Zhang , Ming-Hui Wu , Xin Fang , Li-Rong Dai

Audio-visual pre-trained models have gained substantial attention recently and demonstrated superior performance on various audio-visual tasks. This study investigates whether pre-trained audio-visual models demonstrate non-arbitrary…

计算与语言 · 计算机科学 2024-11-13 Wei-Cheng Tseng , Yi-Jen Shih , David Harwath , Raymond Mooney

We investigate unsupervised learning of correspondences between sound events and textual phrases through aligning audio clips with textual captions describing the content of a whole audio clip. We align originally unaligned and unannotated…

音频与语音处理 · 电气工程与系统科学 2022-02-22 Huang Xie , Okko Räsänen , Konstantinos Drossos , Tuomas Virtanen

The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt…

音频与语音处理 · 电气工程与系统科学 2023-10-17 Dimitrios Bralios , Gordon Wichern , François G. Germain , Zexu Pan , Sameer Khurana , Chiori Hori , Jonathan Le Roux

Audio captioning aims to generate text descriptions of audio clips. In the real world, many objects produce similar sounds. How to accurately recognize ambiguous sounds is a major challenge for audio captioning. In this work, inspired by…

音频与语音处理 · 电气工程与系统科学 2023-05-30 Xubo Liu , Qiushi Huang , Xinhao Mei , Haohe Liu , Qiuqiang Kong , Jianyuan Sun , Shengchen Li , Tom Ko , Yu Zhang , Lilian H. Tang , Mark D. Plumbley , Volkan Kılıç , Wenwu Wang

In this paper, we explore the use of pre-trained language models to learn sentiment information of written texts for speech sentiment analysis. First, we investigate how useful a pre-trained language model would be in a 2-step pipeline…

计算与语言 · 计算机科学 2021-06-15 Suwon Shon , Pablo Brusco , Jing Pan , Kyu J. Han , Shinji Watanabe

Autonomous soundscape augmentation systems typically use trained models to pick optimal maskers to effect a desired perceptual change. While acoustic information is paramount to such systems, contextual information, including participant…

声音 · 计算机科学 2024-07-03 Kenneth Ooi , Karn N. Watcharasupat , Bhan Lam , Zhen-Ting Ong , Woon-Seng Gan

We present a novel approach to multilingual audio-visual speech recognition tasks by introducing a single model on a multilingual dataset. Motivated by a human cognitive system where humans can intuitively distinguish different languages…

多媒体 · 计算机科学 2023-10-24 Joanna Hong , Se Jin Park , Yong Man Ro

The goal of voice conversion is to transform the speech of a source speaker to sound like that of a reference speaker while preserving the original content. A key challenge is to extract disentangled linguistic content from the source and…

声音 · 计算机科学 2025-01-15 Jaehun Kim , Ji-Hoon Kim , Yeunju Choi , Tan Dat Nguyen , Seongkyu Mun , Joon Son Chung

In this paper, we show how to use audio to supervise the learning of active speaker detection in video. Voice Activity Detection (VAD) guides the learning of the vision-based classifier in a weakly supervised manner. The classifier uses…

计算机视觉与模式识别 · 计算机科学 2016-03-30 Punarjay Chakravarty , Tinne Tuytelaars

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

声音 · 计算机科学 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

Although 360\textdegree{} cameras ease the capture of panoramic footage, it remains challenging to add realistic 360\textdegree{} audio that blends into the captured scene and is synchronized with the camera motion. We present a method for…

图形学 · 计算机科学 2018-05-15 Dingzeyu Li , Timothy R. Langlois , Changxi Zheng