中文
相关论文

相关论文: Mutual Learning for Acoustic Matching and Dereverb…

200 篇论文

With the advance in self-supervised learning for audio and visual modalities, it has become possible to learn a robust audio-visual speech representation. This would be beneficial for improving the audio-visual speech recognition (AVSR)…

图像与视频处理 · 电气工程与系统科学 2022-07-12 Zi-Qiang Zhang , Jie Zhang , Jian-Shu Zhang , Ming-Hui Wu , Xin Fang , Li-Rong Dai

We present AdVerb, a novel audio-visual dereverberation framework that uses visual cues in addition to the reverberant sound to estimate clean audio. Although audio-only dereverberation is a well-studied problem, our approach incorporates…

计算机视觉与模式识别 · 计算机科学 2023-08-25 Sanjoy Chowdhury , Sreyan Ghosh , Subhrajyoti Dasgupta , Anton Ratnarajah , Utkarsh Tyagi , Dinesh Manocha

The objective of this work is to extract target speaker's voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their performance with promising intelligibility, but maintaining…

音频与语音处理 · 电气工程与系统科学 2023-10-31 Suyeon Lee , Chaeyoung Jung , Youngjoon Jang , Jaehun Kim , Joon Son Chung

Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality only. Our idea is to…

声音 · 计算机科学 2023-03-15 Changan Chen , Wei Sun , David Harwath , Kristen Grauman

In audio-visual navigation (AVN) tasks, an embodied agent must autonomously localize a sound source in unknown and complex 3D environments based on audio-visual signals. Existing methods often rely on static modality fusion strategies and…

人工智能 · 计算机科学 2025-09-23 Jia Li , Yinfeng Yu , Liejun Wang , Fuchun Sun , Wendong Zheng

Learning-based Multi-View Stereo (MVS) methods warp source images into the reference camera frustum to form 3D volumes, which are fused as a cost volume to be regularized by subsequent networks. The fusing step plays a vital role in…

计算机视觉与模式识别 · 计算机科学 2022-04-18 Xiaofeng Wang , Zheng Zhu , Fangbo Qin , Yun Ye , Guan Huang , Xu Chi , Yijia He , Xingang Wang

Creating novel images by fusing visual cues from multiple sources is a fundamental yet underexplored problem in image-to-image generation, with broad applications in artistic creation, virtual reality and visual media. Existing methods…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Zeren Xiong , Yue Yu , Zedong Zhang , Shuo Chen , Jian Yang , Jun Li

An immersive acoustic experience enabled by spatial audio is just as crucial as the visual aspect in creating realistic virtual environments. However, existing methods for room impulse response estimation rely either on data-demanding…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Derong Jin , Ruohan Gao

To reconstruct the 3D geometry from calibrated images, learning-based multi-view stereo (MVS) methods typically perform multi-view depth estimation and then fuse depth maps into a mesh or point cloud. To improve the computational…

计算机视觉与模式识别 · 计算机科学 2025-09-19 Fangjinhua Wang , Qingshan Xu , Yew-Soon Ong , Marc Pollefeys

Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speakers, background noise and room reverberation remains a…

声音 · 计算机科学 2022-04-11 Guinan Li , Jianwei Yu , Jiajun Deng , Xunying Liu , Helen Meng

Speaker verification (SV) has recently attracted considerable research interest due to the growing popularity of virtual assistants. At the same time, there is an increasing requirement for an SV system: it should be robust to short speech…

音频与语音处理 · 电气工程与系统科学 2020-10-07 Youngmoon Jung , Yeunju Choi , Hyungjun Lim , Hoirin Kim

Visual question answering and visual dialogue tasks have been increasingly studied in the multimodal field towards more practical real-world scenarios. A more challenging task, audio visual scene-aware dialogue (AVSD), is proposed to…

计算与语言 · 计算机科学 2019-08-15 Yi-Ting Yeh , Tzu-Chuan Lin , Hsiao-Hua Cheng , Yu-Hsuan Deng , Shang-Yu Su , Yun-Nung Chen

Zero-shot learning enables models to generalise to unseen classes by leveraging semantic information, bridging the gap between training and testing sets with non-overlapping classes. While much research has focused on zero-shot learning in…

声音 · 计算机科学 2025-07-03 Ysobel Sims , Alexandre Mendes , Stephan Chalup

Speech enhancement (SE) aims to improve the quality and intelligibility of speech in noisy environments. Recent studies have shown that incorporating visual cues in audio signal processing can enhance SE performance. Given that human speech…

声音 · 计算机科学 2025-05-27 Meng-Ping Lin , Jen-Cheng Hou , Chia-Wei Chen , Shao-Yi Chien , Jun-Cheng Chen , Xugang Lu , Yu Tsao

Audio Visual Scene-aware Dialog (AVSD) is the task of generating a response for a question with a given scene, video, audio, and the history of previous turns in the dialog. Existing systems for this task employ the transformers or…

计算与语言 · 计算机科学 2020-04-20 Hwanhee Lee , Seunghyun Yoon , Franck Dernoncourt , Doo Soon Kim , Trung Bui , Kyomin Jung

We propose the first joint audio-video generation framework that brings engaging watching and listening experiences simultaneously, towards high-quality realistic videos. To generate joint audio-video pairs, we propose a novel Multi-Modal…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Ludan Ruan , Yiyang Ma , Huan Yang , Huiguo He , Bei Liu , Jianlong Fu , Nicholas Jing Yuan , Qin Jin , Baining Guo

This paper proposes a new unsupervised audio-visual speech enhancement (AVSE) approach that combines a diffusion-based audio-visual speech generative model with a non-negative matrix factorization (NMF) noise model. First, the diffusion…

声音 · 计算机科学 2025-01-16 Jean-Eudes Ayilo , Mostafa Sadeghi , Romain Serizel , Xavier Alameda-Pineda

Audio-visual speech enhancement (AVSE) is a task that uses visual auxiliary information to extract a target speaker's speech from mixed audio. In real-world scenarios, there often exist complex acoustic environments, accompanied by various…

声音 · 计算机科学 2025-11-03 Jiarong Du , Zhan Jin , Peijun Yang , Juan Liu , Zhuo Li , Xin Liu , Ming Li

Video-to-audio (V2A) generation aims to synthesize realistic and semantically aligned audio from silent videos, with potential applications in video editing, Foley sound design, and assistive multimedia. Although the excellent results,…

Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription…

音频与语音处理 · 电气工程与系统科学 2025-01-09 Xinyu Wang , Haotian Jiang , Haolin Huang , Yu Fang , Mengjie Xu , Qian Wang
‹ 上一页 1 2 3 10 下一页 ›