中文
相关论文

相关论文: RespVAD: Voice Activity Detection via Video-Extrac…

200 篇论文

In TV services, dialogue level personalization is key to meeting user preferences and needs. When dialogue and background sounds are not separately available from the production stage, Dialogue Separation (DS) can estimate them to enable…

音频与语音处理 · 电气工程与系统科学 2023-03-24 Matteo Torcoli , Emanuël A. P. Habets

Video anomaly detection (VAD) aims to detect anomalies that deviate from what is expected. In open-world scenarios, the expected events may change as requirements change. For example, not wearing a mask may be considered abnormal during a…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Zihao Liu , Xiaoyu Wu , Jianqin Wu , Xuxu Wang , Linlin Yang

Video Anomaly Detection (VAD) serves as a pivotal technology in the intelligent surveillance systems, enabling the temporal or spatial identification of anomalous events within videos. While existing reviews predominantly concentrate on…

计算机视觉与模式识别 · 计算机科学 2024-02-02 Yang Liu , Dingkang Yang , Yan Wang , Jing Liu , Jun Liu , Azzedine Boukerche , Peng Sun , Liang Song

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current…

多媒体 · 计算机科学 2025-05-29 Yong Ren , Chenxing Li , Le Xu , Hao Gu , Duzhen Zhang , Yujie Chen , Manjie Xu , Ruibo Fu , Shan Yang , Dong Yu

Detecting anchor's voice in live musical streams is an important preprocessing for music and speech signal processing. Existing approaches to voice activity detection (VAD) primarily rely on audio, however, audio-based VAD is difficult to…

声音 · 计算机科学 2020-11-03 Yuanbo Hou , Yi Deng , Bilei Zhu , Zejun Ma , Dick Botteldooren

Rapid advances in speech synthesis and audio editing have made realistic forgeries increasingly accessible, yet existing detection methods remain vulnerable to tampering or depend on visual/wearable sensors. In this paper, we present…

人机交互 · 计算机科学 2026-03-31 Mingda Han , Huanqi Yang , Chaoqun Li , Wenhao Li , Guoming Zhang , Yanni Yang , Yetong Cao , Weitao Xu , Pengfei Hu

Deviations in respiratory rate often precede abnormalities in other vital signs. However, continuously monitoring respiratory rates outside clinical settings remains challenging due to the obtrusive nature and sensitivity to body motions in…

信号处理 · 电气工程与系统科学 2024-11-15 Sebastian Reidy , Manuel Meier , Christian Holz

It is common in everyday spoken communication that we look at the turning head of a talker to listen to his/her voice. Humans see the talker to listen better, so do machines. However, previous studies on audio-visual speaker extraction have…

声音 · 计算机科学 2023-09-14 Qinghua Liu , Meng Ge , Zhizheng Wu , Haizhou Li

Under noisy conditions, automatic speech recognition (ASR) can greatly benefit from the addition of visual signals coming from a video of the speaker's face. However, when multiple candidate speakers are visible this traditionally requires…

音频与语音处理 · 电气工程与系统科学 2022-05-12 Otavio Braga , Olivier Siohan

Breathing is an essential part of human survival, which carries information about a person's physiological and psychological state. Generally, breath boundaries are marked by experts before using for any task. An unsupervised algorithm for…

音频与语音处理 · 电气工程与系统科学 2023-04-10 Shivani Yadav , Dipanjan Gope , Uma Maheswari K. , Prasanta Kumar Ghosh

The integration of visual cues has revitalized the performance of the target speech extraction task, elevating it to the forefront of the field. Nevertheless, this multi-modal learning paradigm often encounters the challenge of modality…

声音 · 计算机科学 2024-05-07 Zhaoxi Mu , Xinyu Yang

This paper delves into the challenging task of Active Speaker Detection (ASD), where the system needs to determine in real-time whether a person is speaking or not in a series of video frames. While previous works have made significant…

计算机视觉与模式识别 · 计算机科学 2024-09-16 Arnav Kundu , Yanzi Jin , Mohammad Sekhavat , Max Horton , Danny Tormoen , Devang Naik

When video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise. While most existing methods use audio-only inputs, improved performance is…

计算机视觉与模式识别 · 计算机科学 2018-06-14 Aviv Gabbay , Asaph Shamir , Shmuel Peleg

Video action detection (VAD) aims to detect actors and classify their actions in a video. We figure that VAD suffers more from classification rather than localization of actors. Hence, we analyze how prevailing methods form features for…

计算机视觉与模式识别 · 计算机科学 2024-09-12 Jinsung Lee , Taeoh Kim , Inwoong Lee , Minho Shim , Dongyoon Wee , Minsu Cho , Suha Kwak

This paper describes an audio-visual speech enhancement (AV-SE) method that estimates from noisy input audio a mixture of the speech of the speaker appearing in an input video (on-screen target speech) and of a selected speaker not…

音频与语音处理 · 电气工程与系统科学 2023-06-13 Tomoya Yoshinaga , Keitaro Tanaka , Shigeo Morishima

In recent years, considerable progress has been made in the non-contact based detection of the respiration rate from video sequences. Common techniques either directly assess the movement of the chest due to breathing or are based on…

图像与视频处理 · 电气工程与系统科学 2019-06-20 Fabian Schrumpf , Christoph Moench , Gerold Bausch , Mirco Fuchs

Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the remote user experience of a video conference by automatically panning, tilting and zooming of a video conferencing camera: users subjectively rate…

音频与语音处理 · 电气工程与系统科学 2022-05-26 Ross Cutler , Ramin Mehran , Sam Johnson , Cha Zhang , Adam Kirk , Oliver Whyte , Adarsh Kowdle

Singing voice detection (SVD), to recognize vocal parts in the song, is an essential task in music information retrieval (MIR). The task remains challenging since singing voice varies and intertwines with the accompaniment music, especially…

音频与语音处理 · 电气工程与系统科学 2022-05-09 Yifu Sun , Xulong Zhang , Yi Yu , Xi Chen , Wei Li

Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario where the target speech is highly overlapped with the…

声音 · 计算机科学 2023-09-18 Junjie Li , Ruijie Tao , Zexu Pan , Meng Ge , Shuai Wang , Haizhou Li

Speech separation aims to separate individual voice from an audio mixture of multiple simultaneous talkers. Although audio-only approaches achieve satisfactory performance, they build on a strategy to handle the predefined conditions,…

声音 · 计算机科学 2020-12-01 Peng Zhang , Jiaming Xu , Jing shi , Yunzhe Hao , Bo Xu