中文
相关论文

相关论文: BIAS: A Body-based Interpretable Active Speaker Ap…

200 篇论文

Autism spectrum disorder (ASD) is a neurodevelopmental condition characterized by challenges in social communication, repetitive behavior, and sensory processing. One important research area in ASD is evaluating children's behavioral…

声音 · 计算机科学 2025-06-03 Tiantian Feng , Anfeng Xu , Xuan Shi , Somer Bishop , Shrikanth Narayanan

The strong relation between face and voice can aid active speaker detection systems when faces are visible, even in difficult settings, when the face of a speaker is not clear or when there are several people in the same scene. By being…

机器学习 · 计算机科学 2021-09-07 Hugo Carneiro , Cornelius Weber , Stefan Wermter

Automatic Speech Understanding (ASU) aims at human-like speech interpretation, providing nuanced intent, emotion, sentiment, and content understanding from speech and language (text) content conveyed in speech. Typically, training a robust…

声音 · 计算机科学 2024-04-30 Tiantian Feng , Xuan Shi , Rahul Gupta , Shrikanth S. Narayanan

Audio-visual active speaker detection (AVASD) is well-developed, and now is an indispensable front-end for several multi-modal applications. However, to the best of our knowledge, the adversarial robustness of AVASD models hasn't been…

声音 · 计算机科学 2022-10-04 Xuanjun Chen , Haibin Wu , Helen Meng , Hung-yi Lee , Jyh-Shing Roger Jang

The Facial Action Coding System (FACS) has been used by numerous studies to investigate the links between facial behavior and mental health. The laborious and costly process of FACS coding has motivated the development of machine learning…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Evangelos Sariyanidi , Lisa Yankowitz , Robert T. Schultz , John D. Herrington , Birkan Tunc , Jeffrey Cohn

Many speech enhancement methods try to learn the relationship between noisy and clean speech, obtained using an acoustic room simulator. We point out several limitations of enhancement methods relying on clean speech targets; the goal of…

计算与语言 · 计算机科学 2018-12-26 Geonmin Kim , Hwaran Lee , Bo-Kyeong Kim , Sang-Hoon Oh , Soo-Young Lee

It is common in everyday spoken communication that we look at the turning head of a talker to listen to his/her voice. Humans see the talker to listen better, so do machines. However, previous studies on audio-visual speaker extraction have…

声音 · 计算机科学 2023-09-14 Qinghua Liu , Meng Ge , Zhizheng Wu , Haizhou Li

Audio-visual speaker diarization aims at detecting "who spoke when" using both auditory and visual signals. Existing audio-visual diarization datasets are mainly focused on indoor environments like meeting rooms or news studios, which are…

计算机视觉与模式识别 · 计算机科学 2022-07-19 Eric Zhongcong Xu , Zeyang Song , Satoshi Tsutsui , Chao Feng , Mang Ye , Mike Zheng Shou

Compared with automatic speech recognition (ASR), the human auditory system is more adept at handling noise-adverse situations, including environmental noise and channel distortion. To mimic this adeptness, auditory models have been widely…

计算与语言 · 计算机科学 2016-09-16 Peng Dai , Xue Teng , Frank Rudzicz , Ing Yann Soon

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

计算与语言 · 计算机科学 2025-04-11 Lakshmipathi Balaji , Karan Singla

Audio-visual automatic speech recognition (AV-ASR) is an extension of ASR that incorporates visual cues, often from the movements of a speaker's mouth. Unlike works that simply focus on the lip motion, we investigate the contribution of…

计算机视觉与模式识别 · 计算机科学 2022-06-16 Valentin Gabeur , Paul Hongsuck Seo , Arsha Nagrani , Chen Sun , Karteek Alahari , Cordelia Schmid

We present BIAS, a fast, biologically inspired model for dynamic visual saliency detection in continuous video streams. Building on the Itti--Koch framework, BIAS incorporates a retina-inspired motion detector to extract temporal features,…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Zhao-ji Zhang , Ya-tang Li

Voice activity detection (VAD) is an essential pre-processing step for tasks such as automatic speech recognition (ASR) and speaker recognition. A basic goal is to remove silent segments within an audio, while a more general VAD system…

音频与语音处理 · 电气工程与系统科学 2020-09-22 Yefei Chen , Shuai Wang , Yanmin Qian , Kai Yu

Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed architecture has been…

声音 · 计算机科学 2022-07-08 Junwen Xiong , Yu Zhou , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha

Non-native speakers (NNSs) often face speaking challenges in real-time multilingual communication, such as struggling to articulate their thoughts. To address this issue, we developed an AI-based speaking assistant (AISA) that provides…

人机交互 · 计算机科学 2025-05-06 Peinuan Qin , Zicheng Zhu , Naomi Yamashita , Yitian Yang , Keita Suga , Yi-Chieh Lee

Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, and human-robot interaction. The absence of a large,…

In challenging real-life conditions such as extreme head-pose, occlusions, and low-resolution images where the visual information fails to estimate visual attention/gaze direction, audio signals could provide important and complementary…

计算机视觉与模式识别 · 计算机科学 2022-08-15 Shreya Ghosh , Abhinav Dhall , Munawar Hayat , Jarrod Knibbe

We propose a novel voice activity detection (VAD) model in a low-resource environment. Our key idea is to model VAD as a denoising task, and construct a network that is designed to identify nuisance features for a speech classification…

音频与语音处理 · 电气工程与系统科学 2022-10-31 Jonathan Svirsky , Ofir Lindenbaum

The goal of this paper is to learn strong lip reading models that can recognise speech in silent videos. Most prior works deal with the open-set visual speech recognition problem by adapting existing automatic speech recognition techniques…

计算机视觉与模式识别 · 计算机科学 2021-12-06 K R Prajwal , Triantafyllos Afouras , Andrew Zisserman

We introduce GaussianSpeech, a novel approach that synthesizes high-fidelity animation sequences of photo-realistic, personalized 3D human head avatars from spoken audio. To capture the expressive, detailed nature of human heads, including…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Shivangi Aneja , Artem Sevastopolsky , Tobias Kirschstein , Justus Thies , Angela Dai , Matthias Nießner