中文
相关论文

相关论文: Look Who's Talking: Active Speaker Detection in th…

200 篇论文

Social interactions play a crucial role in shaping human behavior, relationships, and societies. It encompasses various forms of communication, such as verbal conversation, non-verbal gestures, facial expressions, and body language. In this…

机器学习 · 计算机科学 2026-05-13 Alice Zhang , Callihan Bertley , Dawei Liang , Edison Thomaz

Target-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence Target-Speaker Voice Activity Detection (Seq2Seq-TSVAD)…

音频与语音处理 · 电气工程与系统科学 2023-02-21 Ming Cheng , Weiqing Wang , Yucong Zhang , Xiaoyi Qin , Ming Li

Speaker diarization(SD) is a classic task in speech processing and is crucial in multi-party scenarios such as meetings and conversations. Current mainstream speaker diarization approaches consider acoustic information only, which result in…

计算与语言 · 计算机科学 2023-05-23 Luyao Cheng , Siqi Zheng , Zhang Qinglin , Hui Wang , Yafeng Chen , Qian Chen

Speaker-independent VSR is a complex task that involves identifying spoken words or phrases from video recordings of a speaker's facial movements. Over the years, there has been a considerable amount of research in the field of VSR…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Praneeth Nemani , G. Sai Krishna , Supriya Kundrapu

Spoofing detection for automatic speaker verification (ASV), which is to discriminate between live speech and attacks, has received increasing attentions recently. However, all the previous studies have been done on the clean data without…

机器学习 · 计算机科学 2016-02-10 Xiaohai Tian , Zhizheng Wu , Xiong Xiao , Eng Siong Chng , Haizhou Li

Silent speech interfaces have been recently proposed as a way to enable communication when the acoustic signal is not available. This introduces the need to build visual speech recognition systems for silent and whispered speech. However,…

计算机视觉与模式识别 · 计算机科学 2018-02-20 Stavros Petridis , Jie Shen , Doruk Cetin , Maja Pantic

In this article, we introduce a novel problem of audio-visual autism behavior recognition, which includes social behavior recognition, an essential aspect previously omitted in AI-assisted autism screening research. We define the task at…

Speech is considered as a multi-modal process where hearing and vision are two fundamentals pillars. In fact, several studies have demonstrated that the robustness of Automatic Speech Recognition systems can be improved when audio and…

计算机视觉与模式识别 · 计算机科学 2023-11-22 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

Developing tools to automatically detect check-worthy claims in political debates and speeches can greatly help moderators of debates, journalists, and fact-checkers. While previous work on this problem has focused exclusively on the text…

计算与语言 · 计算机科学 2024-01-19 Petar Ivanov , Ivan Koychev , Momchil Hardalov , Preslav Nakov

Auditory attention decoding (AAD) is a technique used to identify and amplify the talker that a listener is focused on in a noisy environment. This is done by comparing the listener's brainwaves to a representation of all the sound sources…

音频与语音处理 · 电气工程与系统科学 2023-02-14 Cong Han , Vishal Choudhari , Yinghao Aaron Li , Nima Mesgarani

Active Speaker Detection (ASD) aims to identify who is speaking in complex visual scenes. While humans naturally rely on lip-audio synchronization, existing ASD models often misclassify non-speaking instances when lip movements and audio…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Le Thien Phuc Nguyen , Zhuoran Yu , Yong Jae Lee

The goal of this work is to investigate the performance of popular speaker recognition models on speech segments from movies, where often actors intentionally disguise their voice to play a character. We make the following three…

声音 · 计算机科学 2021-02-12 Andrew Brown , Jaesung Huh , Arsha Nagrani , Joon Son Chung , Andrew Zisserman

Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2018-06-20 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Predicting when to initiate speech in real-world environments remains a fundamental challenge for conversational agents. We introduce EgoSpeak, a novel framework for real-time speech initiation prediction in egocentric streaming video. By…

计算机视觉与模式识别 · 计算机科学 2025-02-24 Junhyeok Kim , Min Soo Kim , Jiwan Chung , Jungbin Cho , Jisoo Kim , Sungwoong Kim , Gyeongbo Sim , Youngjae Yu

Robots are becoming everyday devices, increasing their interaction with humans. To make human-machine interaction more natural, cognitive features like Visual Voice Activity Detection (VVAD), which can detect whether a person is speaking or…

计算机视觉与模式识别 · 计算机科学 2024-01-03 Adrian Lubitz , Matias Valdenegro-Toro , Frank Kirchner

Voice Activity Detection (VAD) is the process of automatically determining whether a person is speaking and identifying the timing of their speech in an audiovisual data. Traditionally, this task has been tackled by processing either audio…

计算机视觉与模式识别 · 计算机科学 2024-10-21 Andrea Appiani , Cigdem Beyan

We introduce a distinctive real-time, causal, neural network-based active speaker detection system optimized for low-power edge computing. This system drives a virtual cinematography module and is deployed on a commercial device. The system…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Ilya Gurvich , Ido Leichter , Dharmendar Reddy Palle , Yossi Asher , Alon Vinnikov , Igor Abramovski , Vishak Gopal , Ross Cutler , Eyal Krupka

Visual speech recognition (VSR) aims to recognize the content of speech based on lip movements, without relying on the audio stream. Advances in deep learning and the availability of large audio-visual datasets have led to the development…

计算机视觉与模式识别 · 计算机科学 2022-11-01 Pingchuan Ma , Stavros Petridis , Maja Pantic

In this paper, we present a system that associates faces with voices in a video by fusing information from the audio and visual signals. The thesis underlying our work is that an extremely simple approach to generating (weak) speech…

多媒体 · 计算机科学 2017-06-02 Ken Hoover , Sourish Chaudhuri , Caroline Pantofaru , Malcolm Slaney , Ian Sturdy

Recognizing who is speaking in a crowded scene is a key challenge towards the understanding of the social interactions going on within. Detecting speaking status from body movement alone opens the door for the analysis of social scenes in…

计算机视觉与模式识别 · 计算机科学 2022-11-02 Jose Vargas-Quiros , Laura Cabrera-Quiros , Hayley Hung