English
Related papers

Related papers: FabuLight-ASD: Unveiling Speech Activity via Body …

200 papers

Speaker identification using voice recordings leverages unique acoustic features, but this approach fails when only textual data is available. Few approaches have attempted to tackle the problem of identifying speakers solely from text, and…

Computation and Language · Computer Science 2025-04-22 Rui Ribeiro , Luísa Coheur , Joao P. Carvalho

We introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Sicheng Xu , Guojun Chen , Yu-Xiao Guo , Jiaolong Yang , Chong Li , Zhenyu Zang , Yizhong Zhang , Xin Tong , Baining Guo

The usage of automatic speech recognition (ASR) systems are becoming omnipresent ranging from personal assistant to chatbots, home, and industrial automation systems, etc. Modern robots are also equipped with ASR capabilities for…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-25 Pradip Pramanick , Chayan Sarkar

Voice Activity Detection (VAD) plays a key role in speech processing, often utilizing hand-crafted or neural features. This study examines the effectiveness of Mel-Frequency Cepstral Coefficients (MFCCs) and pre-trained model (PTM)…

Sound · Computer Science 2025-06-03 Kumud Tripathi , Chowdam Venkata Kumar , Pankaj Wasnik

Over the past few years significant progress has been made in the field of presentation attack detection (PAD) for automatic speaker recognition (ASV). This includes the development of new speech corpora, standard evaluation protocols and…

We introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-30 Yafeng Chen , Siqi Zheng , Hui Wang , Luyao Cheng , Tinglong Zhu , Rongjie Huang , Chong Deng , Qian Chen , Shiliang Zhang , Wen Wang , Xihao Li

Understanding camera motion is a fundamental problem in embodied perception and 3D scene understanding. While visual methods have advanced rapidly, they often struggle under visually degraded conditions such as motion blur or occlusions. In…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Daniel Adebi , Sagnik Majumder , Kristen Grauman

This study considers the problem of detecting and locating an active talker's horizontal position from multichannel audio captured by a microphone array. We refer to this as active speaker detection and localization (ASDL). Our goal was to…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-28 Davide Berghi , Philip J. B. Jackson

Speech Activity Detection (SAD), locating speech segments within an audio recording, is a main part of most speech technology applications. Robust SAD is usually more difficult in noisy conditions with varying signal-to-noise ratios (SNR).…

Sound · Computer Science 2021-06-22 Omid Ghahabi , Volker Fischer

Audio Deepfake Detection (ADD) aims to detect spoof speech from bonafide speech. Most prior studies assume that stronger correlations within or across acoustic and emotional features imply authenticity, and thus focus on enhancing or…

Sound · Computer Science 2026-01-21 Jinhua Zhang , Zhenqi Jia , Rui Liu

The intelligent dialogue system, aiming at communicating with humans harmoniously with natural language, is brilliant for promoting the advancement of human-machine interaction in the era of artificial intelligence. With the gradually…

Artificial Intelligence · Computer Science 2022-07-05 Hao Wang , Bin Guo , Yating Zeng , Yasan Ding , Chen Qiu , Ying Zhang , Lina Yao , Zhiwen Yu

In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-09 Shaojin Ding , Quan Wang , Shuo-yiin Chang , Li Wan , Ignacio Lopez Moreno

An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, and where someone is…

Computer Vision and Pattern Recognition · Computer Science 2022-12-26 Rahul Sharma , Krishna Somandepalli , Shrikanth Narayanan

Automated deception detection systems can enhance health, justice, and security in society by helping humans detect deceivers in high-stakes situations across medical and legal domains, among others. This paper presents a novel analysis of…

Computer Vision and Pattern Recognition · Computer Science 2020-10-27 Leena Mathur , Maja J Matarić

Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co-occurring faces are…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-28 Zexu Pan , Shengkui Zhao , Tingting Wang , Kun Zhou , Yukun Ma , Chong Zhang , Bin Ma

Developing machine learning algorithms to understand person-to-person engagement can result in natural user experiences for communal devices such as Amazon Alexa. Among other cues such as voice activity and gaze, a person's audio-visual…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-02 Srinivas Parthasarathy , Shiva Sundaram

Sound event localization and detection (SELD) combines two subtasks: sound event detection (SED) and direction of arrival (DOA) estimation. SELD is usually tackled as an audio-only problem, but visual information has been recently included.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-15 Davide Berghi , Peipei Wu , Jinzheng Zhao , Wenwu Wang , Philip J. B. Jackson

In medical and industrial domains, providing guidance for assembly processes can be critical to ensure efficiency and safety. Errors in assembly can lead to significant consequences such as extended surgery times and prolonged manufacturing…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Hannah Schieber , Shiyu Li , Niklas Corell , Philipp Beckerle , Julian Kreimeier , Daniel Roth

Person identification systems often rely on audio, visual, or behavioral cues, but real-world conditions frequently present with missing or degraded modalities. To address this challenge, we propose a multimodal person identification…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Aref Farhadipour , Teodora Vukovic , Volker Dellwo , Petr Motlicek , Srikanth Madikeri

This paper introduces the Efficient Facial Landmark Detection (EFLD) model, specifically designed for edge devices confronted with the challenges related to power consumption and time latency. EFLD features a lightweight backbone and a…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Ji-Jia Wu