English
Related papers

Related papers: Comparative Analysis of Personalized Voice Activit…

200 papers

State-of-the-art Active Speaker Detection (ASD) approaches heavily rely on audio and facial features to perform, which is not a sustainable approach in wild scenarios. Although these methods achieve good results in the standard…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Tiago Roxo , Joana C. Costa , Pedro R. M. Inácio , Hugo Proença

Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality only. Our idea is to…

Sound · Computer Science 2023-03-15 Changan Chen , Wei Sun , David Harwath , Kristen Grauman

Video Anomaly Detection (VAD) has emerged as a pivotal task in computer vision, with broad relevance across multiple fields. Recent advances in deep learning have driven significant progress in this area, yet the field remains fragmented…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Ghazal Alinezhad Noghre , Armin Danesh Pazho , Hamed Tabkhi

Keyword spotting (KWS) and speaker verification (SV) have been studied independently although it is known that acoustic and speaker domains are complementary. In this paper, we propose a multi-task network that performs KWS and SV…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-10 Myunghun Jung , Youngmoon Jung , Jahyun Goo , Hoirin Kim

Own voice pickup technology for hearable devices facilitates communication in noisy environments. Own voice reconstruction (OVR) systems enhance the quality and intelligibility of the recorded noisy own voice signals. Since disturbances…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-04 Mattes Ohlenbusch , Christian Rollwage , Simon Doclo , Jan Rennies

Visual Anomaly Detection (VAD) has gained significant research attention for its ability to identify anomalous images and pinpoint the specific areas responsible for the anomaly. A key advantage of VAD is its unsupervised nature, which…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Manuel Barusco , Francesco Borsatti , Davide Dalle Pezze , Francesco Paissan , Elisabetta Farella , Gian Antonio Susto

Social Anxiety Disorder (SAD) significantly impacts individuals' daily lives and relationships. The conventional methods for SAD detection involve physical consultations and self-reported questionnaires, but they have limitations such as…

Computers and Society · Computer Science 2025-01-13 Nilesh Kumar Sahu , Nandigramam Sai Harshit , Rishabh Uikey , Haroon R. Lone

Research has shown that trust is an essential aspect of human-computer interaction directly determining the degree to which the person is willing to use the system. An automatic prediction of the level of trust that a user has on a certain…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-03 Leonardo Pepino , Pablo Riera , Lara Gauder , Agustín Gravano , Luciana Ferrer

To better model the contextual information and increase the generalization ability of Speech Activity Detection (SAD) system, this paper leverages a multi-lingual Automatic Speech Recognition (ASR) system to perform SAD. Sequence…

Sound · Computer Science 2021-04-13 Seyyed Saeed Sarfjoo , Srikanth Madikeri , Petr Motlicek

Video anomaly detection (VAD) holds immense importance across diverse domains such as surveillance, healthcare, and environmental monitoring. While numerous surveys focus on conventional VAD methods, they often lack depth in exploring…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Moshira Abdalla , Sajid Javed , Muaz Al Radi , Anwaar Ulhaq , Naoufel Werghi

Speaker diarization, the process of identifying "who spoke when" in audio recordings, is essential for understanding classroom dynamics. However, classroom settings present distinct challenges, including poor recording quality, high levels…

Sound · Computer Science 2025-05-28 Ali Sartaz Khan , Tolulope Ogunremi , Ahmed Adel Attia , Dorottya Demszky

Voice has become an increasingly popular User Interaction (UI) channel, mainly contributing to the ongoing trend of wearables, smart vehicles, and home automation systems. Voice assistants such as Siri, Google Now and Cortana, have become…

Cryptography and Security · Computer Science 2017-01-18 Huan Feng , Kassem Fawaz , Kang G. Shin

Performing an adequate evaluation of sound event detection (SED) systems is far from trivial and is still subject to ongoing research. The recently proposed polyphonic sound detection (PSD)-receiver operating characteristic (ROC) and PSD…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-01 Janek Ebbers , Romain Serizel , Reinhold Haeb-Umbach

This paper addresses the issue of active speaker detection (ASD) in noisy environments and formulates a robust active speaker detection (rASD) problem. Existing ASD approaches leverage both audio and visual modalities, but non-speech sounds…

Multimedia · Computer Science 2024-04-02 Siva Sai Nagender Vasireddy , Chenxu Zhang , Xiaohu Guo , Yapeng Tian

Speaker verification (SV) has recently attracted considerable research interest due to the growing popularity of virtual assistants. At the same time, there is an increasing requirement for an SV system: it should be robust to short speech…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-07 Youngmoon Jung , Yeunju Choi , Hyungjun Lim , Hoirin Kim

For face presentation attack detection (PAD), most of the spoofing cues are subtle, local image patterns (e.g., local image distortion, 3D mask edge and cut photo edges). The representations of existing PAD works with simple global pooling…

Computer Vision and Pattern Recognition · Computer Science 2022-02-22 Jiong Wang , Zhou Zhao , Weike Jin , Xinyu Duan , Zhen Lei , Baoxing Huai , Yiling Wu , Xiaofei He

Open-vocabulary Temporal Action Detection (Open-vocab TAD) is an advanced video analysis approach that expands Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) capabilities. Closed-vocab TAD is typically confined to localizing…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Trung Thanh Nguyen , Yasutomo Kawanishi , Takahiro Komamizu , Ichiro Ide

Millions of visually impaired people depend on relatives and friends to perform their everyday tasks. One relevant step towards self-sufficiency is to provide them with means to verify the value and operation presented in payment machines.…

Computer Vision and Pattern Recognition · Computer Science 2018-12-17 Guilherme Folego , Filipe Costa , Bruno Costa , Alan Godoy , Luiz Pita

Active speaker detection (ASD) in multimodal environments is crucial for various applications, from video conferencing to human-robot interaction. This paper introduces FabuLight-ASD, an advanced ASD model that integrates facial, audio, and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Hugo Carneiro , Stefan Wermter

Personalized speech enhancement (PSE) models achieve promising results compared with unconditional speech enhancement models due to their ability to remove interfering speech in addition to background noise. Unlike unconditional speech…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-08 Hassan Taherian , Sefik Emre Eskimez , Takuya Yoshioka