中文
相关论文

相关论文: Comparative Analysis of Personalized Voice Activit…

200 篇论文

State-of-the-art Active Speaker Detection (ASD) approaches heavily rely on audio and facial features to perform, which is not a sustainable approach in wild scenarios. Although these methods achieve good results in the standard…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Tiago Roxo , Joana C. Costa , Pedro R. M. Inácio , Hugo Proença

Reverberation not only degrades the quality of speech for human perception, but also severely impacts the accuracy of automatic speech recognition. Prior work attempts to remove reverberation based on the audio modality only. Our idea is to…

声音 · 计算机科学 2023-03-15 Changan Chen , Wei Sun , David Harwath , Kristen Grauman

Video Anomaly Detection (VAD) has emerged as a pivotal task in computer vision, with broad relevance across multiple fields. Recent advances in deep learning have driven significant progress in this area, yet the field remains fragmented…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Ghazal Alinezhad Noghre , Armin Danesh Pazho , Hamed Tabkhi

Keyword spotting (KWS) and speaker verification (SV) have been studied independently although it is known that acoustic and speaker domains are complementary. In this paper, we propose a multi-task network that performs KWS and SV…

音频与语音处理 · 电气工程与系统科学 2020-08-10 Myunghun Jung , Youngmoon Jung , Jahyun Goo , Hoirin Kim

Own voice pickup technology for hearable devices facilitates communication in noisy environments. Own voice reconstruction (OVR) systems enhance the quality and intelligibility of the recorded noisy own voice signals. Since disturbances…

音频与语音处理 · 电气工程与系统科学 2026-03-04 Mattes Ohlenbusch , Christian Rollwage , Simon Doclo , Jan Rennies

Visual Anomaly Detection (VAD) has gained significant research attention for its ability to identify anomalous images and pinpoint the specific areas responsible for the anomaly. A key advantage of VAD is its unsupervised nature, which…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Manuel Barusco , Francesco Borsatti , Davide Dalle Pezze , Francesco Paissan , Elisabetta Farella , Gian Antonio Susto

Social Anxiety Disorder (SAD) significantly impacts individuals' daily lives and relationships. The conventional methods for SAD detection involve physical consultations and self-reported questionnaires, but they have limitations such as…

计算机与社会 · 计算机科学 2025-01-13 Nilesh Kumar Sahu , Nandigramam Sai Harshit , Rishabh Uikey , Haroon R. Lone

Research has shown that trust is an essential aspect of human-computer interaction directly determining the degree to which the person is willing to use the system. An automatic prediction of the level of trust that a user has on a certain…

音频与语音处理 · 电气工程与系统科学 2020-08-03 Leonardo Pepino , Pablo Riera , Lara Gauder , Agustín Gravano , Luciana Ferrer

To better model the contextual information and increase the generalization ability of Speech Activity Detection (SAD) system, this paper leverages a multi-lingual Automatic Speech Recognition (ASR) system to perform SAD. Sequence…

声音 · 计算机科学 2021-04-13 Seyyed Saeed Sarfjoo , Srikanth Madikeri , Petr Motlicek

Video anomaly detection (VAD) holds immense importance across diverse domains such as surveillance, healthcare, and environmental monitoring. While numerous surveys focus on conventional VAD methods, they often lack depth in exploring…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Moshira Abdalla , Sajid Javed , Muaz Al Radi , Anwaar Ulhaq , Naoufel Werghi

Speaker diarization, the process of identifying "who spoke when" in audio recordings, is essential for understanding classroom dynamics. However, classroom settings present distinct challenges, including poor recording quality, high levels…

声音 · 计算机科学 2025-05-28 Ali Sartaz Khan , Tolulope Ogunremi , Ahmed Adel Attia , Dorottya Demszky

Voice has become an increasingly popular User Interaction (UI) channel, mainly contributing to the ongoing trend of wearables, smart vehicles, and home automation systems. Voice assistants such as Siri, Google Now and Cortana, have become…

密码学与安全 · 计算机科学 2017-01-18 Huan Feng , Kassem Fawaz , Kang G. Shin

Performing an adequate evaluation of sound event detection (SED) systems is far from trivial and is still subject to ongoing research. The recently proposed polyphonic sound detection (PSD)-receiver operating characteristic (ROC) and PSD…

音频与语音处理 · 电气工程与系统科学 2022-02-01 Janek Ebbers , Romain Serizel , Reinhold Haeb-Umbach

This paper addresses the issue of active speaker detection (ASD) in noisy environments and formulates a robust active speaker detection (rASD) problem. Existing ASD approaches leverage both audio and visual modalities, but non-speech sounds…

多媒体 · 计算机科学 2024-04-02 Siva Sai Nagender Vasireddy , Chenxu Zhang , Xiaohu Guo , Yapeng Tian

Speaker verification (SV) has recently attracted considerable research interest due to the growing popularity of virtual assistants. At the same time, there is an increasing requirement for an SV system: it should be robust to short speech…

音频与语音处理 · 电气工程与系统科学 2020-10-07 Youngmoon Jung , Yeunju Choi , Hyungjun Lim , Hoirin Kim

For face presentation attack detection (PAD), most of the spoofing cues are subtle, local image patterns (e.g., local image distortion, 3D mask edge and cut photo edges). The representations of existing PAD works with simple global pooling…

计算机视觉与模式识别 · 计算机科学 2022-02-22 Jiong Wang , Zhou Zhao , Weike Jin , Xinyu Duan , Zhen Lei , Baoxing Huai , Yiling Wu , Xiaofei He

Open-vocabulary Temporal Action Detection (Open-vocab TAD) is an advanced video analysis approach that expands Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) capabilities. Closed-vocab TAD is typically confined to localizing…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Trung Thanh Nguyen , Yasutomo Kawanishi , Takahiro Komamizu , Ichiro Ide

Millions of visually impaired people depend on relatives and friends to perform their everyday tasks. One relevant step towards self-sufficiency is to provide them with means to verify the value and operation presented in payment machines.…

计算机视觉与模式识别 · 计算机科学 2018-12-17 Guilherme Folego , Filipe Costa , Bruno Costa , Alan Godoy , Luiz Pita

Active speaker detection (ASD) in multimodal environments is crucial for various applications, from video conferencing to human-robot interaction. This paper introduces FabuLight-ASD, an advanced ASD model that integrates facial, audio, and…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Hugo Carneiro , Stefan Wermter

Personalized speech enhancement (PSE) models achieve promising results compared with unconditional speech enhancement models due to their ability to remove interfering speech in addition to background noise. Unlike unconditional speech…

音频与语音处理 · 电气工程与系统科学 2022-11-08 Hassan Taherian , Sefik Emre Eskimez , Takuya Yoshioka