中文
相关论文

相关论文: BC-VAD: A Robust Bone Conduction Voice Activity De…

200 篇论文

The frequent breakdowns and malfunctions of industrial equipment have driven increasing interest in utilizing cost-effective and easy-to-deploy sensors, such as microphones, for effective condition monitoring of machinery. Microphones offer…

Spatio-temporal action detection (STAD) is an important fine-grained video understanding task. Current methods require box and label supervision for all action classes in advance. However, in real-world applications, it is very likely to…

计算机视觉与模式识别 · 计算机科学 2024-05-20 Tao Wu , Shuqiu Ge , Jie Qin , Gangshan Wu , Limin Wang

Direct speech-to-text translation (ST) models are usually trained on corpora segmented at sentence level, but at inference time they are commonly fed with audio split by a voice activity detector (VAD). Since VAD segmentation is not…

计算与语言 · 计算机科学 2020-08-06 Marco Gaido , Mattia Antonino Di Gangi , Matteo Negri , Mauro Cettolo , Marco Turchi

Deep convolutional neural networks (CNNs) have been applied to extracting speaker embeddings with significant success in speaker verification. Incorporating the attention mechanism has shown to be effective in improving the model…

音频与语音处理 · 电气工程与系统科学 2022-11-01 Jingyu Li , Yusheng Tian , Tan Lee

Estimating noise information exactly is crucial for noise aware training in speech applications including speech enhancement (SE) which is our focus in this paper. To estimate noise-only frames, we employ voice activity detection (VAD) to…

音频与语音处理 · 电气工程与系统科学 2020-12-04 Joohyung Lee , Youngmoon Jung , Myunghun Jung , Hoirin Kim

Developing microphone array technologies for a small number of microphones is important due to the constraints of many devices. One direction to address this situation consists of virtually augmenting the number of microphone signals, e.g.,…

音频与语音处理 · 电气工程与系统科学 2021-01-13 Tsubasa Ochiai , Marc Delcroix , Tomohiro Nakatani , Rintaro Ikeshita , Keisuke Kinoshita , Shoko Araki

Visual speech recognition (VSR), which decodes spoken words from video data, offers significant benefits, particularly when audio is unavailable. However, the high dimensionality of video data leads to prohibitive computational costs that…

计算机视觉与模式识别 · 计算机科学 2025-02-10 Iason Ioannis Panagos , Giorgos Sfikas , Christophoros Nikou

In this paper we propose a robust loudspeaker beamforming algorithm which is used to enhance the performance of voice driven applications in scenarios where the loudspeakers introduce the majority of the noise, e.g. when music is playing…

音频与语音处理 · 电气工程与系统科学 2025-05-21 Dimme de Groot , Baturalp Karslioglu , Odette Scharenborg , Jorge Martinez

Speech Emotion Recognition (SER) often operates on speech segments detected by a Voice Activity Detection (VAD) model. However, VAD models may output flawed speech segments, especially in noisy environments, resulting in degraded…

声音 · 计算机科学 2024-10-18 Natsuo Yamashita , Masaaki Yamamoto , Yohei Kawaguchi

When speaking in presence of background noise, humans reflexively change their way of speaking in order to improve the intelligibility of their speech. This reflex is known as Lombard effect. Collecting speech in Lombard conditions is…

音频与语音处理 · 电气工程与系统科学 2019-11-05 Daniel Michelsanti , Zheng-Hua Tan , Sigurdur Sigurdsson , Jesper Jensen

Video Anomaly Detection (VAD) is an essential yet challenging task in signal processing. Since certain anomalies cannot be detected by isolated analysis of either temporal or spatial information, the interaction between these two types of…

计算机视觉与模式识别 · 计算机科学 2023-07-07 Zhiyuan Ning , Zhangxun Li , Zhengliang Guo , Zile Wang , Liang Song

Class Activation Maps (CAMs) are one of the important methods for visualizing regions used by deep learning models. Yet their robustness to different noise remains underexplored. In this work, we evaluate and report the resilience of…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Syamantak Sarkar , Revoti P. Bora , Bhupender Kaushal , Sudhish N George , Kiran Raja

Millimeter Wave (mmWave) radar has emerged as a promising modality for speech sensing, offering advantages over traditional microphones. Prior works have demonstrated that radar captures motion signals related to vocal vibrations, but there…

音频与语音处理 · 电气工程与系统科学 2025-03-21 Isabella Lenz , Yu Rong , Daniel Bliss , Julie Liss , Visar Berisha

Video anomaly detection (VAD) aims to identify abnormal events in videos. Traditional VAD methods generally suffer from the high costs of labeled data and full training, thus some recent works have explored leveraging frozen multi-modal…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Zhaolin Cai , Fan Li , Huiyu Duan , Lijun He , Guangtao Zhai

As speech processing systems in mobile and edge devices become more commonplace, the demand for unintrusive speech quality monitoring increases. Deep learning methods provide high-quality estimates of objective and subjective speech quality…

The vulnerability against presentation attacks is a crucial problem undermining the wide-deployment of face recognition systems. Though presentation attack detection (PAD) systems try to address this problem, the lack of generalization and…

计算机视觉与模式识别 · 计算机科学 2022-02-22 Anjith George , David Geissbuhler , Sebastien Marcel

Most of the existing studies on voice conversion (VC) are conducted in acoustically matched conditions between source and target signal. However, the robustness of VC methods in presence of mismatch remains unknown. In this paper, we report…

声音 · 计算机科学 2016-12-23 Monisankha Pal , Dipjyoti Paul , Md Sahidullah , Goutam Saha

Visual Anomaly Detection (VAD) has gained significant research attention for its ability to identify anomalous images and pinpoint the specific areas responsible for the anomaly. A key advantage of VAD is its unsupervised nature, which…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Manuel Barusco , Francesco Borsatti , Davide Dalle Pezze , Francesco Paissan , Elisabetta Farella , Gian Antonio Susto

We propose a single neural network architecture for two tasks: on-line keyword spotting and voice activity detection. We develop novel inference algorithms for an end-to-end Recurrent Neural Network trained with the Connectionist Temporal…

计算与语言 · 计算机科学 2016-11-30 Chris Lengerich , Awni Hannun

Passive acoustic mapping (PAM) is a promising tool for monitoring acoustic cavitation activities in the applications of ultrasound therapy. Data-adaptive beamformers for PAM have better image quality compared to the time exposure acoustics…

人工智能 · 计算机科学 2024-12-04 Yi Zeng , Jinwei Li , Hui Zhu , Shukuan Lu , Jianfeng Li , Xiran Cai
‹ 上一页 1 8 9 10 下一页 ›