中文
相关论文

相关论文: Voice Activity Detection for Ultrasound-based Sile…

200 篇论文

State-of-the-art Active Speaker Detection (ASD) approaches mainly use audio and facial features as input. However, the main hypothesis in this paper is that body dynamics is also highly correlated to "speaking" (and "listening") actions and…

计算机视觉与模式识别 · 计算机科学 2024-12-12 Tiago Roxo , Joana C. Costa , Pedro Inácio , Hugo Proença

Advances in machine learning and contactless sensors have enabled the understanding complex human behaviors in a healthcare setting. In particular, several deep learning systems have been introduced to enable comprehensive analysis of…

计算机视觉与模式识别 · 计算机科学 2024-03-06 Pengbo Wei , David Ahmedt-Aristizabal , Harshala Gammulle , Simon Denman , Mohammad Ali Armin

Meetings are a common activity in professional contexts, and it remains challenging to endow vocal assistants with advanced functionalities to facilitate meeting management. In this context, a task like active speaker detection can provide…

计算机视觉与模式识别 · 计算机科学 2022-06-22 Lionel Pibre , Francisco Madrigal , Cyrille Equoy , Frédéric Lerasle , Thomas Pellegrini , Julien Pinquier , Isabelle Ferrané

Automatic speech recognition (ASR) is a key area in computational linguistics, focusing on developing technologies that enable computers to convert spoken language into text. This field combines linguistics and machine learning. ASR models,…

计算与语言 · 计算机科学 2024-06-27 Anish Saha , A. G. Ramakrishnan

We study the problem of detecting talking activities in collaborative learning videos. Our approach uses head detection and projections of the log-magnitude of optical flow vectors to reduce the problem to a simple classification of small…

计算机视觉与模式识别 · 计算机科学 2021-10-18 Wenjing Shi , Marios S. Pattichis , Sylvia Celedón-Pattichis , Carlos LópezLeiva

Advances of deep learning for Artificial Neural Networks(ANNs) have led to significant improvements in the performance of digital signal processing systems implemented on digital chips. Although recent progress in low-power chips is…

音频与语音处理 · 电气工程与系统科学 2021-05-20 Giorgia Dellaferrera , Flavio Martinelli , Milos Cernak

Speaker diarization for real-life scenarios is an extremely challenging problem. Widely used clustering-based diarization approaches perform rather poorly in such conditions, mainly due to the limited ability to handle overlapping speech.…

Thousands of individuals need surgical removal of their larynx due to critical diseases every year and therefore, require an alternative form of communication to articulate speech sounds after the loss of their voice box. This work…

图像与视频处理 · 电气工程与系统科学 2020-07-01 Pramit Saha , Yadong Liu , Bryan Gick , Sidney Fels

Singing voice detection (SVD), to recognize vocal parts in the song, is an essential task in music information retrieval (MIR). The task remains challenging since singing voice varies and intertwines with the accompaniment music, especially…

音频与语音处理 · 电气工程与系统科学 2022-05-09 Yifu Sun , Xulong Zhang , Yi Yu , Xi Chen , Wei Li

In a speech recognition system, voice activity detection (VAD) is a crucial frontend module. Addressing the issues of poor noise robustness in traditional binary VAD systems based on DFSMN, the paper further proposes semantic VAD based on…

声音 · 计算机科学 2023-12-25 Lingyun Zuo , Keyu An , Shiliang Zhang , Zhijie Yan

Auditory attention decoding (AAD) is a technique used to identify and amplify the talker that a listener is focused on in a noisy environment. This is done by comparing the listener's brainwaves to a representation of all the sound sources…

音频与语音处理 · 电气工程与系统科学 2023-02-14 Cong Han , Vishal Choudhari , Yinghao Aaron Li , Nima Mesgarani

The problem of identifying voice commands has always been a challenge due to the presence of noise and variability in speed, pitch, etc. We will compare the efficacies of several neural network architectures for the speech recognition…

机器学习 · 统计学 2020-11-25 Sanjay Krishna Gouda , Salil Kanetkar , David Harrison , Manfred K Warmuth

In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling audio-visual synchronization cue, which depends on the video…

音频与语音处理 · 电气工程与系统科学 2023-06-13 Yidi Jiang , Ruijie Tao , Zexu Pan , Haizhou Li

This paper presents an audio-visual approach for voice separation which produces state-of-the-art results at a low latency in two scenarios: speech and singing voice. The model is based on a two-stage network. Motion cues are obtained with…

声音 · 计算机科学 2022-07-20 Juan F. Montesinos , Venkatesh S. Kadandale , Gloria Haro

In this paper, we propose the use of self-supervised pretraining on a large unlabelled data set to improve the performance of a personalized voice activity detection (VAD) model in adverse conditions. We pretrain a long short-term memory…

声音 · 计算机科学 2024-01-24 Holger Severin Bovbjerg , Jesper Jensen , Jan Østergaard , Zheng-Hua Tan

We introduce a deep learning model for speech denoising, a long-standing challenge in audio analysis arising in numerous applications. Our approach is based on a key observation about human speech: there is often a short pause between each…

声音 · 计算机科学 2020-10-26 Ruilin Xu , Rundi Wu , Yuko Ishiwaka , Carl Vondrick , Changxi Zheng

Speaker verification (SV) has recently attracted considerable research interest due to the growing popularity of virtual assistants. At the same time, there is an increasing requirement for an SV system: it should be robust to short speech…

音频与语音处理 · 电气工程与系统科学 2020-10-07 Youngmoon Jung , Yeunju Choi , Hyungjun Lim , Hoirin Kim

Voice recognition and speaker identification are vital for applications in security and personal assistants. This paper presents a lightweight 1D-Convolutional Neural Network (1D-CNN) designed to perform speaker identification on minimal…

声音 · 计算机科学 2024-11-25 Irfan Nafiz Shahan , Pulok Ahmed Auvi

This paper delves into the challenging task of Active Speaker Detection (ASD), where the system needs to determine in real-time whether a person is speaking or not in a series of video frames. While previous works have made significant…

计算机视觉与模式识别 · 计算机科学 2024-09-16 Arnav Kundu , Yanzi Jin , Mohammad Sekhavat , Max Horton , Danny Tormoen , Devang Naik

Estimating noise information exactly is crucial for noise aware training in speech applications including speech enhancement (SE) which is our focus in this paper. To estimate noise-only frames, we employ voice activity detection (VAD) to…

音频与语音处理 · 电气工程与系统科学 2020-12-04 Joohyung Lee , Youngmoon Jung , Myunghun Jung , Hoirin Kim