English
Related papers

Related papers: Deep Neural Network Voice Activity Detector for Do…

200 papers

This paper presents a new task named weakly-supervised group activity recognition (GAR) which differs from conventional GAR tasks in that only video-level labels are available, yet the important persons within each frame are not provided…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Rui Yan , Lingxi Xie , Jinhui Tang , Xiangbo Shu , Qi Tian

The rapid development of deep learning and generative AI technologies has profoundly transformed the digital contact landscape, creating realistic Deepfake that poses substantial challenges to public trust and digital media integrity. This…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Ying Xu , Marius Pedersen , Kiran Raja

Speech signals are subjected to more acoustic interference and emotional factors than other signals. Noisy emotion-riddled speech data is a challenge for real-time speech processing applications. It is essential to find an effective way to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-25 Shibani Hamsa , Ismail Shahin , Youssef Iraqi , Ernesto Damiani , Naoufel Werghi

Understanding human behavior and monitoring mental health are essential to maintaining the community and society's safety. As there has been an increase in mental health problems during the COVID-19 pandemic due to uncontrolled mental…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-28 Nelly Elsayed , Zag ElSayed , Navid Asadizanjani , Murat Ozer , Ahmed Abdelgawad , Magdy Bayoumi

This work focuses on reliable detection of bird sound emissions as recorded in the open field. Acoustic detection of avian sounds can be used for the automatized monitoring of multiple bird taxa and querying in long-term recordings for…

Sound · Computer Science 2016-09-28 Ilyas Potamitis

Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a…

Sound · Computer Science 2025-05-30 Zhaokai Sun , Li Zhang , Qing Wang , Pan Zhou , Lei Xie

Singing Voice Detection (SVD) has been an active area of research in music information retrieval (MIR). Currently, two deep neural network-based methods, one based on CNN and the other on RNN, exist in literature that learn optimized…

Sound · Computer Science 2021-08-23 Soumava Paul , Gurunath Reddy M , K Sreenivasa Rao , Partha Pratim Das

In many VoIP systems, Voice Activity Detection (VAD) is often used on VoIP traffic to suppress packets of silence in order to reduce the bandwidth consumption of phone calls. Unfortunately, although VoIP traffic is fully encrypted and…

Cryptography and Security · Computer Science 2023-06-02 Roy Laurens , Edo Christianto , Bruce Caulkins , Cliff C. Zou

The adoption of advanced deep learning architectures in stuttering detection (SD) tasks is challenging due to the limited size of the available datasets. To this end, this work introduces the application of speech embeddings extracted from…

Sound · Computer Science 2023-06-02 Shakeel A. Sheikh , Md Sahidullah , Fabrice Hirsch , Slim Ouni

Activity recognition has become a popular research branch in the field of pervasive computing in recent years. A large number of experiments can be obtained that activity sensor-based data's characteristic in activity recognition is…

Computer Vision and Pattern Recognition · Computer Science 2018-05-21 Li Xue , Si Xiandong , Nie Lanshun , Li Jiazhen , Ding Renjie , Zhan Dechen , Chu Dianhui

This paper addresses the issue of active speaker detection (ASD) in noisy environments and formulates a robust active speaker detection (rASD) problem. Existing ASD approaches leverage both audio and visual modalities, but non-speech sounds…

Multimedia · Computer Science 2024-04-02 Siva Sai Nagender Vasireddy , Chenxu Zhang , Xiaohu Guo , Yapeng Tian

Dysarthric speech recognition is a challenging task due to acoustic variability and limited amount of available data. Diverse conditions of dysarthric speakers account for the acoustic variability, which make the variability difficult to be…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-17 Xurong Xie , Rukiye Ruzi , Xunying Liu , Lan Wang

Spatio-temporal action detection (STAD) aims to classify the actions present in a video and localize them in space and time. It has become a particularly active area of research in computer vision because of its explosively emerging…

Computer Vision and Pattern Recognition · Computer Science 2023-08-04 Peng Wang , Fanwei Zeng , Yuntao Qian

This work focuses on reliable detection and segmentation of bird vocalizations as recorded in the open field. Acoustic detection of avian sounds can be used for the automatized monitoring of multiple bird taxa and querying in long-term…

Audio and Speech Processing · Electrical Eng. & Systems 2017-11-20 Lefteris Fanioudakis , Ilyas Potamitis

Social interactions play a crucial role in shaping human behavior, relationships, and societies. It encompasses various forms of communication, such as verbal conversation, non-verbal gestures, facial expressions, and body language. In this…

Machine Learning · Computer Science 2026-05-13 Alice Zhang , Callihan Bertley , Dawei Liang , Edison Thomaz

Batteryless or so called passive wearables are providing new and innovative methods for human activity recognition (HAR), especially in healthcare applications for older people. Passive sensors are low cost, lightweight, unobtrusive and…

Machine Learning · Computer Science 2019-06-07 Alireza Abedin , S. Hamid Rezatofighi , Qinfeng Shi , Damith C. Ranasinghe

The ongoing biodiversity crisis, driven by factors such as land-use change and global warming, emphasizes the need for effective ecological monitoring methods. Acoustic monitoring of biodiversity has emerged as an important monitoring tool.…

Sound · Computer Science 2023-12-18 Drew Priebe , Burooj Ghani , Dan Stowell

Recent research shows that deep neural networks (DNNs) can be used to extract deep speaker vectors (d-vectors) that preserve speaker characteristics and can be used in speaker verification. This new method has been tested on text-dependent…

Computation and Language · Computer Science 2015-05-26 Lantian Li , Dong Wang , Zhiyong Zhang , Thomas Fang Zheng

Speaker diarization, the process of identifying "who spoke when" in audio recordings, is essential for understanding classroom dynamics. However, classroom settings present distinct challenges, including poor recording quality, high levels…

Sound · Computer Science 2025-05-28 Ali Sartaz Khan , Tolulope Ogunremi , Ahmed Adel Attia , Dorottya Demszky

The Automatic Speaker Verification (ASV) system is vulnerable to fraudulent activities using audio deepfakes, also known as logical-access voice spoofing attacks. These deepfakes pose a concerning threat to voice biometrics due to recent…

Sound · Computer Science 2023-10-09 Awais Khan , Khalid Mahmood Malik
‹ Prev 1 4 5 6 7 8 10 Next ›