English
Related papers

Related papers: Utilizing synthetic training data for the supervis…

200 papers

Automatic identification of animal species by their vocalization is an important and challenging task. Although many kinds of audio monitoring system have been proposed in the literature, they suffer from several disadvantages such as…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-25 Weitao Xu , Xiang Zhang , Lina Yao , Wanli Xue , Bo Wei

Singing voice synthesis (SVS) has seen remarkable advancements in recent years. However, compared to speech and general audio data, publicly available singing datasets remain limited. In practice, this data scarcity often leads to…

Sound · Computer Science 2025-12-17 Yiwen Zhao , Jiatong Shi , Yuxun Tang , William Chen , Shinji Watanabe

With the huge technological advances introduced by deep learning in audio & speech processing, many novel synthetic speech techniques achieved incredible realistic results. As these methods generate realistic fake human voices, they can be…

Convolutional neural networks (CNN) are one of the best-performing neural network architectures for environmental sound classification (ESC). Recently, temporal attention mechanisms have been used in CNN to capture the useful information…

Sound · Computer Science 2020-05-22 Helin Wang , Yuexian Zou , Dading Chong , Wenwu Wang

Key challenges in developing underwater acoustic localization methods are related to the combined effects of high reverberation in intricate environments. To address such challenges, recent studies have shown that with a properly designed…

Signal Processing · Electrical Eng. & Systems 2023-05-30 Amir Weiss , Andrew C. Singer , Gregory W. Wornell

Automatically detecting sound units of humpback whales in complex time-varying background noises is a current challenge for scientists. In this paper, we explore the applicability of Convolution Neural Network (CNN) method for this task. In…

Machine Learning · Statistics 2017-04-03 Cazau Dorian , Riwal Lefort , Julien Bonnel , Jean-Luc Zarader , Olivier Adam

In computer vision, convolutional neural networks (CNN) such as ConvNeXt, have been able to surpass state-of-the-art transformers, partly thanks to depthwise separable convolutions (DSC). DSC, as an approximation of the regular convolution,…

Automated classification of animal sounds is a prerequisite for large-scale monitoring of biodiversity. Convolutional Neural Networks (CNNs) are among the most promising algorithms but they are slow, often achieve poor classification in the…

Next to decision tree and k-nearest neighbours algorithms deep convolutional neural networks (CNNs) are widely used to classify audio data in many domains like music, speech or environmental sounds. To train a specific CNN various spectral…

Sound · Computer Science 2025-09-16 Friedrich Wolf-Monheim

Self-supervised learning (SSL) foundation models have emerged as powerful, domain-agnostic, general-purpose feature extractors applicable to a wide range of tasks. Such models pre-trained on human speech have demonstrated high…

Machine Learning · Computer Science 2025-01-22 Eklavya Sarkar , Mathew Magimai. -Doss

This paper introduces a set of acoustic modeling techniques for utterance verification (UV) based continuous speech recognition (CSR). Utterance verification in this work implies the ability to determine when portions of a hypothesized word…

Formal Languages and Automata Theory · Computer Science 2014-01-22 M. Tharun Prasath

State-of-the-art singing voice separation is based on deep learning making use of CNN structures with skip connections (like U-net model, Wave-U-Net model, or MSDENSELSTM). A key to the success of these models is the availability of a large…

Sound · Computer Science 2019-06-25 Alice Cohen-Hadria , Axel Roebel , Geoffroy Peeters

This paper describes the extension and optimization of our previous work on very deep convolutional neural networks (CNNs) for effective recognition of noisy speech in the Aurora 4 task. The appropriate number of convolutional layers, the…

Computation and Language · Computer Science 2016-10-04 Yanmin Qian , Philip C Woodland

Convolutional Neural Networks (CNNs) have achieved comparable error rates to well-trained human on ILSVRC2014 image classification task. To achieve better performance, the complexity of CNNs is continually increasing with deeper and bigger…

Computer Vision and Pattern Recognition · Computer Science 2014-12-30 Wei Yu , Kuiyuan Yang , Yalong Bai , Hongxun Yao , Yong Rui

Automatic speech recognition systems are part of people's daily lives, embedded in personal assistants and mobile phones, helping as a facilitator for human-machine interaction while allowing access to information in a practically intuitive…

Sound · Computer Science 2021-10-05 Julio Cesar Duarte , Sérgio Colcher

This paper introduces a novel method for predicting tool wear in CNC turning operations, combining ultrasonic microphone arrays and convolutional neural networks (CNNs). High-frequency acoustic emissions between 0 kHz and 60 kHz are…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-14 Jan Steckel , Arne Aerts , Erik Verreycken , Dennis Laurijssen , Walter Daems

Deep learning offers potential for various healthcare applications, yet requires extensive datasets of curated medical images where data privacy, cost, and distribution mismatch across various acquisition centers could become major…

Image and Video Processing · Electrical Eng. & Systems 2024-02-02 Kasra Naftchi-Ardebili , Karanpartap Singh , Reza Pourabolghasem , Pejman Ghanouni , Gerald R. Popelka , Kim Butts Pauly

The scattering framework offers an optimal hierarchical convolutional decomposition according to its kernels. Convolutional Neural Net (CNN) can be seen as an optimal kernel decomposition, nevertheless it requires large amount of training…

Sound · Computer Science 2017-01-24 Herve Glotin , Julien Ricard , Randall Balestriero

Recent successful applications of convolutional neural networks (CNNs) to audio classification and speech recognition have motivated the search for better input representations for more efficient training. Visual displays of an audio…

Computer Vision and Pattern Recognition · Computer Science 2017-06-23 M. Huzaifah

Understanding evolution of vocal communication in social animals is an important research problem. In that context, beyond humans, there is an interest in analyzing vocalizations of other social animals such as, meerkats, marmosets, apes.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-29 Imen Ben Mahmoud , Eklavya Sarkar , Marta Manser , Mathew Magimai. -Doss