English
Related papers

Related papers: WhAM: Towards A Translative Model of Sperm Whale V…

200 papers

An algorithm for detecting tonal vocalizations from estuarine dolphin (Sotalia guianensis) specimens without interference of a human operator is developed. The raw audio data collected from a passive monitoring sensor in the Canan\'eia…

Sound · Computer Science 2019-09-11 O. M. Serra , F. P. R. Martins , L. R. Padovese

Recent advancements in Large Audio Language Models (LALMs) have demonstrated exceptional performance in speech recognition and translation. However, existing models often suffer from a disconnect between perception and expression, resulting…

Sound · Computer Science 2026-03-02 Yueran Hou , Peilei Jia , Zihan Sun , Qihang Lu , Wenbing Yang , Yingming Gao , Ya Li , Jun Gao

Systematic evaluation of speech separation and enhancement models under moving sound source conditions requires extensive and diverse data. However, real-world datasets often lack sufficient data for training and evaluation, and synthetic…

Sound · Computer Science 2025-03-07 Kai Li , Wendi Sang , Chang Zeng , Runxuan Yang , Guo Chen , Xiaolin Hu

This paper proposes a WaveNet-based neural excitation model (ExcitNet) for statistical parametric speech synthesis systems. Conventional WaveNet-based neural vocoding systems significantly improve the perceptual quality of synthesized…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-23 Eunwoo Song , Kyungguen Byun , Hong-Goo Kang

Large general-purpose transformer models have recently become the mainstay in the realm of speech analysis. In particular, Whisper achieves state-of-the-art results in relevant tasks such as speech recognition, translation, language…

Sound · Computer Science 2024-05-07 Antonio Bevilacqua , Paolo Saviano , Alessandro Amirante , Simon Pietro Romano

This paper proposes a hierarchical spatial-temporal model for modelling the spectrograms of animal calls. The motivation stems from analyzing recordings of the so-called grunt calls emitted by various lemur species. Our goal is to identify…

Wideband Audio Waveform Evaluation Networks (WAWEnets) are convolutional neural networks that operate directly on wideband audio waveforms in order to produce evaluations of those waveforms. In the present work these evaluations give…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-21 Andrew Catellier , Stephen Voran

Wake word (WW) spotting is challenging in far-field not only because of the interference in signal transmission but also the complexity in acoustic environments. Traditional WW model training requires large amount of in-domain WW-specific…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-15 Yixin Gao , Yuriy Mishchenko , Anish Shah , Spyros Matsoukas , Shiv Vitaladevuni

Humans encode information into sounds by controlling articulators and decode information from sounds using the auditory apparatus. This paper introduces CiwaGAN, a model of human spoken language acquisition that combines unsupervised…

Sound · Computer Science 2023-09-15 Gašper Beguš , Thomas Lu , Alan Zhou , Peter Wu , Gopala K. Anumanchipalli

Recent advances in visually-induced audio generation are based on sampling short, low-fidelity, and one-class sounds. Moreover, sampling 1 second of audio from the state-of-the-art model takes minutes on a high-end GPU. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2021-10-19 Vladimir Iashin , Esa Rahtu

Effective conservation of maritime environments and wildlife management of endangered species require the implementation of efficient, accurate and scalable solutions for environmental monitoring. Ecoacoustics offers the advantages of…

Sound · Computer Science 2025-07-29 Burla Nur Korkmaz , Roee Diamant , Gil Danino , Alberto Testolin

Underwater acoustic (UWA) communication plays a key role in the process of exploring and studying the ocean. In this paper, a modified non-stationary wideband channel model for UWA communication in shallow water scenarios is proposed. In…

Signal Processing · Electrical Eng. & Systems 2021-08-17 Xiuming Zhu , Cheng-Xiang Wang , Ruofei Ma

This paper proposes a data-efficient, semi-supervised, two-pass framework for segmenting bird vocalizations. The framework utilizes a binary classification model to categorize frames of an input audio recording into the background or bird…

Audio and Speech Processing · Electrical Eng. & Systems 2019-02-27 Anshul Thakur , Padmanabhan Rajan

Modern speech synthesis uses neural vocoders to model raw waveform samples directly. This increased versatility has expanded the scope of vocoders from speech to other domains, such as music. We address another interesting domain of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-22 Rhythm Bhatia , Tomi H. Kinnunen

In this work, we propose WaveFlow, a small-footprint generative flow for raw audio, which is directly trained with maximum likelihood. It handles the long-range structure of 1-D waveform with a dilated 2-D convolutional architecture, while…

Sound · Computer Science 2020-06-26 Wei Ping , Kainan Peng , Kexin Zhao , Zhao Song

This paper presents VoiceLDM, a model designed to produce audio that accurately follows two distinct natural language text prompts: the description prompt and the content prompt. The former provides information about the overall…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-26 Yeonghyeon Lee , Inmo Yeon , Juhan Nam , Joon Son Chung

Singing voice conversion aims to convert singer's voice from source to target without changing singing content. Parallel training data is typically required for the training of singing voice conversion system, that is however not practical…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-04 Junchen Lu , Kun Zhou , Berrak Sisman , Haizhou Li

Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper, we tackle…

While contemporary speech separation technologies adeptly process lengthy mixed audio waveforms, they are frequently challenged by the intricacies of real-world environments, including noisy and reverberant settings, which can result in…

Sound · Computer Science 2025-05-27 Zhaoxi Mu , Xinyu Yang , Gang Wang

The rapid advances in text-to-speech (TTS) technologies have made audio deepfakes increasingly realistic and accessible, raising significant security and trust concerns. While existing research has largely focused on detecting…

Sound · Computer Science 2026-02-03 Alabi Ahmed , Vandana Janeja , Sanjay Purushotham