English
Related papers

Related papers: WhAM: Towards A Translative Model of Sperm Whale V…

200 papers

Animal acoustic communication often exhibits temporal dependence, with calls triggering or suppressing subsequent calls within and across call types, individuals, or species. While Hawkes processes provide a natural framework for modeling…

We introduce VampNet, a masked acoustic token modeling approach to music synthesis, compression, inpainting, and variation. We use a variable masking schedule during training which allows us to sample coherent music from the model by…

Sound · Computer Science 2023-07-13 Hugo Flores Garcia , Prem Seetharaman , Rithesh Kumar , Bryan Pardo

Besides humans, several marine mammal species exhibit prerequisites to evolve language: high cognitive abilities, flexibility in vocal production and advanced social interactions. Here, we describe and analyse the vocal repertoire of…

Quantitative Methods · Quantitative Biology 2015-06-15 Heike Vester , Kurt Hammerschmidt , Marc Timme , Sarah Hallerberg

The study of animal communication often involves categorizing units into types (e.g. syllables in songbirds, or notes in humpback whales). While this approach is useful in many cases, it necessarily flattens the complexity and nuance…

Sound · Computer Science 2025-12-23 Mason Youngblood

The marine ecosystem is changing at an alarming rate, exhibiting biodiversity loss and the migration of tropical species to temperate basins. Monitoring the underwater environments and their inhabitants is of fundamental importance to…

Sound · Computer Science 2022-01-17 Michele Mancusi , Nicola Zonca , Emanuele Rodolà , Silvia Zuffi

This paper introduces WaveNet, a deep neural network for generating raw audio waveforms. The model is fully probabilistic and autoregressive, with the predictive distribution for each audio sample conditioned on all previous ones;…

Token-based text-to-speech (TTS) models have emerged as a promising avenue for generating natural and realistic speech, yet they grapple with low pronunciation accuracy, speaking style and timbre inconsistency, and a substantial need for…

Sound · Computer Science 2024-03-12 Chunhui Wang , Chang Zeng , Bowen Zhang , Ziyang Ma , Yefan Zhu , Zifeng Cai , Jian Zhao , Zhonglin Jiang , Yong Chen

The generation of co-speech gestures for digital humans is an emerging area in the field of virtual human creation. Prior research has made progress by using acoustic and semantic information as input and adopting classify method to…

Sound · Computer Science 2024-04-16 Fan Zhang , Naye Ji , Fuxing Gao , Siyuan Zhao , Zhaohan Wang , Shunman Li

A challenge in marine bioacoustic analysis is the detection of animal signals, like calls, whistles and clicks, for behavioral studies. Manual labeling is too time-consuming to process sufficient data to get reasonable results. Thus, an…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-23 Christopher Hauer

Cross-lingual alignment in pretrained language models enables knowledge transfer across languages. Similar alignment has been reported in Whisper-style speech encoders, based on spoken translation retrieval using representational…

Computation and Language · Computer Science 2026-04-07 Ryan Soh-Eun Shim , Domenico De Cristofaro , Chengzhi Martin Hu , Alessandro Vietti , Barbara Plank

Recent research on word-level confidence estimation for speech recognition systems has primarily focused on lightweight models known as Confidence Estimation Modules (CEMs), which rely on hand-engineered features derived from Automatic…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-20 Vaibhav Aggarwal , Shabari S Nair , Yash Verma , Yash Jogi

Vocoders, encoding speech signals into acoustic features and allowing for speech signal reconstruction from them, have been studied for decades. Recently, the rise of deep learning has particularly driven the development of neural vocoders…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-03 Shaowen Chen , Tomoki Toda

Speech foundation models, such as OpenAI's Whisper, become the state of the art in speech understanding due to their strong accuracy and generalizability. Yet, their applications are mostly limited to processing pre-recorded speech, whereas…

Sound · Computer Science 2025-04-23 Rongxiang Wang , Zhiming Xu , Felix Xiaozhu Lin

Audio Language Models (ALM) have emerged as the dominant paradigm for speech and music generation by representing audio as sequences of discrete tokens. Yet, unlike text tokens, which are invertible, audio tokens are extracted from lossy…

Sound · Computer Science 2026-01-14 Simon Rouard , Manu Orsini , Axel Roebel , Neil Zeghidour , Alexandre Défossez

Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, prevailing methods augment acoustic…

Sound · Computer Science 2026-01-28 Xin Zhang , Lin Li , Xiangni Lu , Jianquan Liu , Kong Aik Lee

This paper reports on the development of a large-scale speech recognition model, Whale. Similar to models such as Whisper and OWSM, Whale leverages both a large model size and a diverse, extensive dataset. Whale's architecture integrates…

Computation and Language · Computer Science 2025-06-03 Yosuke Kashiwagi , Hayato Futami , Emiru Tsunoo , Satoshi Asakawa

While significant advances have been made with respect to the separation of overlapping speech signals, studies have been largely constrained to mixtures of clean, near anechoic speech, not representative of many real-world scenarios.…

Sound · Computer Science 2020-02-17 Matthew Maciejewski , Gordon Wichern , Emmett McQuinn , Jonathan Le Roux

Whistle contour extraction aims to derive animal whistles from time-frequency spectrograms as polylines. For toothed whales, whistle extraction results can serve as the basis for analyzing animal abundance, species identity, and social…

Computer Vision and Pattern Recognition · Computer Science 2023-04-07 Pu Li , Marie Roch , Holger Klinck , Erica Fleishman , Douglas Gillespie , Eva-Marie Nosal , Yu Shiu , Xiaobai Liu

This Ph.D. thesis focuses on developing a system for high-quality speech synthesis and voice conversion. Vocoder-based speech analysis, manipulation, and synthesis plays a crucial role in various kinds of statistical parametric speech…

Sound · Computer Science 2021-01-26 Mohammed Salah Al-Radhi