English
Related papers

Related papers: Towards Evaluating Generative Audio: Insights from…

200 papers

This paper proposes a novel acoustic word embedding called Acoustic Neighbor Embeddings where speech or text of arbitrary length are mapped to a vector space of fixed, reduced dimensions by adapting stochastic neighbor embedding (SNE) to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-10 Woojay Jeon

Neural audio codecs have recently enabled high-fidelity reconstruction at high compression rates, especially for speech. However, speech and non-speech audio exhibit fundamentally different spectral characteristics: speech energy…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-11 Haoran Wang , Jiatong Shi , Jinchuan Tian , Bohan Li , Kai Yu , Shinji Watanabe

Modern generative and multimodal models increasingly rely on compact latent representations that trade and balance semantic richness with high-fidelity reconstruction. We introduce SALAD-VAE, a continuous and highly compact semantic Audio…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-10 Sebastian Braun , Hannes Gamper , Dimitra Emmanouilidou

Neural audio codecs (NACs) have garnered significant attention as key technologies for audio compression as well as audio representation for speech language models. While mainstream NAC models are predominantly convolution-based, the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-29 Haibin Wu , Naoyuki Kanda , Sefik Emre Eskimez , Jinyu Li

Music recommender systems frequently utilize network-based models to capture relationships between music pieces, artists, and users. Although these relationships provide valuable insights for predictions, new music pieces or artists often…

Sound · Computer Science 2024-09-16 Florian Grötschla , Luca Strässle , Luca A. Lanzendörfer , Roger Wattenhofer

Embedding acoustic information into fixed length representations is of interest for a whole range of applications in speech and audio technology. Two novel unsupervised approaches to generate acoustic embeddings by modelling of acoustic…

Computation and Language · Computer Science 2021-02-08 Yanpei Shi , Thomas Hain

Neural audio signal codecs have attracted significant attention in recent years. In essence, the impressive low bitrate achieved by such encoders is enabled by learning an abstract representation that captures the properties of encoded…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-06 Mhd Modar Halimeh , Matteo Torcoli , Philipp Grundhuber , Emanuël A. P. Habets

Accurate, low-latency endpointing is crucial for effective spoken dialogue systems. While traditional endpointers often rely on spectrum-based audio features, this work proposes real-time speech endpointing for multi-turn dialogues using…

Sound · Computer Science 2025-06-23 Sathvik Udupa , Shinji Watanabe , Petr Schwarz , Jan Cernocky

Automated Audio Captioning (AAC) involves generating natural language descriptions of audio content, using encoder-decoder architectures. An audio encoder produces audio embeddings fed to a decoder, usually a Transformer decoder, for…

Sound · Computer Science 2023-09-04 Étienne Labbé , Thomas Pellegrini , Julien Pinquier

The emergence of new spoofing attacks poses an increasing challenge to audio security. Current detection methods often falter when faced with unseen spoofing attacks. Traditional strategies, such as retraining with new data, are not always…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-16 Feiyi Dong , Qingchen Tang , Yichen Bai , Zihan Wang

We introduce BANC, a neural binaural audio codec designed for efficient speech compression in single and two-speaker scenarios while preserving the spatial location information of each speaker. Our key contributions are as follows: 1) The…

Sound · Computer Science 2024-11-26 Anton Ratnarajah , Shi-Xiong Zhang , Dong Yu

Neural audio codecs have significantly advanced audio compression by efficiently converting continuous audio signals into discrete tokens. These codecs preserve high-quality sound and enable sophisticated sound generation through generative…

Sound · Computer Science 2025-02-12 Xiaoyu Bie , Xubo Liu , Gaël Richard

State-of-the-art audio classification often employs a zero-shot approach, which involves comparing audio embeddings with embeddings from text describing the respective audio class. These embeddings are usually generated by neural networks…

Sound · Computer Science 2025-07-29 James Taylor , Wolfgang Mack

Audio representation learning based on deep neural networks (DNNs) emerged as an alternative approach to hand-crafted features. For achieving high performance, DNNs often need a large amount of annotated data which can be difficult and…

Machine Learning · Computer Science 2020-07-09 Xavier Favory , Konstantinos Drossos , Tuomas Virtanen , Xavier Serra

Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensuring accurate semantic modeling while maintaining high-fidelity…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-03 Yanzhou Ren , Noboru Harada , Daiki Takeuchi , Siyu Chen , Wei Liu , Xiao Zhang , Liyuan Zhang , Takehiro Moriya , Shoji Makino

We propose the use of Non-Negative Autoencoders (NAEs) for sound deconstruction and user-guided manipulation of sounds for creative purposes. NAEs offer a versatile and scalable extension of traditional Non-Negative Matrix Factorization…

Sound · Computer Science 2025-10-13 Juan José Burred , Carmine-Emanuele Cella

Audio Chord Estimation (ACE) holds a pivotal role in music information research, having garnered attention for over two decades due to its relevance for music transcription and analysis. Despite notable advancements, challenges persist in…

Sound · Computer Science 2025-09-03 Andrea Poltronieri , Xavier Serra , Martín Rocamora

Audio codecs are typically transform-domain based and efficiently code stationary audio signals, but they struggle with speech and signals containing dense transient events such as applause. Specifically, with these two classes of signals…

Audio and Speech Processing · Electrical Eng. & Systems 2020-01-28 Arijit Biswas , Dai Jia

With the ever-rising quality of deep generative models, it is increasingly important to be able to discern whether the audio data at hand have been recorded or synthesized. Although the detection of fake speech signals has been studied…

Sound · Computer Science 2024-06-14 Hafsa Ouajdi , Oussama Hadder , Modan Tailleur , Mathieu Lagrange , Laurie M. Heller

Neural codecs have become crucial to recent speech and audio generation research. In addition to signal compression capabilities, discrete codecs have also been found to enhance downstream training efficiency and compatibility with…