English
Related papers

Related papers: A2SB: Audio-to-Audio Schrodinger Bridges

200 papers

This work introduces audio2chart, a framework for the automatic generation of Guitar Hero style charts directly from raw audio. The task is formalized as a sequence prediction problem, where models are trained to generate discrete chart…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-06 Riccardo Tripodi

Scaling spoken language modeling requires speech tokens that are both efficient and universal. Recent work has proposed syllables as promising speech tokens at low temporal resolution, but existing models are constrained to English and fail…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-02 Cheol Jun Cho , Nicholas Lee , Alan W Black , Gopala K. Anumanchipalli

In this paper, we present a vocoder-free framework for audio super-resolution that employs a flow matching generative model to capture the conditional distribution of complex-valued spectral coefficients. Unlike conventional two-stage…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-06 Woongjib Choi , Sangmin Lee , Hyungseob Lim , Hong-Goo Kang

We consider audio decoding as an inverse problem and solve it through diffusion posterior sampling. Explicit conditioning functions are developed for input signal measurements provided by an example of a transform domain perceptual audio…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-13 Pedro J. Villasana T. , Lars Villemoes , Janusz Klejsa , Per Hedelin

Signal-dependent beamformers are advantageous over signal-independent beamformers when the acoustic scenario - be it real-world or simulated - is straightforward in terms of the number of sound sources, the ambient sound field and their…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-01 Sina Hafezi , Alastair H. Moore , Pierre H. Guiraud , Patrick A. Naylor , Jacob Donley , Vladimir Tourbabin , Thomas Lunner

We propose a Vocos-based bandwidth extension model that enhances audio at 8-48 kHz by generating missing high-frequency content. Inputs are resampled to 48 kHz and processed by a neural vocoder backbone, enabling a single network to support…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-10 Yatharth Sharma

Language models have been successfully used to model natural signals, such as images, speech, and music. A key component of these models is a high quality neural compression model that can compress high-dimensional natural signals into…

Sound · Computer Science 2023-10-30 Rithesh Kumar , Prem Seetharaman , Alejandro Luebs , Ishaan Kumar , Kundan Kumar

Robust audio anti-spoofing has been increasingly challenging due to the recent advancements on deepfake techniques. While spectrograms have demonstrated their capability for anti-spoofing, complementary information presented in multi-order…

Sound · Computer Science 2024-10-04 Penghui Wen , Kun Hu , Wenxi Yue , Sen Zhang , Wanlei Zhou , Zhiyong Wang

Audio restoration consists in inverting degradations of a digital audio signal to recover what would have been the pristine quality signal before the degradation occurred. This is valuable in contexts such as archives of music recordings,…

Sound · Computer Science 2026-01-22 Carlos Hernandez-Olivan , Hendrik Vincent Koops , Hao Hao Tan , Elio Quinton

Audio-to-score alignment is an important pre-processing step for in-depth analysis of classical music. In this paper, we apply novel transposition-invariant audio features to this task. These low-dimensional features represent local pitch…

Sound · Computer Science 2018-07-20 Andreas Arzt , Stefan Lattner

Conventional spoofing detection systems have heavily relied on the use of handcrafted features derived from speech data. However, a notable shift has recently emerged towards the direct utilization of raw speech waveforms, as demonstrated…

Large-scale mobile communication systems tend to contain legacy transmission channels with narrowband bottlenecks, resulting in characteristic "telephone-quality" audio. While higher quality codecs exist, due to the scale and heterogeneity…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-12 Archit Gupta , Brendan Shillingford , Yannis Assael , Thomas C. Walters

Text-to-speech (TTS) systems offer the opportunity to compensate for a hearing loss at the source rather than correcting for it at the receiving end. This removes limitations such as time constraints for algorithms that amplify a sound in a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-23 Josef Schlittenlacher , Thomas Baer

Noise reduction techniques based on deep learning have demonstrated impressive performance in enhancing the overall quality of recorded speech. While these approaches are highly performant, their application in audio engineering can be…

Sound · Computer Science 2023-10-18 Christian J. Steinmetz , Thomas Walther , Joshua D. Reiss

Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrowed stereo image, especially when created in non-professional settings without specialized equipment…

Sound · Computer Science 2026-05-06 Jan Melechovsky , Ambuj Mehrish , Abhinaba Roy , Dorien Herremans

Audio-guided face reenactment aims to generate a photorealistic face that has matched facial expression with the input audio. However, current methods can only reenact a special person once the model is trained or need extra operations such…

Computer Vision and Pattern Recognition · Computer Science 2020-10-27 Jiangning Zhang , Xianfang Zeng , Chao Xu , Jun Chen , Yong Liu , Yunliang Jiang

Musicians and audio engineers sculpt and transform their sounds by connecting multiple processors, forming an audio processing graph. However, most deep-learning methods overlook this real-world practice and assume fixed graph settings. To…

Sound · Computer Science 2023-05-09 Sungho Lee , Jaehyun Park , Seungryeol Paik , Kyogu Lee

Multi-modal brain MRI provides essential complementary information for clinical diagnosis. However, acquiring all modalities in practice is often constrained by time and cost. To address this, various methods have been proposed to generate…

Image and Video Processing · Electrical Eng. & Systems 2026-05-07 Hanyeol Yang , Sunggyu Kim , Mi Kyung Kim , Yongseon Yoo , Yu-Mi Kim , Min-Ho Shin , Insung Chung , Sang Baek Koh , Hyeon Chang Kim , Jong-Min Lee

Audio language models have recently emerged as a promising approach for various audio generation tasks, relying on audio tokenizers to encode waveforms into sequences of discrete symbols. Audio tokenization often poses a necessary…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-11 Zhijun Liu , Shuai Wang , Sho Inoue , Qibing Bai , Haizhou Li

This paper describes an end-to-end (E2E) neural architecture for the audio rendering of small portions of display content on low resource personal computing devices. It is intended to address the problem of accessibility for vision-impaired…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-13 Liu Chen , Michael Deisher , Munir Georges