中文
相关论文

相关论文: A2SB: Audio-to-Audio Schrodinger Bridges

200 篇论文

This work introduces audio2chart, a framework for the automatic generation of Guitar Hero style charts directly from raw audio. The task is formalized as a sequence prediction problem, where models are trained to generate discrete chart…

音频与语音处理 · 电气工程与系统科学 2025-11-06 Riccardo Tripodi

Scaling spoken language modeling requires speech tokens that are both efficient and universal. Recent work has proposed syllables as promising speech tokens at low temporal resolution, but existing models are constrained to English and fail…

音频与语音处理 · 电气工程与系统科学 2026-02-02 Cheol Jun Cho , Nicholas Lee , Alan W Black , Gopala K. Anumanchipalli

In this paper, we present a vocoder-free framework for audio super-resolution that employs a flow matching generative model to capture the conditional distribution of complex-valued spectral coefficients. Unlike conventional two-stage…

音频与语音处理 · 电气工程与系统科学 2026-02-06 Woongjib Choi , Sangmin Lee , Hyungseob Lim , Hong-Goo Kang

We consider audio decoding as an inverse problem and solve it through diffusion posterior sampling. Explicit conditioning functions are developed for input signal measurements provided by an example of a transform domain perceptual audio…

音频与语音处理 · 电气工程与系统科学 2024-09-13 Pedro J. Villasana T. , Lars Villemoes , Janusz Klejsa , Per Hedelin

Signal-dependent beamformers are advantageous over signal-independent beamformers when the acoustic scenario - be it real-world or simulated - is straightforward in terms of the number of sound sources, the ambient sound field and their…

音频与语音处理 · 电气工程与系统科学 2023-12-01 Sina Hafezi , Alastair H. Moore , Pierre H. Guiraud , Patrick A. Naylor , Jacob Donley , Vladimir Tourbabin , Thomas Lunner

We propose a Vocos-based bandwidth extension model that enhances audio at 8-48 kHz by generating missing high-frequency content. Inputs are resampled to 48 kHz and processed by a neural vocoder backbone, enabling a single network to support…

音频与语音处理 · 电气工程与系统科学 2026-03-10 Yatharth Sharma

Language models have been successfully used to model natural signals, such as images, speech, and music. A key component of these models is a high quality neural compression model that can compress high-dimensional natural signals into…

声音 · 计算机科学 2023-10-30 Rithesh Kumar , Prem Seetharaman , Alejandro Luebs , Ishaan Kumar , Kundan Kumar

Robust audio anti-spoofing has been increasingly challenging due to the recent advancements on deepfake techniques. While spectrograms have demonstrated their capability for anti-spoofing, complementary information presented in multi-order…

声音 · 计算机科学 2024-10-04 Penghui Wen , Kun Hu , Wenxi Yue , Sen Zhang , Wanlei Zhou , Zhiyong Wang

Audio restoration consists in inverting degradations of a digital audio signal to recover what would have been the pristine quality signal before the degradation occurred. This is valuable in contexts such as archives of music recordings,…

声音 · 计算机科学 2026-01-22 Carlos Hernandez-Olivan , Hendrik Vincent Koops , Hao Hao Tan , Elio Quinton

Audio-to-score alignment is an important pre-processing step for in-depth analysis of classical music. In this paper, we apply novel transposition-invariant audio features to this task. These low-dimensional features represent local pitch…

声音 · 计算机科学 2018-07-20 Andreas Arzt , Stefan Lattner

Conventional spoofing detection systems have heavily relied on the use of handcrafted features derived from speech data. However, a notable shift has recently emerged towards the direct utilization of raw speech waveforms, as demonstrated…

Large-scale mobile communication systems tend to contain legacy transmission channels with narrowband bottlenecks, resulting in characteristic "telephone-quality" audio. While higher quality codecs exist, due to the scale and heterogeneity…

音频与语音处理 · 电气工程与系统科学 2019-07-12 Archit Gupta , Brendan Shillingford , Yannis Assael , Thomas C. Walters

Text-to-speech (TTS) systems offer the opportunity to compensate for a hearing loss at the source rather than correcting for it at the receiving end. This removes limitations such as time constraints for algorithms that amplify a sound in a…

音频与语音处理 · 电气工程与系统科学 2021-03-23 Josef Schlittenlacher , Thomas Baer

Noise reduction techniques based on deep learning have demonstrated impressive performance in enhancing the overall quality of recorded speech. While these approaches are highly performant, their application in audio engineering can be…

声音 · 计算机科学 2023-10-18 Christian J. Steinmetz , Thomas Walther , Joshua D. Reiss

Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrowed stereo image, especially when created in non-professional settings without specialized equipment…

声音 · 计算机科学 2026-05-06 Jan Melechovsky , Ambuj Mehrish , Abhinaba Roy , Dorien Herremans

Audio-guided face reenactment aims to generate a photorealistic face that has matched facial expression with the input audio. However, current methods can only reenact a special person once the model is trained or need extra operations such…

计算机视觉与模式识别 · 计算机科学 2020-10-27 Jiangning Zhang , Xianfang Zeng , Chao Xu , Jun Chen , Yong Liu , Yunliang Jiang

Musicians and audio engineers sculpt and transform their sounds by connecting multiple processors, forming an audio processing graph. However, most deep-learning methods overlook this real-world practice and assume fixed graph settings. To…

声音 · 计算机科学 2023-05-09 Sungho Lee , Jaehyun Park , Seungryeol Paik , Kyogu Lee

Multi-modal brain MRI provides essential complementary information for clinical diagnosis. However, acquiring all modalities in practice is often constrained by time and cost. To address this, various methods have been proposed to generate…

图像与视频处理 · 电气工程与系统科学 2026-05-07 Hanyeol Yang , Sunggyu Kim , Mi Kyung Kim , Yongseon Yoo , Yu-Mi Kim , Min-Ho Shin , Insung Chung , Sang Baek Koh , Hyeon Chang Kim , Jong-Min Lee

Audio language models have recently emerged as a promising approach for various audio generation tasks, relying on audio tokenizers to encode waveforms into sequences of discrete symbols. Audio tokenization often poses a necessary…

音频与语音处理 · 电气工程与系统科学 2024-06-11 Zhijun Liu , Shuai Wang , Sho Inoue , Qibing Bai , Haizhou Li

This paper describes an end-to-end (E2E) neural architecture for the audio rendering of small portions of display content on low resource personal computing devices. It is intended to address the problem of accessibility for vision-impaired…

音频与语音处理 · 电气工程与系统科学 2023-03-13 Liu Chen , Michael Deisher , Munir Georges