中文
相关论文

相关论文: Improving Inference-Time Optimisation for Vocal Ef…

200 篇论文

In analyses with severe data-limitations, augmenting the target dataset with information from ancillary datasets in the application domain, called source datasets, can lead to significantly improved statistical procedures. However, existing…

机器学习 · 统计学 2025-08-01 Nathan Wycoff , Ali Arab , Lisa O. Singh

Recent studies show that pretrained vision models can boost performance in audio downstream tasks. To enhance the performance further, an additional pretraining stage with large scale audio data is typically required to infuse audio…

声音 · 计算机科学 2024-12-10 Juan Yeo , Jinkwan Jang , Kyubyung Chae , Seongkyu Mun , Taesup Kim

Test-time adaptation (TTA) aims to boost the generalization capability of a trained model by conducting self-/unsupervised learning during the testing phase. While most existing TTA methods for video primarily utilize visual supervisory…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Runhao Zeng , Qi Deng , Ronghao Zhang , Shuaicheng Niu , Jian Chen , Xiping Hu , Victor C. M. Leung

We study timbre transfer as an inference-time editing problem for music audio. Starting from a strong pre-trained latent diffusion model, we introduce a lightweight procedure that requires no additional training: (i) a dimension-wise noise…

声音 · 计算机科学 2026-01-29 Ching Ho Lee , Javier Nistal , Stefan Lattner , Marco Pasini , George Fazekas

Current deep learning techniques for style transfer would not be optimal for design support since their "one-shot" transfer does not fit exploratory design processes. To overcome this gap, we propose parametric transcription, which…

机器学习 · 计算机科学 2021-05-20 Hiromu Yakura , Yuki Koyama , Masataka Goto

Semantic noise initialization has been reported to improve robustness and controllability in image diffusion models. Whether these gains transfer to text-to-video (T2V) generation remains unclear, since temporal coupling can introduce extra…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Yixiao Jing , Chaoyu Zhang , Zixuan Zhong , Peizhou Huang

Direct speech-to-speech translation (S2ST) with discrete self-supervised representations has achieved remarkable accuracy, but is unable to preserve the speaker timbre of the source speech. Meanwhile, the scarcity of high-quality…

声音 · 计算机科学 2024-07-22 Yongqi Wang , Jionghao Bai , Rongjie Huang , Ruiqi Li , Zhiqing Hong , Zhou Zhao

Existing deep learning (DL) based speech enhancement approaches are generally optimised to minimise the distance between clean and enhanced speech features. These often result in improved speech quality however they suffer from a lack of…

声音 · 计算机科学 2021-11-19 Tassadaq Hussain , Mandar Gogate , Kia Dashtipour , Amir Hussain

The aim of this paper is to develop tractable large deviation approximations for the empirical measure of a small noise diffusion. The starting point is the Freidlin-Wentzell theory, which shows how to approximate via a large deviation…

概率论 · 数学 2021-01-11 Paul Dupuis , Guo-Jhen Wu

A recently published method for audio style transfer has shown how to extend the process of image style transfer to audio. This method synthesizes audio "content" and "style" independently using the magnitudes of a short time Fourier…

声音 · 计算机科学 2017-12-01 Parag K. Mital

Current emotional Text-To-Speech (TTS) and style transfer methods rely on reference encoders to control global style or emotion vectors, but do not capture nuanced acoustic details of the reference speech. To this end, we propose a novel…

声音 · 计算机科学 2025-10-03 Jianing Yang , Sheng Li , Takahiro Shinozaki , Yuki Saito , Hiroshi Saruwatari

Recent deep learning models have achieved high performance in speech enhancement; however, it is still challenging to obtain a fast and low-complexity model without significant performance degradation. Previous knowledge distillation…

声音 · 计算机科学 2022-11-01 Wooseok Shin , Hyun Joon Park , Jin Sob Kim , Byung Hoon Lee , Sung Won Han

In this paper, we explore the visual representations produced from a pre-trained text-to-video (T2V) diffusion model for video understanding tasks. We hypothesize that the latent representation learned from a pretrained generative T2V model…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Zixin Zhu , Xuelu Feng , Dongdong Chen , Junsong Yuan , Chunming Qiao , Gang Hua

Recent research has demonstrated impressive results in video-to-speech synthesis which involves reconstructing speech solely from visual input. However, previous works have struggled to accurately synthesize speech due to a lack of…

声音 · 计算机科学 2023-08-16 Jeongsoo Choi , Joanna Hong , Yong Man Ro

We propose an end-to-end music mixing style transfer system that converts the mixing style of an input multitrack to that of a reference song. This is achieved with an encoder pre-trained with a contrastive objective to extract only audio…

音频与语音处理 · 电气工程与系统科学 2023-04-12 Junghyun Koo , Marco A. Martínez-Ramírez , Wei-Hsiang Liao , Stefan Uhlich , Kyogu Lee , Yuki Mitsufuji

The predominant approach to advancing text-to-image generation has been training-time scaling, where larger models are trained on more data using greater computational resources. While effective, this approach is computationally expensive,…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Shufan Li , Konstantinos Kallidromitis , Akash Gokul , Arsh Koneru , Yusuke Kato , Kazuki Kozuka , Aditya Grover

We present LongCat-AudioDiT, a novel, non-autoregressive diffusion-based text-to-speech (TTS) model that achieves state-of-the-art (SOTA) performance. Unlike previous methods that rely on intermediate acoustic representations such as…

声音 · 计算机科学 2026-04-01 Detai Xin , Shujie Hu , Chengzuo Yang , Chen Huang , Guoqiao Yu , Guanglu Wan , Xunliang Cai

In the context of environmental sound classification, the adaptability of systems is key: which sound classes are interesting depends on the context and the user's needs. Recent advances in text-to-audio retrieval allow for zero-shot audio…

声音 · 计算机科学 2023-08-21 Saksham Singh Kushwaha , Magdalena Fuentes

With the demand for autonomous control and personalized speech generation, the style control and transfer in Text-to-Speech (TTS) is becoming more and more important. In this paper, we propose a new TTS system that can perform style…

声音 · 计算机科学 2023-07-12 Wenhao Guan , Tao Li , Yishuang Li , Hukai Huang , Qingyang Hong , Lin Li

Diffusion models have established the state-of-the-art in text-to-image generation, but their performance often relies on a diffusion prior network to translate text embeddings into the visual manifold for easier decoding. These priors are…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Samuele Dell'Erba , Andrew D. Bagdanov