中文
相关论文

相关论文: Prompt-guided Precise Audio Editing with Diffusion…

200 篇论文

Text-driven video editing utilizing generative diffusion models has garnered significant attention due to their potential applications. However, existing approaches are constrained by the limited word embeddings provided in pre-training,…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Mingce Guo , Jingxuan He , Shengeng Tang , Zhangye Wang , Lechao Cheng

It is promising to design a single model that can suppress various distortions and improve speech quality, i.e., universal speech enhancement (USE). Compared to supervised learning-based predictive methods, diffusion-based generative models…

音频与语音处理 · 电气工程与系统科学 2025-06-02 Jie Zhang , Haoyin Yan , Xiaofei Li

This is the technique report for the winning solution of the CVPR2024 GenAI Media Generation Challenge Workshop's Instruction-guided Image Editing track. Instruction-guided image editing has been largely studied in recent years. The most…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Xuan Ju , Junhao Zhuang , Zhaoyang Zhang , Yuxuan Bian , Qiang Xu , Ying Shan

Text-to-image (T2I) diffusion models, with their impressive generative capabilities, have been adopted for image editing tasks, demonstrating remarkable efficacy. However, due to attention leakage and collision between the cross-attention…

计算机视觉与模式识别 · 计算机科学 2025-05-01 Xingxi Yin , Zhi Li , Jingfeng Zhang , Chenglin Li , Yin Zhang

Distributional shift is a central challenge in the deployment of machine learning models as they can be ill-equipped for real-world data. This is particularly evident in text-to-audio generation where the encoded representations are easily…

Text-based editing diffusion models exhibit limited performance when the user's input instruction is ambiguous. To solve this problem, we propose $\textit{Specify ANd Edit}$ (SANE), a zero-shot inference pipeline for diffusion-based editing…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Ekaterina Iakovleva , Fabio Pizzati , Philip Torr , Stéphane Lathuilière

Recently, diffusion models have emerged as a powerful class of generative models. Despite their success, there is still limited understanding of their semantic spaces. This makes it challenging to achieve precise and disentangled image…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Siyi Chen , Huijie Zhang , Minzhe Guo , Yifu Lu , Peng Wang , Qing Qu

Foundation models enable prompt-based classifiers for zero-shot and few-shot learning. Nonetheless, the conventional method of employing fixed prompts suffers from distributional shifts that negatively impact generalizability to unseen…

机器学习 · 计算机科学 2024-10-29 Yingjun Du , Gaowen Liu , Yuzhang Shang , Yuguang Yao , Ramana Kompella , Cees G. M. Snoek

Acoustic echo and background noise pose challenges on speech enhancement in hands-free systems and speakerphones. Discriminatively trained end-to-end methods represent a powerful solution for joint acoustic echo control (AEC) and denoising.…

音频与语音处理 · 电气工程与系统科学 2026-05-12 Haljan Lugo Girao , Ernst Seidel , Pejman Mowlaee , Ziyue Zhao , Tim Fingscheidt

Sound effect editing-modifying audio by adding, removing, or replacing elements-remains constrained by existing approaches that rely solely on low-level signal processing or coarse text prompts, often resulting in limited flexibility and…

多媒体 · 计算机科学 2025-11-27 Xinyue Guo , Xiaoran Yang , Lipan Zhang , Jianxuan Yang , Zhao Wang , Jian Luan

While diffusion models have achieved remarkable success in text-to-image generation, they encounter significant challenges with instruction-driven image editing. Our research highlights a key challenge: these models particularly struggle…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Yujia Hu , Songhua Liu , Zhenxiong Tan , Xingyi Yang , Xinchao Wang

Stutter removal is an essential scenario in the field of speech editing. However, when the speech recording contains stutters, the existing text-based speech editing approaches still suffer from: 1) the over-smoothing problem in the edited…

声音 · 计算机科学 2023-05-24 Ziyue Jiang , Qian Yang , Jialong Zuo , Zhenhui Ye , Rongjie Huang , Yi Ren , Zhou Zhao

Text-to-music generation technology is progressing rapidly, creating new opportunities for musical composition and editing. However, existing music editing methods often fail to preserve the source music's temporal structure, including…

声音 · 计算机科学 2025-11-19 Yi Yang , Haowen Li , Tianxiang Li , Boyu Cao , Xiaohan Zhang , Liqun Chen , Qi Liu

Precise spatial control in diffusion-based style transfer remains challenging. This challenge arises because diffusion models treat style as a global feature and lack explicit spatial grounding of style representations, making it difficult…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Bowen Chen , Jake Zuena , Alan C. Bovik , Divya Kothandaraman

Text-guided diffusion models have significantly advanced image editing, enabling highly realistic and local modifications based on textual prompts. While these developments expand creative possibilities, their malicious use poses…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Valentina Bazyleva , Nicolo Bonettini , Gaurav Bharaj

With the advent of diffusion models, Text-to-Image (T2I) generation has seen substantial advancements. Current T2I models allow users to specify object colors using linguistic color names, and some methods aim to personalize color-object…

图形学 · 计算机科学 2025-08-13 Qianru Qiu , Jiafeng Mao , Xueting Wang

In recent years, the burgeoning interest in diffusion models has led to significant advances in image and speech generation. Nevertheless, the direct synthesis of music waveforms from unrestricted textual prompts remains a relatively…

声音 · 计算机科学 2023-09-22 Pengfei Zhu , Chao Pang , Yekun Chai , Lei Li , Shuohuan Wang , Yu Sun , Hao Tian , Hua Wu

Diffusion models have become the go-to method for text-to-image generation, producing high-quality images from pure noise. However, the inner workings of diffusion models is still largely a mystery due to their black-box nature and complex,…

计算机视觉与模式识别 · 计算机科学 2026-01-23 Berk Tinaz , Zalan Fabian , Mahdi Soltanolkotabi

Text-to-image diffusion models are well-known for their ability to generate realistic images based on textual prompts. However, the existing works have predominantly focused on English, lacking support for non-English text-to-image models.…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Jian Ma , Chen Chen , Qingsong Xie , Haonan Lu

We study timbre transfer as an inference-time editing problem for music audio. Starting from a strong pre-trained latent diffusion model, we introduce a lightweight procedure that requires no additional training: (i) a dimension-wise noise…

声音 · 计算机科学 2026-01-29 Ching Ho Lee , Javier Nistal , Stefan Lattner , Marco Pasini , George Fazekas