中文
相关论文

相关论文: MEDIC: Zero-shot Music Editing with Disentangled I…

200 篇论文

Text-guided image generation and editing using diffusion models have achieved remarkable advancements. Among these, tuning-free methods have gained attention for their ability to perform edits without extensive model adjustments, offering…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Wenyi Mo , Tianyu Zhang , Yalong Bai , Bing Su , Ji-Rong Wen

Editing images with diffusion models under strict training-free constraints remains a significant challenge. While recent optimisation-based methods achieve strong zero-shot edits from text, they struggle to preserve identity and capture…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Niki Foteinopoulou , Ignas Budvytis , Stephan Liwicki

The human voice conveys not just words but also emotional states and individuality. Emotional voice conversion (EVC) modifies emotional expressions while preserving linguistic content and speaker identity, improving applications like…

音频与语音处理 · 电气工程与系统科学 2025-09-29 Hsing-Hang Chou , Yun-Shao Lin , Ching-Chin Sung , Yu Tsao , Chi-Chun Lee

Controllable music generation methods are critical for human-centered AI-based music creation, but are currently limited by speed, quality, and control design trade-offs. Diffusion Inference-Time T-optimization (DITTO), in particular,…

声音 · 计算机科学 2024-05-31 Zachary Novack , Julian McAuley , Taylor Berg-Kirkpatrick , Nicholas Bryan

Diffusion-based image editing is a composite process of preserving the source image content and generating new content or applying modifications. While current editing approaches have made improvements under text guidance, most of them have…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Tianrui Huang , Pu Cao , Lu Yang , Chun Liu , Mengjie Hu , Zhiwei Liu , Qing Song

Given the remarkable results of motion synthesis with diffusion models, a natural question arises: how can we effectively leverage these models for motion editing? Existing diffusion-based motion editing methods overlook the profound…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Sigal Raab , Inbar Gat , Nathan Sala , Guy Tevet , Rotem Shalev-Arkushin , Ohad Fried , Amit H. Bermano , Daniel Cohen-Or

One-shot controllable video editing (OCVE) is an important yet challenging task, aiming to propagate user edits that are made -- using any image editing tool -- on the first frame of a video to all subsequent frames, while ensuring content…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Zhengbo Zhang , Yuxi Zhou , Duo Peng , Joo-Hwee Lim , Zhigang Tu , De Wen Soh , Lin Geng Foo

The recent proliferation of diffusion models has made style mimicry effortless, enabling users to imitate unique artistic styles without authorization. In deployed platforms, this raises copyright and intellectual-property risks and calls…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Tong Zhang , Ru Zhang , Jianyi Liu

Sparse-view computed tomography (CT) reconstruction is fundamentally challenging due to undersampling, leading to an ill-posed inverse problem. Traditional iterative methods incorporate handcrafted or learned priors to regularize the…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Leon Suarez-Rodriguez , Roman Jacome , Romario Gualdron-Hurtado , Ana Mantilla-Dulcey , Henry Arguello

The advent of Video Diffusion Transformers (Video DiTs) marks a milestone in video generation. However, directly applying existing video editing methods to Video DiTs often incurs substantial computational overhead, due to…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Lingling Cai , Kang Zhao , Hangjie Yuan , Xiang Wang , Yingya Zhang , Kejie Huang

Diffusion models have shown remarkable progress in text-to-audio generation. However, text-guided audio editing remains in its early stages. This task focuses on modifying the target content within an audio signal while preserving the rest,…

声音 · 计算机科学 2026-04-17 Liting Gao , Yi Yuan , Yaru Chen , Yuelan Cheng , Zhenbo Li , Juan Wen , Shubin Zhang , Wenwu Wang

Balancing fidelity and editability is essential in text-based image editing (TIE), where failures commonly lead to over- or under-editing issues. Existing methods typically rely on attention injections for structure preservation and…

计算机视觉与模式识别 · 计算机科学 2025-04-09 Qi Mao , Lan Chen , Yuchao Gu , Mike Zheng Shou , Ming-Hsuan Yang

Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speaking style of the source speaker or mimic the distinctive…

声音 · 计算机科学 2025-06-02 Kaidi Wang , Wenhao Guan , Ziyue Jiang , Hukai Huang , Peijie Chen , Weijie Wu , Qingyang Hong , Lin Li

Diffusion-based Image Editing (DIE) is an emerging research hot-spot, which often applies a semantic mask to control the target area for diffusion-based editing. However, most existing solutions obtain these masks via manual operations or…

计算机视觉与模式识别 · 计算机科学 2024-01-24 Siyu Zou , Jiji Tang , Yiyi Zhou , Jing He , Chaoyi Zhao , Rongsheng Zhang , Zhipeng Hu , Xiaoshuai Sun

Audio source separation is fundamental for machines to understand complex acoustic environments and underpins numerous audio applications. Current supervised deep learning approaches, while powerful, are limited by the need for extensive,…

Reducing the radiation dose in computed tomography (CT) is important to mitigate radiation-induced risks. One option is to employ a well-trained model to compensate for incomplete information and map sparse-view measurements to the CT…

图像与视频处理 · 电气工程与系统科学 2023-03-29 Xiaoyue Li , Kai Shang , Gaoang Wang , Mark D. Butala

We propose a unified model for three inter-related tasks: 1) to \textit{separate} individual sound sources from a mixed music audio, 2) to \textit{transcribe} each sound source to MIDI notes, and 3) to\textit{ synthesize} new pieces based…

声音 · 计算机科学 2021-08-10 Liwei Lin , Qiuqiang Kong , Junyan Jiang , Gus Xia

This study proposes a zero-shot image segmentation framework for detecting erythema (redness of the skin) using edit-friendly inversion in diffusion models. The method synthesizes reference images of the same patient that are free from…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Konstantinos Moutselos , Ilias Maglogiannis

Text-to-image diffusion models have achieved remarkable success in generating high-quality and diverse images. Building on these advancements, diffusion models have also demonstrated exceptional performance in text-guided image editing. A…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Mingyu Kang , Yong Suk Choi

Recently, diffusion-based generative models have achieved remarkable success for image generation and edition. However, existing diffusion-based video editing approaches lack the ability to offer precise control over generated content that…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Paul Couairon , Clément Rambour , Jean-Emmanuel Haugeard , Nicolas Thome