中文
相关论文

相关论文: Mustango: Toward Controllable Text-to-Music Genera…

200 篇论文

Text-to-image diffusion model is a popular paradigm that synthesizes personalized images by providing a text prompt and a random Gaussian noise. While people observe that some noises are ``golden noises'' that can achieve better text-image…

机器学习 · 计算机科学 2025-07-18 Zikai Zhou , Shitong Shao , Lichen Bai , Shufei Zhang , Zhiqiang Xu , Bo Han , Zeke Xie

In music-driven dance motion generation, most existing methods use hand-crafted features and neglect that music foundation models have profoundly impacted cross-modal content generation. To bridge this gap, we propose a diffusion-based…

声音 · 计算机科学 2025-02-28 Xinran Liu , Zhenhua Feng , Diptesh Kanojia , Wenwu Wang

Text-to-Image (T2I) diffusion/flow models have recently achieved remarkable progress in visual fidelity and text alignment. However, they remain limited when users need to precisely control image layouts, something that natural language…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Amadou S. Sangare , Adrien Maglo , Mohamed Chaouch , Bertrand Luvison

Dance plays an important role as an artistic form and expression in human culture, yet automatically generating dance sequences is a significant yet challenging endeavor. Existing approaches often neglect the critical aspect of…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Hongsong Wang , Ying Zhu , Xin Geng , Liang Wang

The diffusion model has been proven a powerful generative model in recent years, yet remains a challenge in generating visual text. Several methods alleviated this issue by incorporating explicit text position and content as guidance on…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Jingye Chen , Yupan Huang , Tengchao Lv , Lei Cui , Qifeng Chen , Furu Wei

We propose the first joint audio-video generation framework that brings engaging watching and listening experiences simultaneously, towards high-quality realistic videos. To generate joint audio-video pairs, we propose a novel Multi-Modal…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Ludan Ruan , Yiyang Ma , Huan Yang , Huiguo He , Bei Liu , Jianlong Fu , Nicholas Jing Yuan , Qin Jin , Baining Guo

Text-to-image synthesis has achieved high-quality results with recent advances in diffusion models. However, text input alone has high spatial ambiguity and limited user controllability. Most existing methods allow spatial control through…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Yuki Endo

Diffusion models have gained increasing attention for their impressive generation abilities but currently struggle with rendering accurate and coherent text. To address this issue, we introduce TextDiffuser, focusing on generating images…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Jingye Chen , Yupan Huang , Tengchao Lv , Lei Cui , Qifeng Chen , Furu Wei

Generating the motion of orchestral conductors from a given piece of symphony music is a challenging task since it requires a model to learn semantic music features and capture the underlying distribution of real conducting motion. Prior…

音频与语音处理 · 电气工程与系统科学 2023-11-14 Zhuoran Zhao , Jinbin Bai , Delong Chen , Debang Wang , Yubo Pan

Text-to-image generation has made remarkable progress with the emergence of diffusion models. However, it is still a difficult task to generate images for street views based on text, mainly because the road topology of street scenes is…

计算机视觉与模式识别 · 计算机科学 2024-02-08 Jinming Su , Songen Gu , Yiting Duan , Xingyue Chen , Junfeng Luo

We introduce Reflectance Diffusion, a new neural text-to-texture model capable of generating high-fidelity SVBRDF maps from textual descriptions. Our method leverages a tandem neural approach, consisting of two modules, to accurately model…

图形学 · 计算机科学 2026-01-21 Bowen Xue , Giuseppe Claudio Guarnera , Shuang Zhao , Zahra Montazeri

While modern text-to-image diffusion models generate high-fidelity images, they offer limited control over the spatial and geometric structure of the output. To address this, we introduce and evaluate two ControlNets specialized for…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Julien Boudier , Hugo Caselles-Dupré

Recent advancements in diffusion models have showcased their impressive capacity to generate visually striking images. Nevertheless, ensuring a close match between the generated image and the given prompt remains a persistent challenge. In…

计算机视觉与模式识别 · 计算机科学 2023-09-11 Yupeng Zhou , Daquan Zhou , Zuo-Liang Zhu , Yaxing Wang , Qibin Hou , Jiashi Feng

The design of diffusion-based audio generation systems has been investigated from diverse perspectives, such as data space, network architecture, and conditioning techniques, while most of these innovations require model re-training. In…

声音 · 计算机科学 2026-04-10 Junyou Wang , Zehua Chen , Binjie Yuan , Kaiwen Zheng , Chang Li , Yuxuan Jiang , Jun Zhu

Recommender systems have become indispensable in music streaming services, enhancing user experiences by personalizing playlists and facilitating the serendipitous discovery of new music. However, the existing recommender systems overlook…

信息检索 · 计算机科学 2023-08-29 Yunhak Oh , Sukwon Yun , Dongmin Hyun , Sein Kim , Chanyoung Park

In recent years, image generation has shown a great leap in performance, where diffusion models play a central role. Although generating high-quality images, such models are mainly conditioned on textual descriptions. This begs the…

声音 · 计算机科学 2023-05-23 Guy Yariv , Itai Gat , Lior Wolf , Yossi Adi , Idan Schwartz

To enhance the controllability of text-to-image diffusion models, current ControlNet-like models have explored various control signals to dictate image attributes. However, existing methods either handle conditions inefficiently or use a…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Qingdong He , Jinlong Peng , Pengcheng Xu , Boyuan Jiang , Xiaobin Hu , Donghao Luo , Yong Liu , Yabiao Wang , Chengjie Wang , Xiangtai Li , Jiangning Zhang

Most music generation models directly generate a single music mixture. To allow for more flexible and controllable generation, the Multi-Source Diffusion Model (MSDM) has been proposed to model music as a mixture of multiple instrumental…

音频与语音处理 · 电气工程与系统科学 2025-06-18 Zhongweiyang Xu , Debottam Dutta , Yu-Lin Wei , Romit Roy Choudhury

We introduce a film score generation framework to harmonize visual pixels and music melodies utilizing a latent diffusion model. Our framework processes film clips as input and generates music that aligns with a general theme while offering…

多媒体 · 计算机科学 2024-11-13 F. Qi , L. Ni , C. Xu

Fine-tuning large-scale text-to-video diffusion models to add new generative controls, such as those over physical camera parameters (e.g., shutter speed or aperture), typically requires vast, high-fidelity datasets that are difficult to…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Shihan Cheng , Nilesh Kulkarni , David Hyde , Dmitriy Smirnov