中文
相关论文

相关论文: D3PIA: A Discrete Denoising Diffusion Model for Pi…

200 篇论文

We contribute a pop-song automation framework for lead melody generation and accompaniment arrangement. The framework reflects the major procedures of human music composition, generating both lead melody and piano accompaniment by a unified…

声音 · 计算机科学 2018-12-31 Ziyu Wang , Gus Xia

Discrete diffusion models are a class of generative models that produce samples from an approximated data distribution within a discrete state space. Often, there is a need to target specific regions of the data distribution. Current…

机器学习 · 计算机科学 2025-09-03 Cheuk Kit Lee , Paul Jeha , Jes Frellsen , Pietro Lio , Michael Samuel Albergo , Francisco Vargas

Interior design is a complex and creative discipline involving aesthetics, functionality, ergonomics, and materials science. Effective solutions must meet diverse requirements, typically producing multiple deliverables such as renderings…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Yuxuan Yang , Tao Geng

Diffusion-based text-to-image models have achieved remarkable results in synthesizing diverse images from text prompts and can capture specific artistic styles via style personalization. However, their entangled latent space and lack of…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Jaehyun Lee , Wonhark Park , Wonsik Shin , Hyunho Lee , Hyoung Min Na , Nojun Kwak

Diffusion-based generative models have achieved remarkable success in image generation. Their guidance formulation allows an external model to plug-and-play control the generation process for various tasks without finetuning the diffusion…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Hyojun Go , Yunsung Lee , Jin-Young Kim , Seunghyun Lee , Myeongho Jeong , Hyun Seung Lee , Seungtaek Choi

The paper presents a novel approach for vector-floorplan generation via a diffusion model, which denoises 2D coordinates of room/door corners with two inference objectives: 1) a single-step noise as the continuous quantity to precisely…

计算机视觉与模式识别 · 计算机科学 2022-11-28 Mohammad Amin Shabani , Sepidehsadat Hosseini , Yasutaka Furukawa

Handwriting stroke generation is crucial for improving the performance of tasks such as handwriting recognition and writers order recovery. In handwriting stroke generation, it is significantly important to imitate the sample calligraphic…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Sidra Hanif , Longin Jan Latecki

Generating visual layouts is an essential ingredient of graphic design. The ability to condition layout generation on a partial subset of component attributes is critical to real-world applications that involve user interaction. Recently,…

计算机视觉与模式识别 · 计算机科学 2023-03-08 Elad Levi , Eli Brosh , Mykola Mykhailych , Meir Perez

Score-based generative models and diffusion probabilistic models have been successful at generating high-quality samples in continuous domains such as images and audio. However, due to their Langevin-inspired sampling mechanisms, their…

声音 · 计算机科学 2021-11-29 Gautam Mittal , Jesse Engel , Curtis Hawthorne , Ian Simon

Score distillation of 2D diffusion models has proven to be a powerful mechanism to guide 3D optimization, for example enabling text-based 3D generation or single-view reconstruction. A common limitation of existing score distillation…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Yanbo Xu , Jayanth Srinivasa , Gaowen Liu , Shubham Tulsiani

Adapting learning materials to the level of skill of a student is important in education. In the context of music training, one essential ability is sight-reading -- playing unfamiliar scores at first sight -- which benefits from…

声音 · 计算机科学 2025-09-23 Pedro Ramoneda , Masahiro Suzuki , Akira Maezawa , Xavier Serra

We present SlotAdapt, an object-centric learning method that combines slot attention with pretrained diffusion models by introducing adapters for slot-based conditioning. Our method preserves the generative power of pretrained diffusion…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Adil Kaan Akan , Yucel Yemez

Deep generative models dominate the existing literature in layout pattern generation. However, leaving the guarantee of legality to an inexplicable neural network could be problematic in several applications. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Zixiao Wang , Yunheng Shen , Wenqian Zhao , Yang Bai , Guojin Chen , Farzan Farnia , Bei Yu

Diffusion models have emerged as a promising approach for text generation, with recent works falling into two main categories: discrete and continuous diffusion models. Discrete diffusion models apply token corruption independently using…

计算与语言 · 计算机科学 2025-05-29 Bocheng Li , Zhujin Gao , Linli Xu

This paper introduces a novel method for emulating piano sounds. We propose to exploit the sines, transient, and noise decomposition to design a differentiable spectral modeling synthesizer replicating piano notes. Three sub-modules learn…

声音 · 计算机科学 2025-02-04 Riccardo Simionato , Stefano Fasciani

Diffusion probabilistic models (DPMs) have become a popular approach to conditional generation, due to their promising results and support for cross-modal synthesis. A key desideratum in conditional synthesis is to achieve high…

计算机视觉与模式识别 · 计算机科学 2023-02-17 Ye Zhu , Yu Wu , Kyle Olszewski , Jian Ren , Sergey Tulyakov , Yan Yan

Existing music generation models are mostly language-based, neglecting the frequency continuity property of notes, resulting in inadequate fitting of rare or never-used notes and thus reducing the diversity of generated samples. We argue…

声音 · 计算机科学 2024-08-06 Shipei Liu , Xiaoya Fan , Guowei Wu

Music arrangement generation is a subtask of automatic music generation, which involves reconstructing and re-conceptualizing a piece with new compositional techniques. Such a generation process inevitably requires reference from the…

声音 · 计算机科学 2020-08-18 Ziyu Wang , Ke Chen , Junyan Jiang , Yiyi Zhang , Maoran Xu , Shuqi Dai , Xianbin Gu , Gus Xia

We introduce a novel resampling criterion using lift scores, for improving compositional generation in diffusion models. By leveraging the lift scores, we evaluate whether generated samples align with each single condition and then compose…

机器学习 · 计算机科学 2025-05-27 Chenning Yu , Sicun Gao

Diffusion models have demonstrated exceptional performances in various fields of generative modeling, but suffer from slow sampling speed due to their iterative nature. While this issue is being addressed in continuous domains, discrete…

机器学习 · 计算机科学 2025-05-12 Satoshi Hayakawa , Yuhta Takida , Masaaki Imaizumi , Hiromi Wakaki , Yuki Mitsufuji