English
Related papers

Related papers: Text Conditioned Symbolic Drumbeat Generation usin…

200 papers

Recently, there has been a growing interest in text-to-speech (TTS) methods that can be trained with minimal supervision by combining two types of discrete speech representations and using two sequence-to-sequence tasks to decouple TTS.…

Sound · Computer Science 2023-12-19 Chunyu Qiang , Hao Li , Hao Ni , He Qu , Ruibo Fu , Tao Wang , Longbiao Wang , Jianwu Dang

Masked Diffusion Models (MDMs) provide an efficient non-causal alternative to autoregressive generation but often struggle with token dependencies and semantic incoherence due to their reliance on discrete marginal distributions. We address…

Computation and Language · Computer Science 2026-04-20 Roy Uziel , Omer Belhasin , Itay Levy , Akhiad Bercovich , Ran El-Yaniv , Ran Zilberstein , Michael Elad

Generating image variations, where a model produces variations of an input image while preserving the semantic context has gained increasing attention. Current image variation techniques involve adapting a text-to-image model to reconstruct…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Manoj Kumar , Neil Houlsby , Emiel Hoogeboom

Autoregressive language models decode left-to-right with irreversible commitments, limiting revision during multi-step reasoning. We propose \textbf{VDLM}, a modular variable diffusion language model that separates semantic planning from…

Computation and Language · Computer Science 2026-02-19 Shuhui Qu

Objective: While recent advances in text-conditioned generative models have enabled the synthesis of realistic medical images, progress has been largely confined to 2D modalities such as chest X-rays. Extending text-to-image generation to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Daniele Molino , Camillo Maria Caruso , Filippo Ruffini , Paolo Soda , Valerio Guarrasi

Dancing with music is always an essential human art form to express emotion. Due to the high temporal-spacial complexity, long-term 3D realist dance generation synchronized with music is challenging. Existing methods suffer from the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Siqi Yang , Zejun Yang , Zhisheng Wang

Text-to-Image synthesis is the task of generating an image according to a specific text description. Generative Adversarial Networks have been considered the standard method for image synthesis virtually since their introduction. Denoising…

Computer Vision and Pattern Recognition · Computer Science 2023-05-18 Konstantina Nikolaidou , George Retsinas , Vincent Christlein , Mathias Seuret , Giorgos Sfikas , Elisa Barney Smith , Hamam Mokayed , Marcus Liwicki

Our goal is to train a generative model of 3D hand motions, conditioned on natural language descriptions specifying motion characteristics such as handshapes, locations, finger/hand/arm movements. To this end, we automatically build pairs…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Léore Bensabath , Mathis Petrovich , Gül Varol

Scene text detection techniques have garnered significant attention due to their wide-ranging applications. However, existing methods have a high demand for training data, and obtaining accurate human annotations is labor-intensive and…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Ling Fu , Zijie Wu , Yingying Zhu , Yuliang Liu , Xiang Bai

Scene-text image synthesis techniques that aim to naturally compose text instances on background scene images are very appealing for training deep neural networks due to their ability to provide accurate and comprehensive annotation…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 Zhengmi Tang , Tomo Miyazaki , Shinichiro Omachi

Breakthroughs in text-to-music generation models are transforming the creative landscape, equipping musicians with innovative tools for composition and experimentation like never before. However, controlling the generation process to…

Sound · Computer Science 2025-06-19 Teysir Baoueb , Xiaoyu Bie , Xi Wang , Gaël Richard

Text-conditioned image generation has gained significant attention in recent years and are processing increasingly longer and comprehensive text prompt. In everyday life, dense and intricate text appears in contexts like advertisements,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Alex Jinpeng Wang , Dongxing Mao , Jiawei Zhang , Weiming Han , Zhuobai Dong , Linjie Li , Yiqi Lin , Zhengyuan Yang , Libo Qin , Fuwei Zhang , Lijuan Wang , Min Li

Prior material creation methods had limitations in producing diverse results mainly because reconstruction-based methods relied on real-world measurements and generation-based methods were trained on relatively small material datasets. To…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Linxuan Xin , Zheng Zhang , Jinfu Wei , Wei Gao , Duan Gao

Spurred by the potential of deep learning, computational music generation has gained renewed academic interest. A crucial issue in music generation is that of user control, especially in scenarios where the music generation process is…

Sound · Computer Science 2019-08-05 Stefan Lattner , Maarten Grachten

Despite remarkable progress in autoregressive language models, alternative generative paradigms beyond left-to-right generation are still being actively explored. Discrete diffusion models, with the capacity for parallel generation, have…

Computation and Language · Computer Science 2025-03-10 Minkai Xu , Tomas Geffner , Karsten Kreis , Weili Nie , Yilun Xu , Jure Leskovec , Stefano Ermon , Arash Vahdat

We present a novel alignment-before-generation approach to tackle the challenging task of generating general 3D shapes based on 2D images or texts. Directly learning a conditional generative model from images or texts to 3D shapes is prone…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Zibo Zhao , Wen Liu , Xin Chen , Xianfang Zeng , Rui Wang , Pei Cheng , Bin Fu , Tao Chen , Gang Yu , Shenghua Gao

We present SoundLoCD, a novel text-to-sound generation framework, which incorporates a LoRA-based conditional discrete contrastive latent diffusion model. Unlike recent large-scale sound generation models, our model can be efficiently…

Sound · Computer Science 2024-05-27 Xinlei Niu , Jing Zhang , Christian Walder , Charles Patrick Martin

Diffusion models have revolutionized image generation, yet several challenges restrict their application to large-image domains, such as digital pathology and satellite imagery. Given that it is infeasible to directly train a model on…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Srikar Yellapragada , Alexandros Graikos , Kostas Triaridis , Prateek Prasanna , Rajarsi R. Gupta , Joel Saltz , Dimitris Samaras

Transformer-based Language Models (LMs) have achieved impressive results on natural language understanding tasks, but they can also generate toxic text such as insults, threats, and profanity, limiting their real-world applications. To…

Computation and Language · Computer Science 2023-07-06 Jin Myung Kwak , Minseon Kim , Sung Ju Hwang

The generation of realistic medical images from text descriptions has significant potential to address data scarcity challenges in healthcare AI while preserving patient privacy. This paper presents a comprehensive study of text-to-image…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Mikhail Chaichuk , Sushant Gautam , Steven Hicks , Elena Tutubalina