中文
相关论文

相关论文: C3Net: Compound Conditioned ControlNet for Multimo…

200 篇论文

The control of nonlinear systems with unknown dynamics has been a significant field of research for many years. This paper presents a novel data-driven optimal adaptive control structure with less control effort and faster adaptation than…

系统与控制 · 电气工程与系统科学 2022-06-28 Mohammad Mahmoudi , Nasser Sadati

Recent image generation approaches often address subject, style, and structure-driven conditioning in isolation, leading to feature entanglement and limited task transferability. In this paper, we introduce 3SGen, a task-aware unified…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Xinyang Song , Libin Wang , Weining Wang , Zhiwei Li , Jianxin Sun , Dandan Zheng , Jingdong Chen , Qi Li , Zhenan Sun

Autoregressive conditional image generation algorithms are capable of generating photorealistic images that are consistent with given textual or image conditions, and have great potential for a wide range of applications. Nevertheless, the…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Qiaoying Qu , Shiyu Shen

Medication recommendation targets to provide a proper set of medicines according to patients' diagnoses, which is a critical task in clinics. Currently, the recommendation is manually conducted by doctors. However, for complicated cases,…

机器学习 · 计算机科学 2022-02-21 Rui Wu , Zhaopeng Qiu , Jiacheng Jiang , Guilin Qi , Xian Wu

The goal of exemplar-based texture synthesis is to generate texture images that are visually similar to a given exemplar. Recently, promising results have been reported by methods relying on convolutional neural networks (ConvNets)…

计算机视觉与模式识别 · 计算机科学 2019-12-18 Zi-Ming Wang , Meng-Han Li , Gui-Song Xia

Most existing neural network models for music generation use recurrent neural networks. However, the recent WaveNet model proposed by DeepMind shows that convolutional neural networks (CNNs) can also generate realistic musical waveforms in…

声音 · 计算机科学 2017-07-19 Li-Chia Yang , Szu-Yu Chou , Yi-Hsuan Yang

Automated retinal image medical description generation is crucial for streamlining medical diagnosis and treatment planning. Existing challenges include the reliance on learned retinal image representations, difficulties in handling…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Nagur Shareef Shaik , Teja Krishna Cherukuri , Dong Hye Ye

Text-to-music generation models are now capable of generating high-quality music audio in broad styles. However, text control is primarily suitable for the manipulation of global musical attributes like genre, mood, and tempo, and is less…

声音 · 计算机科学 2023-11-14 Shih-Lun Wu , Chris Donahue , Shinji Watanabe , Nicholas J. Bryan

Mobile robots and autonomous vehicles rely on multi-modal sensor setups to perceive and understand their surroundings. Aside from cameras, LiDAR sensors represent a central component of state-of-the-art perception systems. In addition to…

计算机视觉与模式识别 · 计算机科学 2018-04-27 Florian Piewak , Peter Pinggera , Manuel Schäfer , David Peter , Beate Schwarz , Nick Schneider , David Pfeiffer , Markus Enzweiler , Marius Zöllner

We introduce MeronymNet, a novel hierarchical approach for controllable, part-based generation of multi-category objects using a single unified model. We adopt a guided coarse-to-fine strategy involving semantically conditioned generation…

计算机视觉与模式识别 · 计算机科学 2021-10-19 Rishabh Baghel , Abhishek Trivedi , Tejas Ravichandran , Ravi Kiran Sarvadevabhatla

A unified diffusion framework for multi-modal generation and understanding has the transformative potential to achieve seamless and controllable image diffusion and other cross-modal tasks. In this paper, we introduce MMGen, a unified…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Jiepeng Wang , Zhaoqing Wang , Hao Pan , Yuan Liu , Dongdong Yu , Changhu Wang , Wenping Wang

Our ability to sample realistic natural images, particularly faces, has advanced by leaps and bounds in recent years, yet our ability to exert fine-tuned control over the generative process has lagged behind. If this new technology is to…

计算机视觉与模式识别 · 计算机科学 2020-10-20 Marek Kowalski , Stephan J. Garbin , Virginia Estellers , Tadas Baltrušaitis , Matthew Johnson , Jamie Shotton

Multimodal multitask learning has attracted an increasing interest in recent years. Singlemodal models have been advancing rapidly and have achieved astonishing results on various tasks across multiple domains. Multimodal learning offers…

计算机视觉与模式识别 · 计算机科学 2023-07-04 Ye Xue , Diego Klabjan , Jean Utke

Despite the remarkable progress of 3D generation, achieving controllability, i.e., ensuring consistency between generated 3D content and input conditions like edge and depth, remains a significant challenge. Existing methods often struggle…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Hongbin Xu , Chaohui Yu , Feng Xiao , Jiazheng Xing , Hai Ci , Weitao Chen , Fan Wang , Ming Li

Text-to-video (T2V) diffusion models have recently achieved impressive visual quality, yet most systems still generate silent clips and treat audio as a secondary concern. Existing audio-video generation pipelines typically decompose the…

Video style transfer is getting more attention in AI community for its numerous applications such as augmented reality and animation productions. Compared with traditional image style transfer, performing this task on video presents new…

计算机视觉与模式识别 · 计算机科学 2021-01-21 Yingying Deng , Fan Tang , Weiming Dong , Haibin Huang , Chongyang Ma , Changsheng Xu

Given large amount of real photos for training, Convolutional neural network shows excellent performance on object recognition tasks. However, the process of collecting data is so tedious and the background are also limited which makes it…

计算机视觉与模式识别 · 计算机科学 2022-05-10 Yida Wang , Weihong Deng

We introduce FacadeNet, a deep learning approach for synthesizing building facade images from diverse viewpoints. Our method employs a conditional GAN, taking a single view of a facade along with the desired viewpoint information and…

计算机视觉与模式识别 · 计算机科学 2023-11-06 Yiangos Georgiou , Marios Loizou , Tom Kelly , Melinos Averkiou

We present CM3Leon (pronounced "Chameleon"), a retrieval-augmented, token-based, decoder-only multi-modal language model capable of generating and infilling both text and images. CM3Leon uses the CM3 multi-modal architecture but…

We present Composable Diffusion (CoDi), a novel generative model capable of generating any combination of output modalities, such as language, image, video, or audio, from any combination of input modalities. Unlike existing generative AI…

计算机视觉与模式识别 · 计算机科学 2023-05-22 Zineng Tang , Ziyi Yang , Chenguang Zhu , Michael Zeng , Mohit Bansal