English
Related papers

Related papers: Seedream 4.0: Toward Next-generation Multimodal Im…

200 papers

Due to the fascinating generative performance of text-to-image diffusion models, growing text-to-3D generation works explore distilling the 2D generative priors into 3D, using the score distillation sampling (SDS) loss, to bypass the data…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Yu-Jie Yuan , Leif Kobbelt , Jiwen Liu , Yuan Zhang , Pengfei Wan , Yu-Kun Lai , Lin Gao

Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro, Seedance 2.0 adopts a unified, highly efficient, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Team Seedance , De Chen , Liyang Chen , Xin Chen , Ying Chen , Zhuo Chen , Zhuowei Chen , Feng Cheng , Tianheng Cheng , Yufeng Cheng , Mojie Chi , Xuyan Chi , Jian Cong , Qinpeng Cui , Fei Ding , Qide Dong , Yujiao Du , Haojie Duanmu , Junliang Fan , Jiarui Fang , Jing Fang , Zetao Fang , Chengjian Feng , Yu Gao , Diandian Gu , Dong Guo , Hanzhong Guo , Qiushan Guo , Boyang Hao , Hongxiang Hao , Haoxun He , Jiaao He , Qian He , Tuyen Hoang , Heng Hu , Ruoqing Hu , Yuxiang Hu , Jiancheng Huang , Weilin Huang , Zhaoyang Huang , Zhongyi Huang , Jishuo Jin , Ming Jing , Ashley Kim , Shanshan Lao , Yichong Leng , Bingchuan Li , Gen Li , Haifeng Li , Huixia Li , Jiashi Li , Ming Li , Xiaojie Li , Xingxing Li , Yameng Li , Yiying Li , Yu Li , Yueyan Li , Chao Liang , Han Liang , Jianzhong Liang , Ying Liang , Wang Liao , J. H. Lien , Shanchuan Lin , Xi Lin , Feng Ling , Yue Ling , Fangfang Liu , Jiawei Liu , Jihao Liu , Jingtuo Liu , Shu Liu , Sichao Liu , Wei Liu , Xue Liu , Zuxi Liu , Ruijie Lu , Lecheng Lyu , Jingting Ma , Tianxiang Ma , Xiaonan Nie , Jingzhe Ning , Junjie Pan , Xitong Pan , Ronggui Peng , Xueqiong Qu , Yuxi Ren , Yuchen Shen , Guang Shi , Lei Shi , Yinglong Song , Fan Sun , Li Sun , Renfei Sun , Wenjing Tang , Boyang Tao , Zirui Tao , Dongliang Wang , Feng Wang , Hulin Wang , Ke Wang , Qingyi Wang , Rui Wang , Shuai Wang , Shulei Wang , Weichen Wang , Xuanda Wang , Yanhui Wang , Yue Wang , Yuping Wang , Yuxuan Wang , Zijie Wang , Ziyu Wang , Guoqiang Wei , Meng Wei , Di Wu , Guohong Wu , Hanjie Wu , Huachao Wu , Jian Wu , Jie Wu , Ruolan Wu , Shaojin Wu , Xiaohu Wu , Xinglong Wu , Yonghui Wu , Ruiqi Xia , Xin Xia , Xuefeng Xiao , Shuang Xu , Bangbang Yang , Jiaqi Yang , Runkai Yang , Tao Yang , Yihang Yang , Zhixian Yang , Ziyan Yang , Fulong Ye , Bingqian Yi , Xing Yin , Yongbin You , Linxiao Yuan , Weihong Zeng , Xuejiao Zeng , Yan Zeng , Siyu Zhai , Zhonghua Zhai , Bowen Zhang , Chenlin Zhang , Heng Zhang , Jun Zhang , Manlin Zhang , Peiyuan Zhang , Shuo Zhang , Xiaohe Zhang , Xiaoying Zhang , Xinyan Zhang , Xinyi Zhang , Yichi Zhang , Zixiang Zhang , Haiyu Zhao , Huating Zhao , Liming Zhao , Yian Zhao , Guangcong Zheng , Jianbin Zheng , Xiaozheng Zheng , Zerong Zheng , Kuan Zhu , Feilong Zuo

Prevailing autoregressive (AR) models for text-to-image generation either rely on heavy, computationally-intensive diffusion models to process continuous image tokens, or employ vector quantization (VQ) to obtain discrete tokens with…

Text-to-image generation with visual autoregressive~(VAR) models has recently achieved impressive advances in generation fidelity and inference efficiency. While control mechanisms have been explored for diffusion models, enabling precise…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Keli Liu , Zhendong Wang , Wengang Zhou , Shaodong Xu , Ruixiao Dong , Houqiang Li

Recent advances in Vision-Language Models (VLMs) have enabled unified understanding across text and images, yet equipping these models with robust image generation capabilities remains challenging. Existing approaches often rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Xiangyi Chen , Théophane Vallaeys , Maha Elbayad , John Nguyen , Jakob Verbeek

We introduce MVControl, a novel neural network architecture that enhances existing pre-trained multi-view 2D diffusion models by incorporating additional input conditions, e.g. edge maps. Our approach enables the generation of controllable…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Zhiqi Li , Yiming Chen , Lingzhe Zhao , Peidong Liu

Despite significant advancements in customizing text-to-image and video generation models, generating images and videos that effectively integrate multiple personalized concepts remains a challenging task. To address this, we present…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Gihyun Kwon , Jong Chul Ye

Recent advancements in Virtual Try-On (VTO) have demonstrated exceptional efficacy in generating realistic images and preserving garment details, largely attributed to the robust generative capabilities of text-to-image (T2I) diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Zhenchen Wan , Yanwu Xu , Zhaoqing Wang , Feng Liu , Tongliang Liu , Mingming Gong

Generative Artificial Intelligence (AI) has created unprecedented opportunities for creative expression, education, and research. Text-to-image systems such as DALL.E, Stable Diffusion, and Midjourney can now convert ideas into visuals…

Artificial Intelligence · Computer Science 2025-12-16 Dang Phuong Nam , Nguyen Kieu , Pham Thanh Hieu

Diffusion model based Text-to-Image has achieved impressive achievements recently. Although current technology for synthesizing images is highly advanced and capable of generating images with high fidelity, it is still possible to give the…

Computer Vision and Pattern Recognition · Computer Science 2024-02-22 Yuxiang Tuo , Wangmeng Xiang , Jun-Yan He , Yifeng Geng , Xuansong Xie

Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition. One plausible contributing factor is that natural image datasets provide limited supervision for low-level…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Guanyu Zhou , Yida Yin , Wenhao Chai , Shengbang Tong , Xingyu Fu , Zhuang Liu

Large Text-to-Image(T2I) diffusion models have shown a remarkable capability to produce photorealistic and diverse images based on text inputs. However, existing works only support limited language input, e.g., English, Chinese, and…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Fulong Ye , Guang Liu , Xinya Wu , Ledell Wu

In spite of the rapidly evolving landscape of text-to-image generation, the synthesis and manipulation of multiple entities while adhering to specific relational constraints pose enduring challenges. This paper introduces an innovative…

Computer Vision and Pattern Recognition · Computer Science 2024-01-22 YuTeng Ye , Jiale Cai , Hang Zhou , Guanwen Li , Youjia Zhang , Zikai Song , Chenxing Gao , Junqing Yu , Wei Yang

Most of the recent generative image super-resolution (SR) methods rely on adapting large text-to-image (T2I) diffusion models pretrained on web-scale text-image data. While effective, this paradigm starts from a generic T2I generator,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Rongyuan Wu , Lingchen Sun , Zhengqiang Zhang , Xiangtao Kong , Jixin Zhao , Shihao Wang , Lei Zhang

Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks; however, effectively integrating image segmentation into these models remains a significant challenge. In this paper, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Mengcheng Lan , Chaofeng Chen , Yue Zhou , Jiaxing Xu , Yiping Ke , Xinjiang Wang , Litong Feng , Wayne Zhang

Generating novel views from a single image remains a challenging task due to the complexity of 3D scenes and the limited diversity in the existing multi-view datasets to train a model on. Recent research combining large-scale text-to-image…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Junyoung Seo , Kazumi Fukuda , Takashi Shibuya , Takuya Narihira , Naoki Murata , Shoukang Hu , Chieh-Hsin Lai , Seungryong Kim , Yuki Mitsufuji

Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators. However, modeling the vast spatiotemporal space remains computationally demanding, particularly when…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Jiancheng Huang , Gengwei Zhang , Zequn Jie , Siyu Jiao , Yinlong Qian , Ling Chen , Yunchao Wei , Lin Ma

In the early stages of architectural design, shoebox models are typically used as a simplified representation of building structures but require extensive operations to transform them into detailed designs. Generative artificial…

Graphics · Computer Science 2025-03-06 Xusheng Du , Ruihan Gui , Zhengyang Wang , Ye Zhang , Haoran Xie

Lately, researchers in artificial intelligence have been really interested in how language and vision come together, giving rise to the development of multimodal models that aim to seamlessly integrate textual and visual information.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Rajat Chawla , Arkajit Datta , Tushar Verma , Adarsh Jha , Anmol Gautam , Ayush Vatsal , Sukrit Chaterjee , Mukunda NS , Ishaan Bhola

Removing unwanted concepts from large-scale text-to-image (T2I) diffusion models while maintaining their overall generative quality remains an open challenge. This difficulty is especially pronounced in emerging paradigms, such as Stable…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Daiheng Gao , Shilin Lu , Shaw Walters , Wenbo Zhou , Jiaming Chu , Jie Zhang , Bang Zhang , Mengxi Jia , Jian Zhao , Zhaoxin Fan , Weiming Zhang
‹ Prev 1 8 9 10 Next ›