English
Related papers

Related papers: Seedream 4.0: Toward Next-generation Multimodal Im…

200 papers

Text-to-image diffusion models (T2I) have demonstrated unprecedented capabilities in creating realistic and aesthetic images. On the contrary, text-to-video diffusion models (T2V) still lag far behind in frame quality and text alignment,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Yabo Zhang , Yuxiang Wei , Xianhui Lin , Zheng Hui , Peiran Ren , Xuansong Xie , Xiangyang Ji , Wangmeng Zuo

Despite the astonishing progress in generative AI, 4D dynamic object generation remains an open challenge. With limited high-quality training data and heavy computing requirements, the combination of hallucinating unseen geometry together…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Lu Sang , Zehranaz Canfes , Dongliang Cao , Riccardo Marin , Florian Bernard , Daniel Cremers

We present xGen-VideoSyn-1, a text-to-video (T2V) generation model capable of producing realistic scenes from textual descriptions. Building on recent advancements, such as OpenAI's Sora, we explore the latent diffusion model (LDM)…

We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from…

Computer Vision and Pattern Recognition · Computer Science 2022-09-30 Uriel Singer , Adam Polyak , Thomas Hayes , Xi Yin , Jie An , Songyang Zhang , Qiyuan Hu , Harry Yang , Oron Ashual , Oran Gafni , Devi Parikh , Sonal Gupta , Yaniv Taigman

Recent strides in video generation have paved the way for unified audio-visual generation. In this work, we present Seedance 1.5 pro, a foundational model engineered specifically for native, joint audio-video generation. Leveraging a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Team Seedance , Heyi Chen , Siyan Chen , Xin Chen , Yanfei Chen , Ying Chen , Zhuo Chen , Feng Cheng , Tianheng Cheng , Xinqi Cheng , Xuyan Chi , Jian Cong , Jing Cui , Qinpeng Cui , Qide Dong , Junliang Fan , Jing Fang , Zetao Fang , Chengjian Feng , Han Feng , Mingyuan Gao , Yu Gao , Dong Guo , Qiushan Guo , Boyang Hao , Qingkai Hao , Bibo He , Qian He , Tuyen Hoang , Ruoqing Hu , Xi Hu , Weilin Huang , Zhaoyang Huang , Zhongyi Huang , Donglei Ji , Siqi Jiang , Wei Jiang , Yunpu Jiang , Zhuo Jiang , Ashley Kim , Jianan Kong , Zhichao Lai , Shanshan Lao , Yichong Leng , Ai Li , Feiya Li , Gen Li , Huixia Li , JiaShi Li , Liang Li , Ming Li , Shanshan Li , Tao Li , Xian Li , Xiaojie Li , Xiaoyang Li , Xingxing Li , Yameng Li , Yifu Li , Yiying Li , Chao Liang , Han Liang , Jianzhong Liang , Ying Liang , Zhiqiang Liang , Wang Liao , Yalin Liao , Heng Lin , Kengyu Lin , Shanchuan Lin , Xi Lin , Zhijie Lin , Feng Ling , Fangfang Liu , Gaohong Liu , Jiawei Liu , Jie Liu , Jihao Liu , Shouda Liu , Shu Liu , Sichao Liu , Songwei Liu , Xin Liu , Xue Liu , Yibo Liu , Zikun Liu , Zuxi Liu , Junlin Lyu , Lecheng Lyu , Qian Lyu , Han Mu , Xiaonan Nie , Jingzhe Ning , Xitong Pan , Yanghua Peng , Lianke Qin , Xueqiong Qu , Yuxi Ren , Kai Shen , Guang Shi , Lei Shi , Yan Song , Yinglong Song , Fan Sun , Li Sun , Renfei Sun , Yan Sun , Zeyu Sun , Wenjing Tang , Yaxue Tang , Zirui Tao , Feng Wang , Furui Wang , Jinran Wang , Junkai Wang , Ke Wang , Kexin Wang , Qingyi Wang , Rui Wang , Sen Wang , Shuai Wang , Tingru Wang , Weichen Wang , Xin Wang , Yanhui Wang , Yue Wang , Yuping Wang , Yuxuan Wang , Ziyu Wang , Guoqiang Wei , Wanru Wei , Di Wu , Guohong Wu , Hanjie Wu , Jian Wu , Jie Wu , Ruolan Wu , Xinglong Wu , Yonghui Wu , Ruiqi Xia , Liang Xiang , Fei Xiao , XueFeng Xiao , Pan Xie , Shuangyi Xie , Shuang Xu , Jinlan Xue , Shen Yan , Bangbang Yang , Ceyuan Yang , Jiaqi Yang , Runkai Yang , Tao Yang , Yang Yang , Yihang Yang , ZhiXian Yang , Ziyan Yang , Songting Yao , Yifan Yao , Zilyu Ye , Bowen Yu , Jian Yu , Chujie Yuan , Linxiao Yuan , Sichun Zeng , Weihong Zeng , Xuejiao Zeng , Yan Zeng , Chuntao Zhang , Heng Zhang , Jingjie Zhang , Kuo Zhang , Liang Zhang , Liying Zhang , Manlin Zhang , Ting Zhang , Weida Zhang , Xiaohe Zhang , Xinyan Zhang , Yan Zhang , Yuan Zhang , Zixiang Zhang , Fengxuan Zhao , Huating Zhao , Yang Zhao , Hao Zheng , Jianbin Zheng , Xiaozheng Zheng , Yangyang Zheng , Yijie Zheng , Jiexin Zhou , Jiahui Zhu , Kuan Zhu , Shenhan Zhu , Wenjia Zhu , Benhui Zou , Feilong Zuo

Inspired by the success of the text-to-image (T2I) generation task, many researchers are devoting themselves to the text-to-video (T2V) generation task. Most of the T2V frameworks usually inherit from the T2I model and add extra-temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Yiran Yang , Jinchao Zhang , Ying Deng , Jie Zhou

With the rapid advancements in diffusion models and 3D generation techniques, dynamic 3D content generation has become a crucial research area. However, achieving high-fidelity 4D (dynamic 3D) generation with strong spatial-temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Jinwei Li , Huan-ang Gao , Wenyi Li , Haohan Chi , Chenyu Liu , Chenxi Du , Yiqian Liu , Mingju Gao , Guiyu Zhang , Zongzheng Zhang , Li Yi , Yao Yao , Jingwei Zhao , Hongyang Li , Yikai Wang , Hao Zhao

Text-to-Image (T2I) generation methods based on diffusion model have garnered significant attention in the last few years. Although these image synthesis methods produce visually appealing results, they frequently exhibit spelling errors…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Yiming Zhao , Zhouhui Lian

Diffusion-driven text-to-image (T2I) generation has achieved remarkable advancements in recent years. To further improve T2I models' capability in numerical and spatial reasoning, layout is employed as an intermedium to bridge large…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yuhao Jia , Wenhan Tan

Spiking neural networks (SNNs) promise highly energy-efficient computing, but their adoption is hindered by a critical scarcity of event-stream data. This work introduces I2E, an algorithmic framework that resolves this bottleneck by…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Ruichen Ma , Liwei Meng , Guanchao Qiao , Ning Ning , Yang Liu , Shaogang Hu

Recent advancements in diffusion-based generative image editing have sparked a profound revolution, reshaping the landscape of image outpainting and inpainting tasks. Despite these strides, the field grapples with inherent challenges,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Yuxi Ren , Jie Wu , Yanzuo Lu , Huafeng Kuang , Xin Xia , Xionghui Wang , Qianqian Wang , Yixing Zhu , Pan Xie , Shiyin Wang , Xuefeng Xiao , Yitong Wang , Min Zheng , Lean Fu

Text-to-image synthesis (T2I) aims to generate photo-realistic images which are semantically consistent with the text descriptions. Existing methods are usually built upon conditional generative adversarial networks (GANs) and initialize an…

Computer Vision and Pattern Recognition · Computer Science 2022-03-25 Kai Hu , Wentong Liao , Michael Ying Yang , Bodo Rosenhahn

Diffusion models have exhibited substantial success in text-to-image generation. However, they often encounter challenges when dealing with complex and dense prompts involving multiple objects, attribute binding, and long descriptions. In…

Computer Vision and Pattern Recognition · Computer Science 2024-08-28 Mushui Liu , Yuhang Ma , Yang Zhen , Jun Dan , Yunlong Yu , Zeng Zhao , Zhipeng Hu , Bai Liu , Changjie Fan

Generating realistic and physically plausible 3D Human-Object Interactions (HOI) remains a key challenge in motion generation. One primary reason is that describing these physical constraints with words alone is difficult. To address this…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Songjin Cai , Linjie Zhong , Ling Guo , Changxing Ding

Text-to-image (T2I) diffusion/flow models have drawn considerable attention recently due to their remarkable ability to deliver flexible visual creations. Still, high-resolution image synthesis presents formidable challenges due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Jiazi Bu , Pengyang Ling , Yujie Zhou , Pan Zhang , Tong Wu , Xiaoyi Dong , Yuhang Zang , Yuhang Cao , Dahua Lin , Jiaqi Wang

Recent advancements in camera-trajectory-guided image-to-video generation offer higher precision and better support for complex camera control compared to text-based approaches. However, they also introduce significant usability challenges,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Teng Li , Guangcong Zheng , Rui Jiang , Shuigen Zhan , Tao Wu , Yehao Lu , Yining Lin , Chuanyun Deng , Yepan Xiong , Min Chen , Lin Cheng , Xi Li

Text-to-image retrieval (T2I retrieval) remains challenging because cross-modal embeddings often behave as bags of concepts, underrepresenting structured visual relationships such as pose and viewpoint. We proposeVisualize-then-Retrieve…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Di Wu , Yixin Wan , Kai-Wei Chang

We demonstrate NeedleDB, an open-source, deployment-ready database system for answering complex natural language queries over image data. Unlike existing approaches that rely on contrastive-learning embeddings (e.g., CLIP), which degrade on…

Databases · Computer Science 2026-03-31 Mahdi Erfanian , Abolfazl Asudeh

This paper presents ThinkDiff, a novel alignment paradigm that empowers text-to-image diffusion models with multimodal in-context understanding and reasoning capabilities by integrating the strengths of vision-language models (VLMs).…

Machine Learning · Computer Science 2025-02-18 Zhenxing Mi , Kuan-Chieh Wang , Guocheng Qian , Hanrong Ye , Runtao Liu , Sergey Tulyakov , Kfir Aberman , Dan Xu

Text-to-video generation has trailed behind text-to-image generation in terms of quality and diversity, primarily due to the inherent complexities of spatio-temporal modeling and the limited availability of video-text datasets. Recent…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Xiefan Guo , Jinlin Liu , Miaomiao Cui , Liefeng Bo , Di Huang