English

Seedream 4.0: Toward Next-generation Multimodal Image Generation

Computer Vision and Pattern Recognition 2025-12-11 v3

Abstract

We introduce Seedream 4.0, an efficient and high-performance multimodal image generation system that unifies text-to-image (T2I) synthesis, image editing, and multi-image composition within a single framework. We develop a highly efficient diffusion transformer with a powerful VAE which also can reduce the number of image tokens considerably. This allows for efficient training of our model, and enables it to fast generate native high-resolution images (e.g., 1K-4K). Seedream 4.0 is pretrained on billions of text-image pairs spanning diverse taxonomies and knowledge-centric concepts. Comprehensive data collection across hundreds of vertical scenarios, coupled with optimized strategies, ensures stable and large-scale training, with strong generalization. By incorporating a carefully fine-tuned VLM model, we perform multi-modal post-training for training both T2I and image editing tasks jointly. For inference acceleration, we integrate adversarial distillation, distribution matching, and quantization, as well as speculative decoding. It achieves an inference time of up to 1.8 seconds for generating a 2K image (without a LLM/VLM as PE model). Comprehensive evaluations reveal that Seedream 4.0 can achieve state-of-the-art results on both T2I and multimodal image editing. In particular, it demonstrates exceptional multimodal capabilities in complex tasks, including precise image editing and in-context reasoning, and also allows for multi-image reference, and can generate multiple output images. This extends traditional T2I systems into an more interactive and multidimensional creative tool, pushing the boundary of generative AI for both creativity and professional applications. We further scale our model and data as Seedream 4.5. Seedream 4.0 and Seedream 4.5 are accessible on Volcano Engine https://www.volcengine.com/experience/ark?launch=seedream.

Keywords

Cite

@article{arxiv.2509.20427,
  title  = {Seedream 4.0: Toward Next-generation Multimodal Image Generation},
  author = {Team Seedream and : and Yunpeng Chen and Yu Gao and Lixue Gong and Meng Guo and Qiushan Guo and Zhiyao Guo and Xiaoxia Hou and Weilin Huang and Yixuan Huang and Xiaowen Jian and Huafeng Kuang and Zhichao Lai and Fanshi Li and Liang Li and Xiaochen Lian and Chao Liao and Liyang Liu and Wei Liu and Yanzuo Lu and Zhengxiong Luo and Tongtong Ou and Guang Shi and Yichun Shi and Shiqi Sun and Yu Tian and Zhi Tian and Peng Wang and Rui Wang and Xun Wang and Ye Wang and Guofeng Wu and Jie Wu and Wenxu Wu and Yonghui Wu and Xin Xia and Xuefeng Xiao and Shuang Xu and Xin Yan and Ceyuan Yang and Jianchao Yang and Zhonghua Zhai and Chenlin Zhang and Heng Zhang and Qi Zhang and Xinyu Zhang and Yuwei Zhang and Shijia Zhao and Wenliang Zhao and Wenjia Zhu},
  journal= {arXiv preprint arXiv:2509.20427},
  year   = {2025}
}

Comments

Seedream 4.0/4.5 Technical Report

R2 v1 2026-07-01T05:54:42.219Z