English
Related papers

Related papers: Seedream 2.0: A Native Chinese-English Bilingual I…

200 papers

We present Seedream 3.0, a high-performance Chinese-English bilingual image generation foundation model. We develop several technical improvements to address existing challenges in Seedream 2.0, including alignment with complicated prompts,…

We introduce Seedream 4.0, an efficient and high-performance multimodal image generation system that unifies text-to-image (T2I) synthesis, image editing, and multi-image composition within a single framework. We develop a highly efficient…

Text-to-Image generation (TTI) technologies are advancing rapidly, especially in the English language communities. However, apart from the user input language barrier problem, English-native TTI models inherently carry biases from their…

Computation and Language · Computer Science 2026-03-19 Shanyuan Liu , Bo Cheng , Yuhang Ma , Liebucha Wu , Ao Ma , Xiaoyu Wu , Dawei Leng , Yuhui Yin

Recent breakthroughs in the field of language-guided image generation have yielded impressive achievements, enabling the creation of high-quality and diverse images based on user instructions.Although the synthesis performance is…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Jian Ma , Mingjun Zhao , Chen Chen , Ruichen Wang , Di Niu , Haonan Lu , Xiaodong Lin

Recent advancements in text-to-image models have significantly enhanced image generation capabilities, yet a notable gap of open-source models persists in bilingual or Chinese language support. To address this need, we present…

Computation and Language · Computer Science 2024-06-19 Xiaojun Wu , Dixiang Zhang , Ruyi Gan , Junyu Lu , Ziwei Wu , Renliang Sun , Jiaxing Zhang , Pingjian Zhang , Yan Song

Modern text-to-image diffusion models encode rich visual priors, but expose them only through one-way text-conditioned generation. Existing unified vision--language models derived from them recover bidirectional capability through…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Eric Tillmann Bill , Enis Simsar , Alessio Tonioni , Thomas Hofmann

We introduce LongCat-Image, a pioneering open-source and bilingual (Chinese-English) foundation model for image generation, designed to address core challenges in multilingual text rendering, photorealism, deployment efficiency, and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Meituan LongCat Team , Hanghang Ma , Haoxian Tan , Jiale Huang , Junqiang Wu , Jun-Yan He , Lishuai Gao , Songlin Xiao , Xiaoming Wei , Xiaoqi Ma , Xunliang Cai , Yayong Guan , Jie Hu

We present Qwen-Image, an image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. To address the challenges of complex text rendering, we design a…

We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite recent progress, existing models still struggle with ultra-long…

Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro, Seedance 2.0 adopts a unified, highly efficient, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Team Seedance , De Chen , Liyang Chen , Xin Chen , Ying Chen , Zhuo Chen , Zhuowei Chen , Feng Cheng , Tianheng Cheng , Yufeng Cheng , Mojie Chi , Xuyan Chi , Jian Cong , Qinpeng Cui , Fei Ding , Qide Dong , Yujiao Du , Haojie Duanmu , Junliang Fan , Jiarui Fang , Jing Fang , Zetao Fang , Chengjian Feng , Yu Gao , Diandian Gu , Dong Guo , Hanzhong Guo , Qiushan Guo , Boyang Hao , Hongxiang Hao , Haoxun He , Jiaao He , Qian He , Tuyen Hoang , Heng Hu , Ruoqing Hu , Yuxiang Hu , Jiancheng Huang , Weilin Huang , Zhaoyang Huang , Zhongyi Huang , Jishuo Jin , Ming Jing , Ashley Kim , Shanshan Lao , Yichong Leng , Bingchuan Li , Gen Li , Haifeng Li , Huixia Li , Jiashi Li , Ming Li , Xiaojie Li , Xingxing Li , Yameng Li , Yiying Li , Yu Li , Yueyan Li , Chao Liang , Han Liang , Jianzhong Liang , Ying Liang , Wang Liao , J. H. Lien , Shanchuan Lin , Xi Lin , Feng Ling , Yue Ling , Fangfang Liu , Jiawei Liu , Jihao Liu , Jingtuo Liu , Shu Liu , Sichao Liu , Wei Liu , Xue Liu , Zuxi Liu , Ruijie Lu , Lecheng Lyu , Jingting Ma , Tianxiang Ma , Xiaonan Nie , Jingzhe Ning , Junjie Pan , Xitong Pan , Ronggui Peng , Xueqiong Qu , Yuxi Ren , Yuchen Shen , Guang Shi , Lei Shi , Yinglong Song , Fan Sun , Li Sun , Renfei Sun , Wenjing Tang , Boyang Tao , Zirui Tao , Dongliang Wang , Feng Wang , Hulin Wang , Ke Wang , Qingyi Wang , Rui Wang , Shuai Wang , Shulei Wang , Weichen Wang , Xuanda Wang , Yanhui Wang , Yue Wang , Yuping Wang , Yuxuan Wang , Zijie Wang , Ziyu Wang , Guoqiang Wei , Meng Wei , Di Wu , Guohong Wu , Hanjie Wu , Huachao Wu , Jian Wu , Jie Wu , Ruolan Wu , Shaojin Wu , Xiaohu Wu , Xinglong Wu , Yonghui Wu , Ruiqi Xia , Xin Xia , Xuefeng Xiao , Shuang Xu , Bangbang Yang , Jiaqi Yang , Runkai Yang , Tao Yang , Yihang Yang , Zhixian Yang , Ziyan Yang , Fulong Ye , Bingqian Yi , Xing Yin , Yongbin You , Linxiao Yuan , Weihong Zeng , Xuejiao Zeng , Yan Zeng , Siyu Zhai , Zhonghua Zhai , Bowen Zhang , Chenlin Zhang , Heng Zhang , Jun Zhang , Manlin Zhang , Peiyuan Zhang , Shuo Zhang , Xiaohe Zhang , Xiaoying Zhang , Xinyan Zhang , Xinyi Zhang , Yichi Zhang , Zixiang Zhang , Haiyu Zhao , Huating Zhao , Liming Zhao , Yian Zhao , Guangcong Zheng , Jianbin Zheng , Xiaozheng Zheng , Zerong Zheng , Kuan Zhu , Feilong Zuo

Recent diffusion and flow matching models have demonstrated strong capabilities in image generation and editing by progressively removing noise through iterative sampling. While this enables flexible inversion for semantic-preserving edits,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yasong Dai , Zeeshan Hayder , David Ahmedt-Aristizabal , Hongdong Li

Text-to-image generation models often struggle with key element loss or semantic confusion in tasks involving Chinese classical poetry.Addressing this issue through fine-tuning models needs considerable training costs. Additionally, manual…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Jing Jiang , Yiran Ling , Binzhu Li , Pengxiang Li , Junming Piao , Yu Zhang

Notable breakthroughs in diffusion modeling have propelled rapid improvements in video generation, yet current foundational model still face critical challenges in simultaneously balancing prompt following, motion plausibility, and visual…

The evolution of visual generative models has long been constrained by fragmented architectures relying on disjoint text encoders and external VAEs. In this report, we present HiDream-O1-Image, a natively unified generative foundation model…

We introduce SeedEdit 3.0, in companion with our T2I model Seedream 3.0, which significantly improves over our previous SeedEdit versions in both aspects of edit instruction following and image content (e.g., ID/IP) preservation on real…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Peng Wang , Yichun Shi , Xiaochen Lian , Zhonghua Zhai , Xin Xia , Xuefeng Xiao , Weilin Huang , Jianchao Yang

We proposed the Chinese Text Adapter-Flux (CTA-Flux). An adaptation method fits the Chinese text inputs to Flux, a powerful text-to-image (TTI) generative model initially trained on the English corpus. Despite the notable image generation…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Yue Gong , Shanyuan Liu , Liuzhuozheng Li , Jian Zhu , Bo Cheng , Liebucha Wu , Xiaoyu Wu , Yuhang Ma , Dawei Leng , Yuhui Yin

Recent advances in large-scale text-to-image generation models have led to a surge in subject-driven text-to-image generation, which aims to produce customized images that align with textual descriptions while preserving the identity of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Kewen Chen , Xiaobin Hu , Wenqi Ren

Recently, large-scale diffusion models, e.g., Stable diffusion and DallE2, have shown remarkable results on image synthesis. On the other hand, large-scale cross-modal pre-trained models (e.g., CLIP, ALIGN, and FILIP) are competent for…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Runhui Huang , Jianhua Han , Guansong Lu , Xiaodan Liang , Yihan Zeng , Wei Zhang , Hang Xu

The field of advanced text-to-image generation is witnessing the emergence of unified frameworks that integrate powerful text encoders, such as CLIP and T5, with Diffusion Transformer backbones. Although there have been efforts to control…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Liang Chen , Shuai Bai , Wenhao Chai , Weichu Xie , Haozhe Zhao , Leon Vinci , Junyang Lin , Baobao Chang

Recent advancements in diffusion-based generative image editing have sparked a profound revolution, reshaping the landscape of image outpainting and inpainting tasks. Despite these strides, the field grapples with inherent challenges,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Yuxi Ren , Jie Wu , Yanzuo Lu , Huafeng Kuang , Xin Xia , Xionghui Wang , Qianqian Wang , Yixing Zhu , Pan Xie , Shiyin Wang , Xuefeng Xiao , Yitong Wang , Min Zheng , Lean Fu
‹ Prev 1 2 3 10 Next ›