English
Related papers

Related papers: Skywork UniPic: Unified Autoregressive Modeling fo…

200 papers

The emergence of unified multimodal understanding and generation models is rapidly attracting attention because of their ability to enhance instruction-following capabilities while minimizing model redundancy. However, there is a lack of a…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Yi Li , Haonan Wang , Qixiang Zhang , Boyu Xiao , Chenchang Hu , Hualiang Wang , Xiaomeng Li

Unified understanding and generation is a highly appealing research direction in multimodal learning. There exist two approaches: one trains a transformer via an auto-regressive paradigm, and the other adopts a two-stage scheme connecting…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Shihao Zhao , Yitong Chen , Zeyinzi Jiang , Bojia Zi , Shaozhe Hao , Yu Liu , Chaojie Mao , Kwan-Yee K. Wong

Low-level vision involves a wide spectrum of tasks, including image restoration, enhancement, stylization, and feature extraction, which differ significantly in both task formulation and output domains. To address the challenge of unified…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Xiangyu Chen , Kaiwen Zhu , Yuandong Pu , Shuo Cao , Xiaohui Li , Wenlong Zhang , Yihao Liu , Yu Qiao , Jiantao Zhou , Chao Dong

The current conditional autoregressive image generation methods have shown promising results, yet their potential remains largely unexplored in the practical unsupervised image translation domain, which operates without explicit…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Yi Liu , Shengqian Li , Zuzeng Lin , Feng Wang , Si Liu

In this paper, we introduce OneReward, a unified reinforcement learning framework that enhances the model's generative capabilities across multiple tasks under different evaluation criteria using only \textit{One Reward} model. By employing…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Yuan Gong , Xionghui Wang , Jie Wu , Shiyin Wang , Yitong Wang , Xinglong Wu

We present UniMIC, a universal multi-modality image compression framework, intending to unify the rate-distortion-perception (RDP) optimization for multiple image codecs simultaneously through excavating cross-modality generative priors.…

Image and Video Processing · Electrical Eng. & Systems 2024-12-10 Yixin Gao , Xin Li , Xiaohan Pan , Runsen Feng , Zongyu Guo , Yiting Lu , Yulin Ren , Zhibo Chen

We present Skywork R1V2, a next-generation multimodal reasoning model and a major leap forward from its predecessor, Skywork R1V. At its core, R1V2 introduces a hybrid reinforcement learning paradigm that jointly leverages the Mixed…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Peiyu Wang , Yichen Wei , Yi Peng , Xiaokun Wang , Weijie Qiu , Wei Shen , Tianyidan Xie , Jiangbo Pei , Jianhao Zhang , Yunzhuo Hao , Xuchen Song , Yang Liu , Yahui Zhou

Text-to-Image (T2I) diffusion models have shown impressive results in generating visually compelling images following user prompts. Building on this, various methods further fine-tune the pre-trained T2I model for specific tasks. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Tsu-Jui Fu , Yusu Qian , Chen Chen , Wenze Hu , Zhe Gan , Yinfei Yang

Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. While image generation and editing have rapidly transitioned from task-specific to unified…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Xuan Ju , Tianyu Wang , Yuqian Zhou , He Zhang , Qing Liu , Nanxuan Zhao , Zhifei Zhang , Yijun Li , Yuanhao Cai , Shaoteng Liu , Daniil Pakhomov , Zhe Lin , Soo Ye Kim , Qiang Xu

Unified Multimodal Large Language Models (MLLMs) require a visual representation that simultaneously supports high-fidelity reconstruction, complex semantic extraction, and generative suitability. However, existing visual tokenizers…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Shaobin Zhuang , Yuang Ai , Jiaming Han , Weijia Mao , Xiaohui Li , Fangyikang Wang , Xiao Wang , Yan Li , Shanchuan Lin , Kun Xu , Zhenheng Yang , Huaibo Huang , Xiangyu Yue , Hao Chen , Yali Wang

We introduce spatially grounded contextual image generation, a controllable image generation task that reframes the conditioning paradigm. Instead of supplying a reference image and a global text prompt through two separate encoders, one…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Jiayun Wang , Yu Wang , Weijie Gan , Zhenting Wang , Wei Wei

Unified image understanding and generation has emerged as a promising paradigm in multimodal artificial intelligence. Despite recent progress, the optimal architectural design for such unified models remains an open challenge. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Teng Li , Quanfeng Lu , Lirui Zhao , Hao Li , Xizhou Zhu , Yu Qiao , Jun Zhang , Wenqi Shao

Image fusion aims to integrate complementary information from multiple source images to produce a more informative and visually consistent representation, benefiting both human perception and downstream vision tasks. Despite recent…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xingyuan Li , Songcheng Du , Yang Zou , HaoYuan Xu , Zhiying Jiang , Jinyuan Liu

Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their potential in supporting computer-use planning tasks, which are closely related to our…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Junxian Li , Kai Liu , Leyang Chen , Weida Wang , Zhixin Wang , Jiaqi Xu , Fan Li , Renjing Pei , Linghe Kong , Yulun Zhang

In recent years, significant progress has been made in both image generation and generated image detection. Despite their rapid, yet largely independent, development, these two fields have evolved distinct architectural paradigms: the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Yanran Zhang , Wenzhao Zheng , Yifei Li , Bingyao Yu , Yu Zheng , Lei Chen , Jiwen Lu , Jie Zhou

Multimodal generation has long been dominated by text-driven pipelines where language dictates vision but cannot reason or create within it. We challenge this paradigm by asking whether all modalities, including textual descriptions,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Junchao Yi , Rui Zhao , Jiahao Tang , Weixian Lei , Linjie Li , Qisheng Su , Zhengyuan Yang , Lijuan Wang , Xiaofeng Zhu , Alex Jinpeng Wang

Most existing vision-language pre-training methods focus on understanding tasks and use BERT-like objectives (masked language modeling and image-text matching) during pretraining. Although they perform well in many understanding downstream…

Computer Vision and Pattern Recognition · Computer Science 2021-12-16 Tianyi Liu , Zuxuan Wu , Wenhan Xiong , Jingjing Chen , Yu-Gang Jiang

We introduce OneDiffusion, a versatile, large-scale diffusion model that seamlessly supports bidirectional image synthesis and understanding across diverse tasks. It enables conditional generation from inputs such as text, depth, pose,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Duong H. Le , Tuan Pham , Sangho Lee , Christopher Clark , Aniruddha Kembhavi , Stephan Mandt , Ranjay Krishna , Jiasen Lu

Recent video generation models demonstrate impressive synthesis capabilities but remain limited by single-modality conditioning, constraining their holistic world understanding. This stems from insufficient cross-modal interaction and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Jiehui Huang , Yuechen Zhang , Xu He , Yuan Gao , Zhi Cen , Bin Xia , Yan Zhou , Xin Tao , Pengfei Wan , Jiaya Jia

Unifying text-image contrastive learning and text-to-image (T2I) generation in a single end-to-end model is challenging because the two objectives demand opposing masking regimes: contrastive alignment needs near-complete visible tokens,…