English
Related papers

Related papers: RealUnify: Do Unified Models Truly Benefit from Un…

200 papers

Spatial intelligence is essential for multimodal large language models, yet current benchmarks largely assess it only from an understanding perspective. We ask whether modern generative or unified multimodal models also possess generative…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Muzhi Zhu , Shunyao Jiang , Huanyi Zheng , Zekai Luo , Hao Zhong , Anzhou Li , Kaijun Wang , Jintao Rong , Yang Liu , Hao Chen , Tao Lin , Chunhua Shen

Emotional understanding and generation are often treated as separate tasks, yet they are inherently complementary and can mutually enhance each other. In this paper, we propose the UniEmo, a unified framework that seamlessly integrates…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Yijie Zhu , Lingsen Zhang , Zitong Yu , Rui Shao , Tao Tan , Liqiang Nie

Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying…

Video generation models have significantly advanced embodied intelligence, unlocking new possibilities for generating diverse robot data that capture perception, reasoning, and action in the physical world. However, synthesizing…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Yufan Deng , Zilin Pan , Hongyu Zhang , Xiaojie Li , Ruoqing Hu , Yufei Ding , Yiming Zou , Yan Zeng , Daquan Zhou

This paper presents Omni-View, which extends the unified multimodal understanding and generation to 3D scenes based on multiview images, exploring the principle that "generation facilitates understanding". Consisting of understanding model,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 JiaKui Hu , Shanshan Zhao , Qing-Guo Chen , Xuerui Qiu , Jialun Liu , Zhao Xu , Weihua Luo , Kaifu Zhang , Yanye Lu

We introduce UEval, a benchmark to evaluate unified models, i.e., models capable of generating both images and text. UEval comprises 1,000 expert-curated questions that require both images and text in the model output, sourced from 8…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Bo Li , Yida Yin , Wenhao Chai , Xingyu Fu , Zhuang Liu

Unified Multimodal Large Models (UMLMs) integrate understanding and generation capabilities within a single architecture. While this architectural unification, driven by the deep fusion of multimodal features, enhances model performance, it…

Artificial Intelligence · Computer Science 2026-04-02 Zixiang Peng , Yongxiu Xu , Qinyi Zhang , Jiexun Shen , Yifan Zhang , Hongbo Xu , Yubin Wang , Gaopeng Gou

With the rapid advancement of generative models, associated privacy concerns have attracted growing attention. To address this, researchers have begun adapting machine unlearning techniques from traditional classification models to…

Machine Learning · Computer Science 2025-07-29 Xiaohua Feng , Jiaming Zhang , Fengyuan Yu , Chengye Wang , Li Zhang , Kaixiang Li , Yuyuan Li , Chaochao Chen , Jianwei Yin

Autonomous vehicle perception typically relies on modular pipelines that decompose the task into detection, tracking, and prediction. While interpretable, these pipelines suffer from error accumulation and limited inter-task synergy.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Loïc Stratil , Felix Fent , Esteban Rivera , Markus Lienkamp

Despite recent progress, medical foundation models still struggle to unify visual understanding and generation, as these tasks have inherently conflicting goals: semantic abstraction versus pixel-level reconstruction. Existing approaches,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Ruiheng Zhang , Jingfeng Yao , Huangxuan Zhao , Hao Yan , Xiao He , Lei Chen , Zhou Wei , Yong Luo , Zengmao Wang , Lefei Zhang , Dacheng Tao , Bo Du

The integration of geometric reconstruction and generative modeling remains a critical challenge in developing AI systems capable of human-like spatial reasoning. This paper proposes Aether, a unified framework that enables geometry-aware…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Aether Team , Haoyi Zhu , Yifan Wang , Jianjun Zhou , Wenzheng Chang , Yang Zhou , Zizun Li , Junyi Chen , Chunhua Shen , Jiangmiao Pang , Tong He

Multi-modal generative AI (Artificial Intelligence) has attracted increasing attention from both academia and industry. Particularly, two dominant families of techniques have emerged: i) Multi-modal large language models (LLMs) demonstrate…

Artificial Intelligence · Computer Science 2025-11-26 Xin Wang , Yuwei Zhou , Bin Huang , Hong Chen , Wenwu Zhu

Deep learning yields great results across many fields, from speech recognition, image classification, to translation. But for each problem, getting a deep model to work well involves research into the architecture and a long period of…

Machine Learning · Computer Science 2017-06-19 Lukasz Kaiser , Aidan N. Gomez , Noam Shazeer , Ashish Vaswani , Niki Parmar , Llion Jones , Jakob Uszkoreit

The capability of Unified Multimodal Models (UMMs) to apply world knowledge across diverse tasks remains a critical, unresolved challenge. Existing benchmarks fall short, offering only siloed, single-task evaluations with limited diagnostic…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Jintao Lin , Bowen Dong , Weikang Shi , Chenyang Lei , Suiyun Zhang , Rui Liu , Xihui Liu

Vision-language large models are moving toward the unification of visual understanding and visual generation tasks. However, whether generation can enhance understanding is still under-explored on large data scale. In this work, we analysis…

Computation and Language · Computer Science 2026-01-01 Fengjiao Chen , Minhao Jing , Weitao Lu , Yan Feng , Xiaoyu Li , Xuezhi Cao

Multimodal generative models have made significant strides in image editing, demonstrating impressive performance on a variety of static tasks. However, their proficiency typically does not extend to complex scenarios requiring dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Zhiqiang Sheng , Xumeng Han , Zhiwei Zhang , Zenghui Xiong , Yifan Ding , Aoxiang Ping , Xiang Li , Tong Guo , Yao Mao

We present JanusFlow, a powerful framework that unifies image understanding and generation in a single model. JanusFlow introduces a minimalist architecture that integrates autoregressive language models with rectified flow, a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Yiyang Ma , Xingchao Liu , Xiaokang Chen , Wen Liu , Chengyue Wu , Zhiyu Wu , Zizheng Pan , Zhenda Xie , Haowei Zhang , Xingkai yu , Liang Zhao , Yisong Wang , Jiaying Liu , Chong Ruan

The rapid evolution of multimodal foundation model has demonstrated significant progresses in vision-language understanding and generation, e.g., our previous work SEED-LLaMA. However, there remains a gap between its capability and the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Yuying Ge , Sijie Zhao , Jinguo Zhu , Yixiao Ge , Kun Yi , Lin Song , Chen Li , Xiaohan Ding , Ying Shan

Unified models aim to support both understanding and generation by encoding images into discrete tokens and processing them alongside text within a single autoregressive framework. This unified design offers architectural simplicity and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Ziyao Wang , Chen Chen , Jingtao Li , Weiming Zhuang , Jiabo Huang , Ang Li , Lingjuan Lyu

Current unified multimodal models typically rely on discrete visual tokenizers to bridge the modality gap. However, discretization inevitably discards fine-grained semantic information, leading to suboptimal performance in visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Yaqi Zhao , Wang Lin , Zijian Zhang , Miles Yang , Jingyuan Chen , Wentao Zhang , Zhao Zhong , Liefeng Bo