English
Related papers

Related papers: Generation Enhances Understanding in Unified Multi…

200 papers

Multi-modal medical image completion has been extensively applied to alleviate the missing modality issue in a wealth of multi-modal diagnostic tasks. However, for most existing synthesis methods, their inferences of missing modalities can…

Image and Video Processing · Electrical Eng. & Systems 2022-07-08 Xiangxi Meng , Yuning Gu , Yongsheng Pan , Nizhuan Wang , Peng Xue , Mengkang Lu , Xuming He , Yiqiang Zhan , Dinggang Shen

We present MMCORE, a unified framework designed for multimodal image generation and editing. MMCORE leverages a pre-trained Vision-Language Model (VLM) to predict semantic visual embeddings via learnable query tokens, which subsequently…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Zijie Li , Yichun Shi , Jingxiang Sun , Ye Wang , Yixuan Huang , Zhiyao Guo , Xiaochen Lian , Peihao Zhu , Yu Tian , Zhonghua Zhai , Peng Wang

Recent advances in the audio language modeling (ALM) domain tackle audio understanding and text-to-audio generation as separate tasks. Very few studies attempt to unify these tasks -- an essential step toward advanced multimodal reasoning.…

The rapid progress of large multimodal models has inspired efforts toward unified frameworks that couple understanding and generation. While such paradigms have shown remarkable success in 2D, extending them to 3D remains largely…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Yongwei Chen , Tianyi Wei , Yushi Lan , Zhaoyang Lyu , Shangchen Zhou , Xudong Xu , Xingang Pan

Retrieval-augmented generation (RAG) enables large language models (LLMs) to access external knowledge, helping mitigate hallucinations and enhance domain-specific expertise. Graph-based RAG enhances structural reasoning by introducing…

Computation and Language · Computer Science 2025-11-26 Linxiao Cao , Ruitao Wang , Jindong Li , Zhipeng Zhou , Menglin Yang

Humans understand the world through the integration of multiple sensory modalities, enabling them to perceive, reason about, and imagine dynamic physical processes. Inspired by this capability, multimodal foundation models (MFMs) have…

Artificial Intelligence · Computer Science 2025-10-07 Xuehai He

Large-scale pre-trained Vision-Language Models (VLMs) have significantly advanced transfer learning across diverse tasks. However, adapting these models with limited few-shot data often leads to overfitting, undermining their ability to…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Yuncheng Guo , Xiaodong Gu

Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVideo, a versatile framework that extends unified modeling to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Cong Wei , Quande Liu , Zixuan Ye , Qiulin Wang , Xintao Wang , Pengfei Wan , Kun Gai , Wenhu Chen

Multi-modal learning is a fast growing area in artificial intelligence. It tries to help machines understand complex things by combining information from different sources, like images, text, and audio. By using the strengths of each…

Machine Learning · Computer Science 2025-12-22 Qihang Jin , Enze Ge , Yuhang Xie , Hongying Luo , Junhao Song , Ziqian Bi , Chia Xin Liang , Jibin Guan , Joe Yeong , Xinyuan Song , Junfeng Hao

Personalized models have demonstrated remarkable success in understanding and generating concepts provided by users. However, existing methods use separate concept tokens for understanding and generation, treating these tasks in isolation.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Ruichuan An , Sihan Yang , Renrui Zhang , Zijun Shen , Ming Lu , Gaole Dai , Hao Liang , Ziyu Guo , Shilin Yan , Yulin Luo , Bocheng Zou , Chaoqun Yang , Wentao Zhang

Recent advances in Retrieval-Augmented Generation (RAG) have significantly improved response accuracy and relevance by incorporating external knowledge into Large Language Models (LLMs). However, existing RAG methods primarily focus on…

Machine Learning · Computer Science 2025-04-22 Qinhan Yu , Zhiyou Xiao , Binghui Li , Zhengren Wang , Chong Chen , Wentao Zhang

Retrieval-Augmented Generation (RAG) has emerged as an effective paradigm for expanding the knowledge capacity of Multimodal Large Language Models (MLLMs) by incorporating external knowledge sources into the generation process, and has been…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Xu Yuan , Liangbo Ning , Qingqing Ye , Wenqi Fan , Qing Li

Understanding and replicating the real world is a critical challenge in Artificial General Intelligence (AGI) research. To achieve this, many existing approaches, such as world models, aim to capture the fundamental principles governing the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Yuqi Hu , Longguang Wang , Xian Liu , Ling-Hao Chen , Yuwei Guo , Yukai Shi , Ce Liu , Anyi Rao , Zeyu Wang , Hui Xiong

Understanding visual art requires reasoning across multiple perspectives -- cultural, historical, and stylistic -- beyond mere object recognition. While recent multimodal large language models (MLLMs) perform well on general image…

Artificial Intelligence · Computer Science 2025-09-08 Shuai Wang , Ivona Najdenkoska , Hongyi Zhu , Stevan Rudinac , Monika Kackovic , Nachoem Wijnberg , Marcel Worring

Existing text-to-image diffusion models primarily generate images from text prompts. However, the inherent conciseness of textual descriptions poses challenges in faithfully synthesizing images with intricate details, such as specific…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Wei Li , Xue Xu , Jiachen Liu , Xinyan Xiao

Effectively retrieving, reasoning, and understanding multimodal information remains a critical challenge for agentic systems. Traditional Retrieval-augmented Generation (RAG) methods rely on linear interaction histories, which struggle to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Qiuchen Wang , Shihang Wang , Yu Zeng , Qiang Zhang , Fanrui Zhang , Zhuoning Guo , Bosi Zhang , Wenxuan Huang , Lin Chen , Zehui Chen , Pengjun Xie , Ruixue Ding

Following rapid advancements in text and image generation, research has increasingly shifted towards 3D generation. Unlike the well-established pixel-based representation in images, 3D representations remain diverse and fragmented,…

Multimodality Representation Learning, as a technique of learning to embed information from different modalities and their correlations, has achieved remarkable success on a variety of applications, such as Visual Question Answering (VQA),…

Artificial Intelligence · Computer Science 2024-03-04 Muhammad Arslan Manzoor , Sarah Albarri , Ziting Xian , Zaiqiao Meng , Preslav Nakov , Shangsong Liang

Retrieval-augmented generation (RAG) is a paradigm that augments large language models (LLMs) with external knowledge to tackle knowledge-intensive question answering. While several benchmarks evaluate Multimodal LLMs (MLLMs) under…

Computation and Language · Computer Science 2025-08-18 Yin Wu , Quanyu Long , Jing Li , Jianfei Yu , Wenya Wang

Recent unified image generation models have achieved remarkable success by employing MLLMs for semantic understanding and diffusion backbones for image generation. However, these models remain fundamentally limited in spatially-aware tasks…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Haiyi Qiu , Kaihang Pan , Jiacheng Li , Juncheng Li , Siliang Tang , Yueting Zhuang
‹ Prev 1 8 9 10 Next ›