English
Related papers

Related papers: PairUni: Pairwise Training for Unified Multimodal …

200 papers

The diagnosis of pathological images is often limited by expert availability and regional disparities, highlighting the importance of automated diagnosis using Vision-Language Models (VLMs). Traditional multimodal models typically emphasize…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Jianyu Wu , Hao Yang , Xinhua Zeng , Guibing He , Zhiyu Chen , Zihui Li , Xiaochuan Zhang , Yangyang Ma , Run Fang , Yang Liu

Unified multimodal models integrate the reasoning capacity of large language models with both image understanding and generation, showing great promise for advanced multimodal intelligence. However, the community still lacks a rigorous…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Hongxiang Li , Yaowei Li , Bin Lin , Yuwei Niu , Yuhang Yang , Xiaoshuang Huang , Jiayin Cai , Xiaolong Jiang , Yao Hu , Long Chen

Existing methods for vision-and-language learning typically require designing task-specific architectures and objectives for each task. For example, a multi-label answer classifier for visual question answering, a region scorer for…

Computation and Language · Computer Science 2021-05-25 Jaemin Cho , Jie Lei , Hao Tan , Mohit Bansal

Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation. However, these two capabilities remain largely independent, as if they are two separate functions encapsulated within the same…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Kaihang Pan , Yang Wu , Wendong Bu , Kai Shen , Juncheng Li , Yingting Wang , Yunfei Li , Siliang Tang , Jun Xiao , Fei Wu , Hang Zhao , Yueting Zhuang

In this work, we present an information-theoretic framework that formulates cross-lingual language model pre-training as maximizing mutual information between multilingual-multi-granularity texts. The unified view helps us to better…

Computation and Language · Computer Science 2021-04-08 Zewen Chi , Li Dong , Furu Wei , Nan Yang , Saksham Singhal , Wenhui Wang , Xia Song , Xian-Ling Mao , Heyan Huang , Ming Zhou

We introduce Ming-Lite-Uni, an open-source multimodal framework featuring a newly designed unified visual generator and a native multimodal autoregressive model tailored for unifying vision and language. Specifically, this project provides…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Inclusion AI , Biao Gong , Cheng Zou , Dandan Zheng , Hu Yu , Jingdong Chen , Jianxin Sun , Junbo Zhao , Jun Zhou , Kaixiang Ji , Lixiang Ru , Libin Wang , Qingpei Guo , Rui Liu , Weilong Chai , Xinyu Xiao , Ziyuan Huang

Emotional understanding and generation are often treated as separate tasks, yet they are inherently complementary and can mutually enhance each other. In this paper, we propose the UniEmo, a unified framework that seamlessly integrates…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Yijie Zhu , Lingsen Zhang , Zitong Yu , Rui Shao , Tao Tan , Liqiang Nie

Recent advances in the audio language modeling (ALM) domain tackle audio understanding and text-to-audio generation as separate tasks. Very few studies attempt to unify these tasks -- an essential step toward advanced multimodal reasoning.…

Recent advances in vision-language-action (VLA) models have motivated the extension of their capabilities to embodied settings, where reinforcement learning (RL) offers a principled way to optimize task success through interaction. However,…

Vision language decision making (VLDM) is a challenging multimodal task. The agent have to understand complex human instructions and complete compositional tasks involving environment navigation and object manipulation. However, the long…

Computer Vision and Pattern Recognition · Computer Science 2023-07-18 Ruipu Luo , Jiwen Zhang , Zhongyu Wei

Recent advances in deep learning have witnessed many successful unsupervised image-to-image translation models that learn correspondences between two visual domains without paired data. However, it is still a great challenge to build robust…

Computer Vision and Pattern Recognition · Computer Science 2023-06-08 Shuai Yang , Liming Jiang , Ziwei Liu , Chen Change Loy

Building a universal Video-Language model for solving various video understanding tasks (\emph{e.g.}, text-video retrieval, video question answering) is an open challenge to the machine learning field. Towards this goal, most recent works…

Computer Vision and Pattern Recognition · Computer Science 2022-12-21 Jingjia Huang , Yinan Li , Jiashi Feng , Xinglong Wu , Xiaoshuai Sun , Rongrong Ji

Large Language Models (LLMs) have shown strong capabilities through two complementary paradigms: Retrieval-Augmented Generation (RAG) for knowledge grounding and Reinforcement Learning from Verifiable Rewards (RLVR) for complex reasoning.…

Computation and Language · Computer Science 2026-04-27 Weitao Li , Boran Xiang , Xiaolong Wang , Zhinan Gou , Weizhi Ma , Yang Liu

Text-to-image generation has evolved beyond single monolithic models to complex multi-component pipelines. These combine fine-tuned generators, adapters, upscaling blocks and even editing steps, leading to significant improvements in image…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Uri Gadot , Rinon Gal , Yftah Ziser , Gal Chechik , Shie Mannor

World models - generative models that simulate environment dynamics conditioned on past observations and actions - are gaining prominence in planning, simulation, and embodied AI. However, evaluating their rollouts remains a fundamental…

Designing effective data manipulation methods is a long standing problem in data lakes. Traditional methods, which rely on rules or machine learning models, require extensive human efforts on training data collection and tuning models.…

Artificial Intelligence · Computer Science 2024-05-13 Yichen Qian , Yongyi He , Rong Zhu , Jintao Huang , Zhijian Ma , Haibin Wang , Yaohua Wang , Xiuyu Sun , Defu Lian , Bolin Ding , Jingren Zhou

Text-to-image generation has advanced rapidly with diffusion models, progressing from CLIP and T5 conditioning to unified systems where a single LLM backbone handles both visual understanding and generation. Despite the architectural…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Sucheng Ren , Chen Chen , Zhenbang Wang , Liangchen Song , Xiangxin Zhu , Alan Yuille , Liang-Chieh Chen , Jiasen Lu

Unified multimodal large language models such as Show-o and Janus have achieved strong performance across both generation and understanding tasks. However, these models typically rely on large-scale datasets and require substantial…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Weijia Mao , Zhenheng Yang , Mike Zheng Shou

Heterogeneous ensembles that can aggregate an unrestricted number and variety of base predictors can effectively address challenging prediction problems. In particular, accurate ensembles that are also parsimonious, i.e., consist of as few…

Machine Learning · Computer Science 2021-03-01 Ana Stanescu , Gaurav Pandey

Deep metric learning (DML) has received much attention in deep learning due to its wide applications in computer vision. Previous studies have focused on designing complicated losses and hard example mining methods, which are mostly…

Machine Learning · Computer Science 2020-06-19 Qi Qi , Yan Yan , Xiaoyu Wang , Tianbao Yang
‹ Prev 1 4 5 6 7 8 10 Next ›