English
Related papers

Related papers: POINTS1.5: Building a Vision-Language Model toward…

200 papers

While Multimodal Large Language Models (MLLMs) have experienced rapid advancements, their visual encoders frequently remain a performance bottleneck. Conventional CLIP-based encoders struggle with dense spatial tasks due to the loss of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Peisen Zhao , Xiaopeng Zhang , Mingxing Xu , Ruoyu Sun , Zewei Du , Dunzheng Wang , Guanghao Zheng , Haohang Xu , Zhibo Zhang , Yuhang Zhang , Yi Ai , Lin Liu , Qi Tian

Recent advances in language modeling have witnessed the rise of highly desirable emergent capabilities, such as reasoning and in-context learning. However, vision models have yet to exhibit comparable progress in these areas. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Jike Zhong , Yuxiang Lai , Xiaofeng Yang , Konstantinos Psounis

Large pre-trained vision-language models like CLIP have shown great potential in learning representations that are transferable across a wide range of downstream tasks. Different from the traditional representation learning that is based…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Kaiyang Zhou , Jingkang Yang , Chen Change Loy , Ziwei Liu

Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and English, have lagged due to the relative scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2024-02-06 Qingpei Guo , Furong Xu , Hanxiao Zhang , Wang Ren , Ziping Ma , Lin Ju , Jian Wang , Jingdong Chen , Ming Yang

Existing open-world universal segmentation approaches usually leverage CLIP and pre-computed proposal masks to treat open-world segmentation tasks as proposal classification. However, 1) these works cannot handle universal segmentation in…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Bowen Dong , Jiaxi Gu , Jianhua Han , Hang Xu , Wangmeng Zuo

Charts are essential to data analysis, transforming raw data into clear visual representations that support human decision-making. Although current vision-language models (VLMs) have made significant progress, they continue to struggle with…

Vision-language models (VLMs) extend the conventional large language models by integrating visual data, enabling richer multimodal reasoning and significantly broadens the practical applications of AI. However, including visual inputs also…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Daulet Toibazar , Kesen Wang , Sherif Mohamed , Abdulaziz Al-Badawi , Abdulrahman Alfulayt , Pedro J. Moreno

We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter vision encoder and a Mixture-of-Experts (MoE) LLM of 20B…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Dong Guo , Faming Wu , Feida Zhu , Fuxing Leng , Guang Shi , Haobin Chen , Haoqi Fan , Jian Wang , Jianyu Jiang , Jiawei Wang , Jingji Chen , Jingjia Huang , Kang Lei , Liping Yuan , Lishu Luo , Pengfei Liu , Qinghao Ye , Rui Qian , Shen Yan , Shixiong Zhao , Shuai Peng , Shuangye Li , Sihang Yuan , Sijin Wu , Tianheng Cheng , Weiwei Liu , Wenqian Wang , Xianhan Zeng , Xiao Liu , Xiaobo Qin , Xiaohan Ding , Xiaojun Xiao , Xiaoying Zhang , Xuanwei Zhang , Xuehan Xiong , Yanghua Peng , Yangrui Chen , Yanwei Li , Yanxu Hu , Yi Lin , Yiyuan Hu , Yiyuan Zhang , Youbin Wu , Yu Li , Yudong Liu , Yue Ling , Yujia Qin , Zanbo Wang , Zhiwu He , Aoxue Zhang , Bairen Yi , Bencheng Liao , Can Huang , Can Zhang , Chaorui Deng , Chaoyi Deng , Cheng Lin , Cheng Yuan , Chenggang Li , Chenhui Gou , Chenwei Lou , Chengzhi Wei , Chundian Liu , Chunyuan Li , Deyao Zhu , Donghong Zhong , Feng Li , Feng Zhang , Gang Wu , Guodong Li , Guohong Xiao , Haibin Lin , Haihua Yang , Haoming Wang , Heng Ji , Hongxiang Hao , Hui Shen , Huixia Li , Jiahao Li , Jialong Wu , Jianhua Zhu , Jianpeng Jiao , Jiashi Feng , Jiaze Chen , Jianhui Duan , Jihao Liu , Jin Zeng , Jingqun Tang , Jingyu Sun , Joya Chen , Jun Long , Junda Feng , Junfeng Zhan , Junjie Fang , Junting Lu , Kai Hua , Kai Liu , Kai Shen , Kaiyuan Zhang , Ke Shen , Ke Wang , Keyu Pan , Kun Zhang , Kunchang Li , Lanxin Li , Lei Li , Lei Shi , Li Han , Liang Xiang , Liangqiang Chen , Lin Chen , Lin Li , Lin Yan , Liying Chi , Longxiang Liu , Mengfei Du , Mingxuan Wang , Ningxin Pan , Peibin Chen , Pengfei Chen , Pengfei Wu , Qingqing Yuan , Qingyao Shuai , Qiuyan Tao , Renjie Zheng , Renrui Zhang , Ru Zhang , Rui Wang , Rui Yang , Rui Zhao , Shaoqiang Xu , Shihao Liang , Shipeng Yan , Shu Zhong , Shuaishuai Cao , Shuangzhi Wu , Shufan Liu , Shuhan Chang , Songhua Cai , Tenglong Ao , Tianhao Yang , Tingting Zhang , Wanjun Zhong , Wei Jia , Wei Weng , Weihao Yu , Wenhao Huang , Wenjia Zhu , Wenli Yang , Wenzhi Wang , Xiang Long , XiangRui Yin , Xiao Li , Xiaolei Zhu , Xiaoying Jia , Xijin Zhang , Xin Liu , Xinchen Zhang , Xinyu Yang , Xiongcai Luo , Xiuli Chen , Xuantong Zhong , Xuefeng Xiao , Xujing Li , Yan Wu , Yawei Wen , Yifan Du , Yihao Zhang , Yining Ye , Yonghui Wu , Yu Liu , Yu Yue , Yufeng Zhou , Yufeng Yuan , Yuhang Xu , Yuhong Yang , Yun Zhang , Yunhao Fang , Yuntao Li , Yurui Ren , Yuwen Xiong , Zehua Hong , Zehua Wang , Zewei Sun , Zeyu Wang , Zhao Cai , Zhaoyue Zha , Zhecheng An , Zhehui Zhao , Zhengzhuo Xu , Zhipeng Chen , Zhiyong Wu , Zhuofan Zheng , Zihao Wang , Zilong Huang , Ziyu Zhu , Zuquan Song

Current vision-language models (VLMs) are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Dewen Zhang , Tahir Hussain , Wangpeng An , Hayaru Shouno

We introduce SuperClass, a super simple classification method for vision-language pre-training on image-text data. Unlike its contrastive counterpart CLIP who contrast with a text encoder, SuperClass directly utilizes tokenized raw text as…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Zilong Huang , Qinghao Ye , Bingyi Kang , Jiashi Feng , Haoqi Fan

Image-text training like CLIP has dominated the pretraining of vision foundation models in recent years. Subsequent efforts have been made to introduce region-level visual learning into CLIP's pretraining but face scalability challenges due…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Xiaohu Jiang , Yixiao Ge , Yuying Ge , Dachuan Shi , Chun Yuan , Ying Shan

Large Vision-Language Models (LVLMs) have shown strong performance across various multimodal tasks by leveraging the reasoning capabilities of Large Language Models (LLMs). However, processing visually complex and information-rich images,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Mincheol Kwon , Minseung Lee , Seonga Choi , Miso Choi , Kyeong-Jin Oh , Hyunyoung Lee , Cheonyoung Park , Yongho Song , Seunghyun Park , Jinkyu Kim

Large Multimodal Models (LMMs) have achieved strong performance in vision-language understanding, yet many existing approaches rely on large-scale architectures and coarse supervision, which limits their ability to generate detailed image…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Jiaxin Fan , Wenpo Song

This work proposes POMP, a prompt pre-training method for vision-language models. Being memory and computation efficient, POMP enables the learned prompt to condense semantic information for a rich set of visual concepts with over…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Shuhuai Ren , Aston Zhang , Yi Zhu , Shuai Zhang , Shuai Zheng , Mu Li , Alex Smola , Xu Sun

In this report, we introduce InternVL 1.5, an open-source multimodal large language model (MLLM) to bridge the capability gap between open-source and proprietary commercial models in multimodal understanding. We introduce three simple…

Multimodal pre-trained models, such as CLIP, are popular for zero-shot classification due to their open-vocabulary flexibility and high performance. However, vision-language models, which compute similarity scores between images and class…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Mia Chiquier , Utkarsh Mall , Carl Vondrick

Multimodal Vision Language Models (VLMs) have emerged as a transformative topic at the intersection of computer vision and natural language processing, enabling machines to perceive and reason about the world through both visual and textual…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Zongxia Li , Xiyang Wu , Hongyang Du , Fuxiao Liu , Huy Nghiem , Guangyao Shi

As a pioneering vision-language model, CLIP (Contrastive Language-Image Pre-training) has achieved significant success across various domains and a wide range of downstream vision-language tasks. However, the text encoders in popular CLIP…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Mothilal Asokan , Kebin Wu , Fatima Albreiki

Previous works show that noisy, web-crawled image-text pairs may limit vision-language pretraining like CLIP and propose learning with synthetic captions as a promising alternative. Our work continues this effort, introducing two simple yet…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Yanqing Liu , Xianhang Li , Zeyu Wang , Bingchen Zhao , Cihang Xie

As Vision Transformers (ViTs) are increasingly adopted in sensitive vision applications, there is a growing demand for improved interpretability. This has led to efforts to forward-align these models with carefully annotated abstract,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Sanchit Sinha , Guangzhi Xiong , Aidong Zhang