English
Related papers

Related papers: Qwen2.5-VL Technical Report

200 papers

Vision-Language Models (VLMs) achieve outstanding performance, yet their huge model size severely hinders deployment on edge devices with limited resources. As an efficient model compression technique, vector quantization (VQ) excels in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Zhong Wang , Zukang Xu , Xing Hu , Dawei Yang

Pre-trained language models (PLMs) have played an increasing role in multimedia research. In terms of vision-language (VL) tasks, they often serve as a language encoder and still require an additional fusion network for VL reasoning,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Shubin Huang , Qiong Wu , Yiyi Zhou , Weijie Chen , Rongsheng Zhang , Xiaoshuai Sun , Rongrong Ji

Most multimodal large language models (MLLMs) treat visual tokens as "a sequence of text", integrating them with text tokens into a large language model (LLM). However, a great quantity of visual tokens significantly increases the demand…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Dongchen Lu , Yuyao Sun , Zilu Zhang , Leping Huang , Jianliang Zeng , Mao Shu , Huo Cao

Beginning with VisualGLM and CogVLM, we are continuously exploring VLMs in pursuit of enhanced vision-language fusion, efficient higher-resolution architecture, and broader modalities and applications. Here we propose the CogVLM2 family, a…

Despite recent advancements in Multi-modal Large Language Models (MLLMs) on diverse understanding tasks, these models struggle to solve problems which require extensive multi-step reasoning. This is primarily due to the progressive dilution…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Byungwoo Jeon , Yoonwoo Jeong , Hyunseok Lee , Minsu Cho , Jinwoo Shin

We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tasks (e.g., VQA, image captioning). GLIPv2 elegantly unifies…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Haotian Zhang , Pengchuan Zhang , Xiaowei Hu , Yen-Chun Chen , Liunian Harold Li , Xiyang Dai , Lijuan Wang , Lu Yuan , Jenq-Neng Hwang , Jianfeng Gao

Vision-Language Models (VLMs) excel at reasoning in linguistic space but struggle with perceptual understanding that requires dense visual perception, e.g., spatial reasoning and geometric awareness. This limitation stems from the fact that…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Yiming Qin , Bomin Wei , Jiaxin Ge , Konstantinos Kallidromitis , Stephanie Fu , Trevor Darrell , XuDong Wang

The integration of visual inputs with large language models (LLMs) has led to remarkable advancements in multi-modal capabilities, giving rise to visual large language models (VLLMs). However, effectively harnessing VLLMs for intricate…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Renjie Pi , Lewei Yao , Jiahui Gao , Jipeng Zhang , Tong Zhang

Spatial perception and reasoning are crucial for Vision-Language-Action (VLA) models to accomplish fine-grained manipulation tasks. However, existing approaches often lack the ability to understand and reason over the essential 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Yixuan Li , Yuhui Chen , Mingcai Zhou , Haoran Li , Zhengtao Zhang , Dongbin Zhao

Robots operating in shared human environments must not only navigate, interact, and detect their surroundings, they must also interpret and respond to dynamic, and often unpredictable, human behaviours. Although recent advances have shown…

Vision-and-Language Navigation (VLN) is a core task where embodied agents leverage their spatial mobility to navigate in 3D environments toward designated destinations based on natural language instructions. Recently, video-language large…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Zihan Wang , Seungjun Lee , Gim Hee Lee

Understanding where drivers direct their visual attention during driving, as characterized by gaze behavior, is critical for developing next-generation advanced driver-assistance systems and improving road safety. This paper tackles this…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Penghao Deng , Jidong J. Yang , Jiachen Bian

Recently, due to the advancement of multimodal technology, people are attempting to use visual large language models (VLLMs) in industrial production. Many deep learning models (DLMs) deployed in the production environment are gradually…

Computation and Language · Computer Science 2026-01-26 Xiang Chen

Lately, researchers in artificial intelligence have been really interested in how language and vision come together, giving rise to the development of multimodal models that aim to seamlessly integrate textual and visual information.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Rajat Chawla , Arkajit Datta , Tushar Verma , Adarsh Jha , Anmol Gautam , Ayush Vatsal , Sukrit Chaterjee , Mukunda NS , Ishaan Bhola

Recent advances in vision-language models (VLMs) have demonstrated the advantages of processing images at higher resolutions and utilizing multi-crop features to preserve native resolution details. However, despite these improvements,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Rahul Thapa , Kezhen Chen , Ian Covert , Rahul Chalamala , Ben Athiwaratkun , Shuaiwen Leon Song , James Zou

As Vision-Language Models (VLMs) grow in sophistication, their ability to perform reasoning is coming under increasing supervision. While they excel at many tasks, their grasp of fundamental scientific principles, such as physics, remains…

Machine Learning · Computer Science 2025-09-11 Pranav Pawar , Kavish Shah , Akshat Bhalani , Komal Kasat , Dev Mittal , Hadi Gala , Deepali Patil , Nikita Raichada , Monali Deshmukh

The limited capacity for fine-grained visual perception presents a critical bottleneck for Vision-Language Models (VLMs) in real-world applications. Addressing this is challenging due to the scarcity of high-quality data and the limitations…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Juntian Zhang , Song Jin , Chuanqi Cheng , Yuhan Liu , Yankai Lin , Xun Zhang , Yufei Zhang , Fei Jiang , Guojun Yin , Wei Lin , Rui Yan

We introduce Nemotron Nano V2 VL, the latest model of the Nemotron vision-language series designed for strong real-world document understanding, long video comprehension, and reasoning tasks. Nemotron Nano V2 VL delivers significant…

Machine Learning · Computer Science 2025-11-10 NVIDIA , : , Amala Sanjay Deshmukh , Kateryna Chumachenko , Tuomas Rintamaki , Matthieu Le , Tyler Poon , Danial Mohseni Taheri , Ilia Karmanov , Guilin Liu , Jarno Seppanen , Guo Chen , Karan Sapra , Zhiding Yu , Adi Renduchintala , Charles Wang , Peter Jin , Arushi Goel , Mike Ranzinger , Lukas Voegtle , Philipp Fischer , Timo Roman , Wei Ping , Boxin Wang , Zhuolin Yang , Nayeon Lee , Shaokun Zhang , Fuxiao Liu , Zhiqi Li , Di Zhang , Greg Heinrich , Hongxu Yin , Song Han , Pavlo Molchanov , Parth Mannan , Yao Xu , Jane Polak Scowcroft , Tom Balough , Subhashree Radhakrishnan , Paris Zhang , Sean Cha , Ratnesh Kumar , Zaid Pervaiz Bhat , Jian Zhang , Darragh Hanley , Pritam Biswas , Jesse Oliver , Kevin Vasques , Roger Waleffe , Duncan Riach , Oluwatobi Olabiyi , Ameya Sunil Mahabaleshwarkar , Bilal Kartal , Pritam Gundecha , Khanh Nguyen , Alexandre Milesi , Eugene Khvedchenia , Ran Zilberstein , Ofri Masad , Natan Bagrov , Nave Assaf , Tomer Asida , Daniel Afrimi , Amit Zuker , Netanel Haber , Zhiyu Cheng , Jingyu Xin , Di Wu , Nik Spirin , Maryam Moosaei , Roman Ageev , Vanshil Atul Shah , Yuting Wu , Daniel Korzekwa , Unnikrishnan Kizhakkemadam Sreekumar , Wanli Jiang , Padmavathy Subramanian , Alejandra Rico , Sandip Bhaskar , Saeid Motiian , Kedi Wu , Annie Surla , Chia-Chih Chen , Hayden Wolff , Matthew Feinberg , Melissa Corpuz , Marek Wawrzos , Eileen Long , Aastha Jhunjhunwala , Paul Hendricks , Farzan Memarian , Benika Hall , Xin-Yu Wang , David Mosallanezhad , Soumye Singhal , Luis Vega , Katherine Cheung , Krzysztof Pawelec , Michael Evans , Katherine Luna , Jie Lou , Erick Galinkin , Akshay Hazare , Kaustubh Purandare , Ann Guan , Anna Warno , Chen Cui , Yoshi Suhara , Shibani Likhite , Seph Mard , Meredith Price , Laya Sleiman , Saori Kaji , Udi Karpas , Kari Briski , Joey Conway , Michael Lightstone , Jan Kautz , Mohammad Shoeybi , Mostofa Patwary , Jonathen Cohen , Oleksii Kuchaiev , Andrew Tao , Bryan Catanzaro

Empowered by Large Language Models (LLMs), recent advancements in Video-based LLMs (VideoLLMs) have driven progress in various video understanding tasks. These models encode video representations through pooling or query aggregation over a…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Yuetian Weng , Mingfei Han , Haoyu He , Xiaojun Chang , Bohan Zhuang

Vision-Language Models (VLMs) have shown promising capabilities in handling various multimodal tasks, yet they struggle in long-context scenarios, particularly in tasks involving videos, high-resolution images, or lengthy image-text…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Junqi Ge , Ziyi Chen , Jintao Lin , Jinguo Zhu , Xihui Liu , Jifeng Dai , Xizhou Zhu
‹ Prev 1 4 5 6 7 8 10 Next ›