English
Related papers

Related papers: XGC-AVis: Towards Audio-Visual Content Understandi…

200 papers

Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Although existing Large Language Model (LLM)-based approaches have advanced long video…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Jiahua Li , Zhanhe Zhang , Chenghao Xu , Zhe Xu , Kun Wei , Xu Yang , Cheng Deng

Multi-modal large language models (MLLMs) have demonstrated considerable potential across various downstream tasks that require cross-domain knowledge. MLLMs capable of processing videos, known as Video-MLLMs, have attracted broad interest…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Jiajun Fei , Dian Li , Zhidong Deng , Zekun Wang , Gang Liu , Hui Wang

Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial…

Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body of work addresses…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Ali Athar , Xueqing Deng , Liang-Chieh Chen

Recently, there is a surge in interest surrounding video large language models (Video LLMs). However, existing benchmarks fail to provide a comprehensive feedback on the temporal perception ability of Video LLMs. On the one hand, most of…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Yuanxin Liu , Shicheng Li , Yi Liu , Yuxiang Wang , Shuhuai Ren , Lei Li , Sishuo Chen , Xu Sun , Lu Hou

Visual compliance verification is a critical yet underexplored problem in computer vision, especially in domains such as media, entertainment, and advertising where content must adhere to complex and evolving policy rules. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Rahul Ghosh , Baishali Chaudhury , Hari Prasanna Das , Meghana Ashok , Ryan Razkenari , Long Chen , Sungmin Hong , Chun-Hao Liu

Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily focus on short audio and video clips ranging from 10 seconds…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Keda Tao , Yuhua Zheng , Jia Xu , Wenjie Du , Kele Shao , Hesong Wang , Xueyi Chen , Xin Jin , Junhan Zhu , Bohan Yu , Weiqiang Wang , Jian Liu , Can Qin , Yulun Zhang , Ming-Hsuan Yang , Huan Wang

We introduce V-Agent, a novel multi-agent platform designed for advanced video search and interactive user-system conversations. By fine-tuning a vision-language model (VLM) with a small video preference dataset and enhancing it with a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 SunYoung Park , Jong-Hyeon Lee , Youngjune Kim , Daegyu Sung , Younghyun Yu , Young-rok Cha , Jeongho Ju

Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of…

Spatial understanding over continuous visual input is crucial for MLLMs to evolve into general-purpose assistants in physical environments. Yet there is still no comprehensive benchmark that holistically assesses the progress toward this…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Jingli Lin , Runsen Xu , Shaohao Zhu , Sihan Yang , Peizhou Cao , Yunlong Ran , Miao Hu , Chenming Zhu , Yiman Xie , Yilin Long , Wenbo Hu , Dahua Lin , Tai Wang , Jiangmiao Pang

Multimodal Large Language Models (MLLMs) struggle with complex video QA benchmarks like HD-EPIC VQA due to ambiguous queries/options, poor long-range temporal reasoning, and non-standardized outputs. We propose a framework integrating…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Sicheng Yang , Yukai Huang , Shitong Sun , Weitong Cai , Jiankang Deng , Jifei Song , Zhensong Zhang

The emergence of advanced multimodal large language models (MLLMs) has significantly enhanced AI assistants' ability to process complex information across modalities. Recently, egocentric videos, by directly capturing user focus, actions,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Taiying Peng , Jiacheng Hua , Miao Liu , Feng Lu

As embodied models become powerful, humans will collaborate with multiple embodied AI agents at their workplace or home in the future. To ensure better communication between human users and the multi-agent system, it is crucial to interpret…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Kangsan Kim , Yanlai Yang , Suji Kim , Woongyeong Yeo , Youngwan Lee , Mengye Ren , Sung Ju Hwang

The recent development of Multimodal Large Language Models (MLLMs) has significantly advanced AI's ability to understand visual modalities. However, existing evaluation benchmarks remain limited to single-turn question answering,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Yaning Pan , Qianqian Xie , Guohui Zhang , Zekun Wang , Yongqian Wen , Yuanxing Zhang , Haoxuan Hu , Zhiyu Pan , Yibing Huang , Zhidong Gan , Yonghong Lin , An Ping , Shihao Li , Yanghai Wang , Tianhao Peng , Jiaheng Liu

Systems that can find correspondences between multiple modalities, such as between speech and images, have great potential to solve different recognition and data analysis tasks in an unsupervised manner. This work studies multimodal…

Computer Vision and Pattern Recognition · Computer Science 2024-03-08 Khazar Khorrami , Okko Räsänen

Accurate visual understanding is imperative for advancing autonomous systems and intelligent robots. Despite the powerful capabilities of vision-language models (VLMs) in processing complex visual scenes, precisely recognizing obscured or…

Computer Vision and Pattern Recognition · Computer Science 2024-06-03 Huaxiang Zhang , Yaojia Mu , Guo-Niu Zhu , Zhongxue Gan

With the advancements in Large Language Models (LLMs), Vision-Language Models (VLMs) have reached a new level of sophistication, showing notable competence in executing intricate cognition and reasoning tasks. However, existing evaluation…

Computer Vision and Pattern Recognition · Computer Science 2023-11-27 Yuanfeng Ji , Chongjian Ge , Weikai Kong , Enze Xie , Zhengying Liu , Zhengguo Li , Ping Luo

There has been a surge of interest in assistive wearable agents: agents embodied in wearable form factors (e.g., smart glasses) who take assistive actions toward a user's goal/query (e.g. "Where did I leave my keys?"). In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Vijay Veerabadran , Fanyi Xiao , Nitin Kamra , Pedro Matias , Joy Chen , Caley Drooff , Brett D Roads , Riley Williams , Ethan Henderson , Xuanyi Zhao , Kevin Carlberg , Joseph Tighe , Karl Ridgeway

The dominant paradigm of monolithic scaling in Vision-Language Models (VLMs) is failing for understanding and reasoning in documents, yielding diminishing returns as it struggles with the inherent need of this domain for document-based…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Xinlei Yu , Chengming Xu , Zhangquan Chen , Yudong Zhang , Shilin Lu , Cheng Yang , Jiangning Zhang , Shuicheng Yan , Xiaobin Hu

Multimodal Large Language Models (MLLMs) are widely regarded as crucial in the exploration of Artificial General Intelligence (AGI). The core of MLLMs lies in their capability to achieve cross-modal alignment. To attain this goal, current…

Computation and Language · Computer Science 2024-11-26 Fei Zhao , Taotian Pang , Chunhui Li , Zhen Wu , Junjie Guo , Shangyu Xing , Xinyu Dai
‹ Prev 1 4 5 6 7 8 10 Next ›