English
Related papers

Related papers: RE-VLM: Event-Augmented Vision-Language Model for …

200 papers

Event cameras are novel sensors that report brightness changes in the form of a stream of asynchronous "events" instead of intensity frames. They offer significant advantages with respect to conventional cameras: high temporal resolution,…

Computer Vision and Pattern Recognition · Computer Science 2019-06-19 Henri Rebecq , René Ranftl , Vladlen Koltun , Davide Scaramuzza

Event stream based scene text recognition is a newly arising research topic in recent years which performs better than the widely used RGB cameras in extremely challenging scenarios, especially the low illumination, fast motion. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Xiao Wang , Jingtao Jiang , Qiang Chen , Lan Chen , Lin Zhu , Yaowei Wang , Yonghong Tian , Jin Tang

As multimodal language models advance, their application to 3D scene understanding is a fast-growing frontier, driving the development of 3D Vision-Language Models (VLMs). Current methods show strong dependence on object detectors,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Anna-Maria Halacheva , Jan-Nico Zaech , Xi Wang , Danda Pani Paudel , Luc Van Gool

Vision-Language Models (VLMs) excel in diverse visual tasks but face challenges in document understanding, which requires fine-grained text processing. While typical visual tasks perform well with low-resolution inputs, reading-intensive…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Mor Shpigel Nacson , Aviad Aberdam , Roy Ganz , Elad Ben Avraham , Alona Golts , Yair Kittenplon , Shai Mazor , Ron Litman

Most adaptive AR storytelling systems define environmental semantics using simple object labels and spatial coordinates, limiting narratives to rigid, pre-defined logic. This oversimplification overlooks the contextual significance of…

Human-Computer Interaction · Computer Science 2025-04-18 Yusi Sun , Haoyan Guan , leith Kin Yep Chan , Yong Hong Kuo

Following the recent popularity of Large Language Models (LLMs), several attempts have been made to extend them to the visual domain. From having a visual assistant that could guide us through unfamiliar environments to generative models…

Recent progress in VLMs has demonstrated impressive capabilities across a variety of tasks in the natural image domain. Motivated by these advancements, the remote sensing community has begun to adopt VLMs for remote sensing vision-language…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Congcong Wen , Yiting Lin , Xiaokang Qu , Nan Li , Yong Liao , Xiang Li , Hui Lin

Action recognition and localization in complex, untrimmed videos remain a formidable challenge in computer vision, largely due to the limitations of existing methods in capturing fine-grained actions, long-term temporal dependencies, and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Liyang Peng , Sihan Zhu , Yunjie Guo

Vision-language models (VLMs) excel in various visual benchmarks but are often constrained by the lack of high-quality visual fine-tuning data. To address this challenge, we introduce VisCon-100K, a novel dataset derived from interleaved…

Computation and Language · Computer Science 2025-02-25 Gokul Karthik Kumar , Iheb Chaabane , Kebin Wu

Recent advancements in dialogue systems have highlighted the significance of integrating multimodal responses, which enable conveying ideas through diverse modalities rather than solely relying on text-based interactions. This enrichment…

Computation and Language · Computer Science 2024-07-08 Chang-Sheng Kao , Yun-Nung Chen

Vision-Language Models (VLMs) are increasingly deployed in real-time applications such as autonomous driving and human-computer interaction, which demand fast and reliable responses based on accurate perception. To meet these requirements,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Chen Qian , Xinran Yu , Zewen Huang , Danyang Li , Qiang Ma , Fan Dang , Xuan Ding , Guangyong Shang , Zheng Yang

This paper explores the potential of event cameras to enable continuous time reinforcement learning. We formalise this problem where a continuous stream of unsynchronised observations is used to produce a corresponding stream of output…

Computer Vision and Pattern Recognition · Computer Science 2023-02-16 Celyn Walters , Simon Hadfield

We study the problem of estimating optical flow from event cameras. One important issue is how to build a high-quality event-flow dataset with accurate event values and flow labels. Previous datasets are created by either capturing real…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Xinglong Luo , Kunming Luo , Ao Luo , Zhengning Wang , Ping Tan , Shuaicheng Liu

Vision-language models (VLMs) have recently emerged as a promising paradigm for video anomaly detection (VAD) due to their strong visual reasoning ability and natural language-based explainability. In this paper, we aim to address a key…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Mitchell Piehl , Muchao Ye

Multimodal large language models (MLLMs) have achieved impressive performance across various tasks such as image captioning and visual question answer(VQA); however, they often struggle to accurately interpret depth information inherent in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Hao Yang , Hongbo Zhang , Yanyan Zhao , Bing Qin

Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps visual features generated by a vision encoder to a shared…

Large vision-language models (VLMs) have shown promising capabilities in scene understanding, enhancing the explainability of driving behaviors and interactivity with users. Existing methods primarily fine-tune VLMs on on-board multi-view…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Nan Song , Bozhou Zhang , Xiatian Zhu , Jiankang Deng , Li Zhang

Multimodal large language models (MLLMs) have achieved impressive performance on visual perception and reasoning tasks with RGB imagery, yet they remain fragile under common degradations, such as fog, blur, or low-light conditions. Infrared…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Abrar Majeedi , Zhiyuan Ruan , Ziyi Zhao , Hongcheng Wang , Jianglin Lu , Yin Li

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, their application to safety-critical driving scenarios remains limited by an inability to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Tomaso Trinci , Henrique Piñeiro Monteagudo , Leonardo Taccari

Event cameras sense intensity changes and have many advantages over conventional cameras. To take advantage of event cameras, some methods have been proposed to reconstruct intensity images from event streams. However, the outputs are still…

Computer Vision and Pattern Recognition · Computer Science 2020-03-18 Lin Wang , Tae-Kyun Kim , Kuk-Jin Yoon