English
Related papers

Related papers: Visual Grounding from Event Cameras

200 papers

Multi-view visual reasoning is essential for intelligent systems that must understand complex environments from sparse and discrete viewpoints, yet existing research has largely focused on single-image or temporally dense video settings. In…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Fucai Ke , Zhixi Cai , Boying Li , Long Chen , Beibei Lin , Weiqing Wang , Pari Delir Haghighi , Gholamreza Haffari , Hamid Rezatofighi

Visual-language grounding aims to establish semantic correspondences between natural language and visual entities, enabling models to accurately identify and localize target objects based on textual instructions. Existing VLG approaches…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Linfei Li , Lin Zhang , Ying Shen

Event cameras are a paradigm shift in camera technology. Instead of full frames, the sensor captures a sparse set of events caused by intensity changes. Since only the changes are transferred, those cameras are able to capture quick…

Computer Vision and Pattern Recognition · Computer Science 2017-03-22 Christian Reinbacher , Gottfried Munda , Thomas Pock

We propose a model to learn visually grounded word embeddings (vis-w2v) to capture visual notions of semantic relatedness. While word embeddings trained using text have been extremely successful, they cannot uncover notions of semantic…

Computer Vision and Pattern Recognition · Computer Science 2016-06-30 Satwik Kottur , Ramakrishna Vedantam , José M. F. Moura , Devi Parikh

Recent DETR-based video grounding models have made the model directly predict moment timestamps without any hand-crafted components, such as a pre-defined proposal or non-maximum suppression, by learning moment queries. However, their…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Jinhyun Jang , Jungin Park , Jin Kim , Hyeongjun Kwon , Kwanghoon Sohn

Event cameras offering high dynamic range and low latency have emerged as disruptive technologies in imaging. Despite growing research on leveraging these benefits for different imaging tasks, a comprehensive study of recently advances and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Yunfan Lu , Xiaogang Xu , Pengteng Li , Yusheng Wang , Yi Cui , Huizai Yao , Hui Xiong

We address the problem of video captioning by grounding language generation on object interactions in the video. Existing work mostly focuses on overall scene understanding with often limited or no emphasis on object interactions to address…

Computer Vision and Pattern Recognition · Computer Science 2017-11-20 Chih-Yao Ma , Asim Kadav , Iain Melvin , Zsolt Kira , Ghassan AlRegib , Hans Peter Graf

In recent years, a substantial body of work in visually grounded natural language processing has focused on real-life multimodal scenarios such as describing content depicted in images or videos. However, comparatively less attention has…

Computation and Language · Computer Science 2025-08-21 Aditya K Surikuchi , Raquel Fernández , Sandro Pezzelle

Event cameras are bio-inspired sensors that capture the per-pixel intensity changes asynchronously and produce event streams encoding the time, pixel position, and polarity (sign) of the intensity changes. Event cameras possess a myriad of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Xu Zheng , Yexin Liu , Yunfan Lu , Tongyan Hua , Tianbo Pan , Weiming Zhang , Dacheng Tao , Lin Wang

Among prerequisites for a synthetic agent to interact with dynamic scenes, the ability to identify independently moving objects is specifically important. From an application perspective, nevertheless, standard cameras may deteriorate…

Computer Vision and Pattern Recognition · Computer Science 2021-11-08 Xiuyuan Lu , Yi Zhou , Shaojie Shen

Visual grounding seeks to localize the image region corresponding to a free-form text description. Recently, the strong multimodal capabilities of Large Vision-Language Models (LVLMs) have driven substantial improvements in visual…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Seil Kang , Jinyeong Kim , Junhyeok Kim , Seong Jae Hwang

Audio LLMs have shown a strong ability to understand audio samples, yet their reliability in complex acoustic scenes remains under-explored. Unlike prior work limited to small scale or less controlled query construction, we present a…

Sound · Computer Science 2026-03-05 Taehan Lee , Jaehan Jung , Hyukjun Lee

We introduce the task of 3D visual grounding in large-scale dynamic scenes based on natural linguistic descriptions and online captured multi-modal visual data, including 2D images and 3D LiDAR point clouds. We present a novel method,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Zhenxiang Lin , Xidong Peng , Peishan Cong , Ge Zheng , Yujin Sun , Yuenan Hou , Xinge Zhu , Sibei Yang , Yuexin Ma

Text-to-motion generation has advanced with diffusion models, yet existing systems often collapse complex multi-action prompts into a single embedding, leading to omissions, reordering, or unnatural transitions. In this work, we shift…

Graphics · Computer Science 2026-02-05 Seong-Eun Hong , JaeYoung Seon , JuYeong Hwang , JongHwan Shin , HyeongYeop Kang

We then introduce a novel hierarchical knowledge distillation strategy that incorporates the similarity matrix, feature representation, and response map-based distillation to guide the learning of the student Transformer network. We also…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Shiao Wang , Xiao Wang , Chao Wang , Liye Jin , Lin Zhu , Bo Jiang , Yonghong Tian , Jin Tang

Large-scale video generation models have demonstrated high visual realism in diverse contexts, spurring interest in their potential as general-purpose world simulators. Existing benchmarks focus on individual subjects rather than scenes…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Aaron Appelle , Jerome P. Lynch

We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language. We ask annotators to describe an image with their voice while simultaneously hovering their mouse over the region they are…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Jordi Pont-Tuset , Jasper Uijlings , Soravit Changpinyo , Radu Soricut , Vittorio Ferrari

Event cameras are vision sensors that record asynchronous streams of per-pixel brightness changes, referred to as "events". They have appealing advantages over frame-based cameras for computer vision, including high temporal resolution,…

Computer Vision and Pattern Recognition · Computer Science 2019-08-21 Daniel Gehrig , Antonio Loquercio , Konstantinos G. Derpanis , Davide Scaramuzza

Event cameras provide a natural and data efficient representation of visual information, motivating novel computational strategies towards extracting visual information. Inspired by the biological vision system, we propose a behavior driven…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Nan Cai , Pia Bideau

Neuromorphic vision or event vision is an advanced vision technology, where in contrast to the visible camera that outputs pixels, the event vision generates neuromorphic events every time there is a brightness change which exceeds a…

Computer Vision and Pattern Recognition · Computer Science 2023-01-11 Waseem Shariff , Muhammad Ali Farooq , Joe Lemley , Peter Corcoran
‹ Prev 1 3 4 5 6 7 10 Next ›