English
Related papers

Related papers: ReXCam: Resource-Efficient, Cross-Camera Video Ana…

200 papers

Training a referring expression comprehension (ReC) model for a new visual domain requires collecting referring expressions, and potentially corresponding bounding boxes, for images in the domain. While large-scale pre-trained models are…

Computer Vision and Pattern Recognition · Computer Science 2022-05-04 Sanjay Subramanian , William Merrill , Trevor Darrell , Matt Gardner , Sameer Singh , Anna Rohrbach

Multi-target multi-camera tracking (MTMCT), i.e., tracking multiple targets across multiple cameras, is a crucial technique for smart city applications. In this paper, we propose an effective and reliable MTMCT framework for vehicles, which…

Computer Vision and Pattern Recognition · Computer Science 2020-09-01 Hung-Min Hsu , Yizhou Wang , Jenq-Neng Hwang

Typical attempts to improve the capability of visual place recognition techniques include the use of multi-sensor fusion and integration of information over time from image sequences. These approaches can improve performance but have…

Robotics · Computer Science 2019-03-11 Stephen Hausler , Adam Jacobson , Michael Milford

A truly interactive world model requires three key ingredients: real-time long-horizon streaming, consistent spatial memory, and precise user control. However, most existing approaches address only one of these aspects in isolation, as…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Yicong Hong , Yiqun Mei , Chongjian Ge , Yiran Xu , Yang Zhou , Sai Bi , Yannick Hold-Geoffroy , Mike Roberts , Matthew Fisher , Eli Shechtman , Kalyan Sunkavalli , Feng Liu , Zhengqi Li , Hao Tan

The growing need for video surveillance in public spaces has created a demand for systems that can track individuals across multiple cameras feeds in real-time. While existing tracking systems have achieved impressive performance using deep…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Vipin Gautam , Shitala Prasad , Sharad Sinha

We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data. Our model leverages over 15 billion curated image-text pairs and 10 billion video-hashtag…

Vision Transformers achieve impressive accuracy across a range of visual recognition tasks. Unfortunately, their accuracy frequently comes with high computational costs. This is a particular issue in video recognition, where models are…

Computer Vision and Pattern Recognition · Computer Science 2023-08-28 Matthew Dutson , Yin Li , Mohit Gupta

Simultaneous Localization and Mapping (SLAM) plays an important role in many robotics fields, including social robots. Many of the available visual SLAM methods are based on the assumption of a static world and struggle in dynamic…

Robotics · Computer Science 2025-10-06 Mobin Habibpour , Alireza Nemati , Ali Meghdari , Alireza Taheri , Shima Nazari

The quality of the video stream is key to neural network-based video analytics. However, low-quality video is inevitably collected by existing surveillance systems because of poor quality cameras or over-compressed/pruned video streaming…

Computer Vision and Pattern Recognition · Computer Science 2023-01-25 Tingting Yuan , Liang Mi , Weijun Wang , Haipeng Dai , Xiaoming Fu

Recent advancements in video semantic segmentation have made substantial progress by exploiting temporal correlations. Nevertheless, persistent challenges, including redundant computation and the reliability of the feature propagation…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Yaoyan Zheng , Hongyu Yang , Di Huang

Retrieving images from the same location as a given query is an important component of multiple computer vision tasks, like Visual Place Recognition, Landmark Retrieval, Visual Localization, 3D reconstruction, and SLAM. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Gabriele Berton , Carlo Masone

We present DeepCache, a principled cache design for deep learning inference in continuous mobile vision. DeepCache benefits model execution efficiency by exploiting temporal locality in input video streams. It addresses a key challenge…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Mengwei Xu , Mengze Zhu , Yunxin Liu , Felix Xiaozhu Lin , Xuanzhe Liu

Real-time world simulation is becoming a key infrastructure for scalable evaluation and online reinforcement learning of autonomous driving systems. Recent driving world models built on autoregressive video diffusion achieve high-fidelity,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Yixiao Zeng , Jianlei Zheng , Chaoda Zheng , Shijia Chen , Mingdian Liu , Tongping Liu , Tengwei Luo , Yu Zhang , Boyang Wang , Linkun Xu , Siyuan Lu , Bo Tian , Xianming Liu

Hundreds of millions of network cameras have been installed throughout the world. Each is capable of providing a vast amount of real-time data. Analyzing the massive data generated by these cameras requires significant computational…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-01-21 Zohar Kapach , Andrew Ulmer , Daniel Merrick , Arshad Alikhan , Yung-Hsiang Lu , Anup Mohan , Ahmed S. Kaseb , George K. Thiruvathukal

Video-based person re-identification (re-id) is a central application in surveillance systems with significant concern in security. Matching persons across disjoint camera views in their video fragments is inherently challenging due to the…

Computer Vision and Pattern Recognition · Computer Science 2018-10-17 Lin Wu , Yang Wang , Junbin Gao , Xue Li

We developed REVEX, a removal-based video explanations framework. This work extends fine-grained explanation frameworks for computer vision data and adapts six existing techniques to video by adding temporal information and local…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 F. Xavier Gaya-Morey , Jose M. Buades-Rubio , I. Scott MacKenzie , Cristina Manresa-Yee

We present Recurrent Vision Transformers (RVTs), a novel backbone for object detection with event cameras. Event cameras provide visual information with sub-millisecond latency at a high-dynamic range and with strong robustness against…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Mathias Gehrig , Davide Scaramuzza

Visual place recognition is a critical task in computer vision, especially for localization and navigation systems. Existing methods often rely on contrastive learning: image descriptors are trained to have small distance for similar images…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 María Leyva-Vallina , Nicola Strisciuglio , Nicolai Petkov

In multiview applications, multiple cameras acquire the same scene from different viewpoints and generally produce correlated video streams. This results in large amounts of highly redundant data. In order to save resources, it is critical…

Multimedia · Computer Science 2015-06-15 Laura Toni , Thomas Maugey , Pascal Frossard

Existing multimodal-based human action recognition approaches are computationally intensive, limiting their deployment in real-time applications. In this work, we present a novel and efficient pose-driven attention-guided multimodal network…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Ahmed Abdelkawy , Asem Ali , Aly Farag