中文
相关论文

相关论文: MVP: Robust Multi-View Practice for Driving Action…

200 篇论文

We present Multi-View Attentive Contextualization (MvACon), a simple yet effective method for improving 2D-to-3D feature lifting in query-based multi-view 3D (MV3D) object detection. Despite remarkable progress witnessed in the field of…

计算机视觉与模式识别 · 计算机科学 2024-05-21 Xianpeng Liu , Ce Zheng , Ming Qian , Nan Xue , Chen Chen , Zhebin Zhang , Chen Li , Tianfu Wu

Action recognition in videos is a challenging task due to the complexity of the spatio-temporal patterns to model and the difficulty to acquire and learn on large quantities of video data. Deep learning, although a breakthrough for image…

计算机视觉与模式识别 · 计算机科学 2016-08-26 César Roberto de Souza , Adrien Gaidon , Eleonora Vig , Antonio Manuel López

The creation of diverse and realistic driving scenarios has become essential to enhance perception and planning capabilities of the autonomous driving system. However, generating long-duration, surround-view consistent driving videos…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Rui Chen , Zehuan Wu , Yichen Liu , Yuxin Guo , Jingcheng Ni , Haifeng Xia , Siyu Xia

Despite increasing interest in computer vision-based distracted driving detection, most existing models rely exclusively on driver-facing views and overlook crucial environmental context that influences driving behavior. This study…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Anthony Dontoh , Stephanie Ivey , Armstrong Aboah

Simultaneous object recognition and pose estimation are two key functionalities for robots to safely interact with humans as well as environments. Although both object recognition and pose estimation use visual input, most state-of-the-art…

机器人学 · 计算机科学 2023-04-10 Tommaso Parisotto , Subhaditya Mukherjee , Hamidreza Kasaei

The data sparsity problem significantly hinders the performance of recommender systems, as traditional models rely on limited historical interactions to learn user preferences and item properties. While incorporating multimodal information…

信息检索 · 计算机科学 2025-05-23 Jinfeng Xu , Zheyu Chen , Jinze Li , Shuo Yang , Hewei Wang , Yijie Li , Mengran Li , Puzhen Wu , Edith C. H. Ngai

Robots can use Visual Imitation Learning (VIL) to learn manipulation tasks from video demonstrations. However, translating visual observations into actionable robot policies is challenging due to the high-dimensional nature of video data.…

机器人学 · 计算机科学 2025-01-22 Ananth Jonnavittula , Sagar Parekh , Dylan P. Losey

Inspired by human vision, we propose a new periphery-fovea multi-resolution driving model that predicts vehicle speed from dash camera videos. The peripheral vision module of the model processes the full video frames in low resolution. Its…

计算机视觉与模式识别 · 计算机科学 2019-03-26 Ye Xia , Jinkyu Kim , John Canny , Karl Zipser , David Whitney

Localization is a key requirement for mobile robot autonomy and human-robot interaction. Vision-based localization is accurate and flexible, however, it incurs a high computational burden which limits its application on many…

机器人学 · 计算机科学 2016-12-30 Ronald Clark , Sen Wang , Hongkai Wen , Niki Trigoni , Andrew Markham

Due to its relevance in intelligent transportation systems, anomaly detection in traffic videos has recently received much interest. It remains a difficult problem due to a variety of factors influencing the video quality of a real-time…

计算机视觉与模式识别 · 计算机科学 2021-04-21 Keval Doshi , Yasin Yilmaz

In this paper, we present a solution to Large-Scale Video Classification Challenge (LSVC2017) [1] that ranked the 1st place. We focused on a variety of modalities that cover visual, motion and audio. Also, we visualized the aggregation…

计算机视觉与模式识别 · 计算机科学 2017-10-31 Chen Chen , Xiaowei Zhao , Yang Liu

Multi-Object Tracking (MOT) has been a long-standing challenge in video understanding. A natural and intuitive approach is to split this task into two parts: object detection and association. Most mainstream methods employ meticulously…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Ruopeng Gao , Ji Qi , Limin Wang

The objective of this paper is self-supervised representation learning, with the goal of solving semi-supervised video object segmentation (a.k.a. dense tracking). We make the following contributions: (i) we propose to improve the existing…

计算机视觉与模式识别 · 计算机科学 2020-06-23 Fangrui Zhu , Li Zhang , Yanwei Fu , Guodong Guo , Weidi Xie

We introduce a novel approach to dynamic obstacle avoidance based on Deep Reinforcement Learning by defining a traffic type independent environment with variable complexity. Filling a gap in the current literature, we thoroughly investigate…

机器学习 · 计算机科学 2021-12-30 Fabian Hart , Martin Waltz , Ostap Okhrin

There has been a paradigm-shift in urban logistic services in the last years; demand for real-time, instant mobility and delivery services grows. This poses new challenges to logistic service providers as the underlying stochastic dynamic…

人工智能 · 计算机科学 2021-03-02 Florentin D Hildebrandt , Barrett Thomas , Marlin W Ulmer

Foot contact is an important cue for human motion capture, understanding, and generation. Existing datasets tend to annotate dense foot contact using visual matching with thresholding or incorporating pressure signals. However, these…

计算机视觉与模式识别 · 计算机科学 2024-04-02 He Zhang , Shenghao Ren , Haolei Yuan , Jianhui Zhao , Fan Li , Shuangpeng Sun , Zhenghao Liang , Tao Yu , Qiu Shen , Xun Cao

To track the target in a video, current visual trackers usually adopt greedy search for target object localization in each frame, that is, the candidate region with the maximum response score will be selected as the tracking result of each…

计算机视觉与模式识别 · 计算机科学 2022-10-26 Xiao Wang , Zhe Chen , Bo Jiang , Jin Tang , Bin Luo , Dacheng Tao

Visual Place Recognition (VPR) is a crucial part of mobile robotics and autonomous driving as well as other computer vision tasks. It refers to the process of identifying a place depicted in a query image using only computer vision. At…

计算机视觉与模式识别 · 计算机科学 2023-03-07 Amar Ali-bey , Brahim Chaib-draa , Philippe Giguère

Running a large open-vocabulary (Open-vocab) detector on every video frame is accurate but expensive. We introduce a training-free pipeline that invokes OWLv2 only on fixed-interval keyframes and propagates detections to intermediate frames…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Binhua Huang , Ni Wang , Wendong Yao , Soumyabrata Dev

This study presents a novel driver drowsiness detection system that combines deep learning techniques with the OpenCV framework. The system utilises facial landmarks extracted from the driver's face as input to Convolutional Neural Networks…

计算机视觉与模式识别 · 计算机科学 2024-06-25 Sandeep Singh Sengar , Aswin Kumar , Owen Singh