中文
相关论文

相关论文: ACT360: An Efficient 360-Degree Action Detection a…

200 篇论文

Effective urban warfare training requires situational awareness and muscle memory, developed through repeated practice in realistic yet controlled environments. A key drill, Enter and Clear the Room (ECR), demands threat assessment,…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Surya Rayala , Marcos Quinones-Grueiro , Naveeduddin Mohammed , Ashwin T S , Benjamin Goldberg , Randall Spain , Paige Lawton , Gautam Biswas

Accelerators implementing Deep Neural Networks for image-based object detection operate on large volumes of data due to fetching images and neural network parameters, especially if they need to process video streams, hence with high power…

硬件体系结构 · 计算机科学 2023-03-01 Martí Caro , Hamid Tabani , Jaume Abella

Self-attention based Transformer models have demonstrated impressive results for image classification and object detection, and more recently for video understanding. Inspired by this success, we investigate the application of Transformer…

计算机视觉与模式识别 · 计算机科学 2022-08-30 Chenlin Zhang , Jianxin Wu , Yin Li

Research into multi-modal perception, human cognition, behavior, and attention can benefit from high-fidelity content that may recreate real-life-like scenes when rendered on head-mounted displays. Moreover, aspects of audiovisual…

图像与视频处理 · 电气工程与系统科学 2022-12-29 Thomas Robotham , Ashutosh Singla , Olli S. Rummukainen , Alexander Raake , Emanuël A. P. Habets

As it requires a huge number of parameters when exposed to high dimensional inputs in video detection and classification, there is a grand challenge to develop a compact yet accurate video comprehension at terminal devices. Current works…

计算机视觉与模式识别 · 计算机科学 2018-06-08 Yuan Cheng , Guangya Li , Hai-Bao Chen , Sheldon X. -D. Tan , Hao Yu

Obstacle avoidance in unmanned aerial vehicles (UAVs), as a fundamental capability, has gained increasing attention with the growing focus on spatial intelligence. However, current obstacle-avoidance methods mainly depend on limited…

机器人学 · 计算机科学 2026-03-09 Xiangkai Zhang , Dizhe Zhang , WenZhuo Cao , Zhaoliang Wan , Yingjie Niu , Lu Qi , Xu Yang , Zhiyong Liu

High precision, lightweight, and real-time responsiveness are three essential requirements for implementing autonomous driving. In this study, we incorporate A-YOLOM, an adaptive, real-time, and lightweight multi-task model designed to…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Jiayuan Wang , Q. M. Jonathan Wu , Ning Zhang

Evaluating whether human action is standard or not and providing reasonable feedback to improve action standardization is very crucial but challenging in real-world scenarios. However, current video understanding methods are mainly…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Mengshi Qi , Yeteng Wu , Xianlin Zhang , Huadong Ma

Autonomous Vehicle (AV) perception systems require more than simply seeing, via e.g., object detection or scene segmentation. They need a holistic understanding of what is happening within the scene for safe interaction with other road…

Detecting anomalies in human-related videos is crucial for surveillance applications. Current methods primarily include appearance-based and action-based techniques. Appearance-based methods rely on low-level visual features such as color,…

计算机视觉与模式识别 · 计算机科学 2024-09-06 Chenglizhao Chen , Xinyu Liu , Mengke Song , Luming Li , Xu Yu , Shanchen Pang

We propose ASL360, an adaptive deep reinforcement learning-based scheduler for on-demand 360$^\circ$ video streaming to mobile VR users in next generation wireless networks. We aim to maximize the overall Quality of Experience (QoE) of the…

网络与互联网体系结构 · 计算机科学 2026-02-24 Alireza Mohammadhosseini , Jacob Chakareski , Nicholas Mastronarde

The YOLO (You Only Look Once) series has been a leading framework in real-time object detection, consistently improving the balance between speed and accuracy. However, integrating attention mechanisms into YOLO has been challenging due to…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Rahima Khanam , Muhammad Hussain

Delving into the realm of egocentric vision, the advancement of referring video object segmentation (RVOS) stands as pivotal in understanding human activities. However, existing RVOS task primarily relies on static attributes such as object…

计算机视觉与模式识别 · 计算机科学 2024-07-11 Liangyang Ouyang , Ruicong Liu , Yifei Huang , Ryosuke Furuta , Yoichi Sato

The canonical approach to video action recognition dictates a neural model to do a classic and standard 1-of-N majority vote task. They are trained to predict a fixed set of predefined categories, limiting their transferable ability on new…

计算机视觉与模式识别 · 计算机科学 2021-09-20 Mengmeng Wang , Jiazheng Xing , Yong Liu

Traditional depth sensors generate accurate real world depth estimates that surpass even the most advanced learning approaches trained only on simulation domains. Since ground truth depth is readily available in the simulation domain but…

计算机视觉与模式识别 · 计算机科学 2021-12-07 Isabella Liu , Edward Yang , Jianyu Tao , Rui Chen , Xiaoshuai Zhang , Qing Ran , Zhu Liu , Hao Su

The development of effective training and evaluation strategies is critical. Conventional methods for assessing surgical proficiency typically rely on expert supervision, either through onsite observation or retrospective analysis of…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Yan Meng , Daniel A. Donoho , Marcelle Altshuler , Omar Arnaout

Pretrained video generation models provide strong priors for robot control, but existing unified world action models still struggle to decode reliable actions without substantial robot-specific training. We attribute this limitation to a…

机器人学 · 计算机科学 2026-04-14 Liaoyuan Fan , Zetian Xu , Chen Cao , Wenyao Zhang , Mingqi Yuan , Jiayu Chen

Safety helmets play a crucial role in protecting workers from head injuries in construction sites, where potential hazards are prevalent. However, currently, there is no approach that can simultaneously achieve both model accuracy and…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Shuqi Shen , Junjie Yang

We present See360, which is a versatile and efficient framework for 360 panoramic view interpolation using latent space viewpoint estimation. Most of the existing view rendering approaches only focus on indoor or synthetic 3D environments…

计算机视觉与模式识别 · 计算机科学 2024-01-09 Zhi-Song Liu , Marie-Paule Cani , Wan-Chi Siu

Robotic imitation learning has advanced from solving static tasks to addressing dynamic interaction scenarios, but testing and evaluation remain costly and challenging due to the need for real-time interaction with dynamic environments. We…