中文
相关论文

相关论文: ACT360: An Efficient 360-Degree Action Detection a…

200 篇论文

In recent years, significant progress has been made in the field of underwater image enhancement (UIE). However, its practical utility for high-level vision tasks, such as underwater object detection (UOD) in Autonomous Underwater Vehicles…

计算机视觉与模式识别 · 计算机科学 2024-03-29 Junjie Wen , Jinqiang Cui , Benyun Zhao , Bingxin Han , Xuchen Liu , Zhi Gao , Ben M. Chen

This study examines how Critical Care Air Transport Team (CCATT) members are trained using mixed-reality simulations that replicate the high-pressure conditions of aeromedical evacuation. Each team - a physician, nurse, and respiratory…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Divya Mereddy , Marcos Quinones-Grueiro , Ashwin T S , Eduardo Davalos , Gautam Biswas , Kent Etherton , Tyler Davis , Katelyn Kay , Jill Lear , Benjamin Goldberg

Temporal action localization aims to predict the boundary and category of each action instance in untrimmed long videos. Most of previous methods based on anchors or proposals neglect the global-local context interaction in entire video…

计算机视觉与模式识别 · 计算机科学 2022-09-16 Yizheng Ouyang , Tianjin Zhang , Weibo Gu , Hongfa Wang

Advances in neural fields are enabling high-fidelity capture of the shape and appearance of dynamic 3D scenes. However, their capabilities lag behind those offered by conventional representations such as 2D videos because of algorithmic…

We introduce OpenVO, a novel framework for Open-world Visual Odometry (VO) with temporal awareness under limited input conditions. OpenVO effectively estimates real-world-scale ego-motion from monocular dashcam footage with varying…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Phuc D. A. Nguyen , Anh N. Nhu , Ming C. Lin

Despite significant progress in the development of human action detection datasets and algorithms, no current dataset is representative of real-world aerial view scenarios. We present Okutama-Action, a new video dataset for aerial view…

计算机视觉与模式识别 · 计算机科学 2017-06-16 Mohammadamin Barekatain , Miquel Martí , Hsueh-Fu Shih , Samuel Murray , Kotaro Nakayama , Yutaka Matsuo , Helmut Prendinger

Optical flow estimation in omnidirectional videos faces two significant issues: the lack of benchmark datasets and the challenge of adapting perspective video-based methods to accommodate the omnidirectional nature. This paper proposes the…

计算机视觉与模式识别 · 计算机科学 2022-08-09 Keshav Bhandari , Bin Duan , Gaowen Liu , Hugo Latapie , Ziliang Zong , Yan Yan

Egocentric temporal action segmentation in videos is a crucial task in computer vision with applications in various fields such as mixed reality, human behavior analysis, and robotics. Although recent research has utilized advanced…

计算机视觉与模式识别 · 计算机科学 2023-05-25 Sakib Reza , Balaji Sundareshan , Mohsen Moghaddam , Octavia Camps

World models have emerged as promising neural simulators for autonomous driving, with the potential to supplement scarce real-world data and enable closed-loop evaluations. However, current research primarily evaluates these models based on…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Hidehisa Arai , Keishi Ishihara , Tsubasa Takahashi , Yu Yamaguchi

Repetitive action counting quantifies the frequency of specific actions performed by individuals. However, existing action-counting datasets have limited action diversity, potentially hampering model performance on unseen actions. To…

计算机视觉与模式识别 · 计算机科学 2024-06-19 Jiada Lu , WeiWei Zhou , Xiang Qian , Dongze Lian , Yanyu Xu , Weifeng Wang , Lina Cao , Shenghua Gao

Recent advances in diffusion transformers have empowered video generation models to generate high-quality video clips from texts or images. However, world models with the ability to predict long-horizon futures from past observations and…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Yixuan Zhu , Jiaqi Feng , Wenzhao Zheng , Yuan Gao , Xin Tao , Pengfei Wan , Jie Zhou , Jiwen Lu

Two factors have proven to be very important to the performance of semantic segmentation models: global context and multi-level semantics. However, generating features that capture both factors always leads to high computational complexity,…

计算机视觉与模式识别 · 计算机科学 2021-03-11 Qi Song , Kangfu Mei , Rui Huang

Recognizing human actions in videos requires spatial and temporal understanding. Most existing action recognition models lack a balanced spatio-temporal understanding of videos. In this work, we propose a novel two-stream architecture,…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Dongho Lee , Jongseo Lee , Jinwoo Choi

Classifying fine-grained actions in fast-paced, close-combat sports such as fencing and boxing presents unique challenges due to the complexity, speed, and nuance of movements. Traditional methods reliant on pose estimation or fancy sensor…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Christopher Lai , Jason Mo , Haotian Xia , Yuan-fang Wang

Omnidirectional camera is a cost-effective and information-rich sensor highly suitable for many marine applications and the ocean scientific community, encompassing several domains such as augmented reality, mapping, motion estimation,…

机器人学 · 计算机科学 2023-10-03 Quan-Dung Pham , Yipeng Zhu , Tan-Sang Ha , K. H. Long Nguyen , Binh-Son Hua , Sai-Kit Yeung

Online Action Detection (OAD) detects actions in streaming videos using past observations. State-of-the-art OAD approaches model past observations and their interactions with an anticipated future. The past is encoded using short- and…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Zhanzhong Pang , Fadime Sener , Angela Yao

Human-object interaction segmentation is a fundamental task of daily activity understanding, which plays a crucial role in applications such as assistive robotics, healthcare, and autonomous systems. Most existing learning-based methods…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Hao Xing , Kai Zhe Boey , Gordon Cheng

The increasing use of 360 images across various domains has emphasized the need for robust depth estimation techniques tailored for omnidirectional images. However, obtaining large-scale labeled datasets for 360 depth estimation remains a…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Dongki Jung , Jaehoon Choi , Yonghan Lee , Dinesh Manocha

Despite recent progress in video large language models (VideoLLMs), a key open challenge remains: how to equip models with chain-of-thought (CoT) reasoning abilities grounded in fine-grained object-level video understanding. Existing…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Yanan Wang , Julio Vizcarra , Zhi Li , Hao Niu , Mori Kurokawa

The field of 360-degree omnidirectional understanding has been receiving increasing attention for advancing spatial intelligence. However, the lack of large-scale and diverse data remains a major limitation. In this work, we propose…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Xian Ge , Yuling Pan , Yuhang Zhang , Xiang Li , Weijun Zhang , Dizhe Zhang , Zhaoliang Wan , Xin Lin , Xiangkai Zhang , Juntao Liang , Jason Li , Wenjie Jiang , Bo Du , Ming-Hsuan Yang , Lu Qi
‹ 上一页 1 8 9 10 下一页 ›