中文
相关论文

相关论文: Human in Events: A Large-Scale Benchmark for Human…

200 篇论文

Along with the increasing use of unmanned aerial vehicles (UAVs), large volumes of aerial videos have been produced. It is unrealistic for humans to screen such big data and understand their contents. Hence methodological research on the…

计算机视觉与模式识别 · 计算机科学 2020-06-26 Lichao Mou , Yuansheng Hua , Pu Jin , Xiao Xiang Zhu

Estimating 3D human motion from an egocentric video sequence plays a critical role in human behavior understanding and has various applications in VR/AR. However, naively learning a mapping between egocentric videos and human motions is…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Jiaman Li , C. Karen Liu , Jiajun Wu

First-person video highlights a camera-wearer's activities in the context of their persistent environment. However, current video understanding approaches reason over visual features from short video clips that are detached from the…

计算机视觉与模式识别 · 计算机科学 2023-11-13 Tushar Nagarajan , Santhosh Kumar Ramakrishnan , Ruta Desai , James Hillis , Kristen Grauman

Predicting pedestrian movement is critical for human behavior analysis and also for safe and efficient human-agent interactions. However, despite significant advancements, it is still challenging for existing approaches to capture the…

计算机视觉与模式识别 · 计算机科学 2022-11-01 Pei Xu , Jean-Bernard Hayet , Ioannis Karamouzas

In recent years, technologies of indoor crowd positioning and movement data analysis have received widespread attention in the fields of reliability management, indoor navigation, and crowd behavior monitoring. However, only a few indoor…

其他计算机科学 · 计算机科学 2021-09-03 Ying Zhao , Xin Zhao , Siming Chen , Zhuo Zhang , Xin Huang

Tracking humans in crowded video sequences is an important constituent of visual scene understanding. Increasing crowd density challenges visibility of humans, limiting the scalability of existing pedestrian trackers to higher crowd…

计算机视觉与模式识别 · 计算机科学 2021-03-26 Ramana Sundararaman , Cedric De Almeida Braga , Eric Marchand , Julien Pettre

Effectively measuring the similarity between two human motions is necessary for several computer vision tasks such as gait analysis, person identi- fication and action retrieval. Nevertheless, we believe that traditional approaches such as…

计算机视觉与模式识别 · 计算机科学 2018-08-07 Huseyin Coskun , David Joseph Tan , Sailesh Conjeti , Nassir Navab , Federico Tombari

We present a review on the current state of publicly available datasets within the human action recognition community; highlighting the revival of pose based methods and recent progress of understanding person-person interaction modeling.…

计算机视觉与模式识别 · 计算机科学 2015-11-20 Michael Edwards , Jingjing Deng , Xianghua Xie

People detection methods are highly sensitive to the perpetual occlusions among the targets. As multi-camera set-ups become more frequently encountered, joint exploitation of the across views information would allow for improved detection…

计算机视觉与模式识别 · 计算机科学 2017-07-31 Tatjana Chavdarova , Pierre Baqué , Stéphane Bouquet , Andrii Maksai , Cijo Jose , Louis Lettry , Pascal Fua , Luc Van Gool , François Fleuret

Visual Emotion Analysis (VEA) aims at predicting people's emotional responses to visual stimuli. This is a promising, yet challenging, task in affective computing, which has drawn increasing attention in recent years. Most of the existing…

计算机视觉与模式识别 · 计算机科学 2023-07-31 Jingyuan Yang , Qirui Huang , Tingting Ding , Dani Lischinski , Daniel Cohen-Or , Hui Huang

Despite rapid progress in large vision-language models (LVLMs), existing video caption benchmarks remain limited in evaluating their alignment with human understanding. Most rely on a single annotation per video and lexical similarity-based…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Shiyu Hu , Xuchen Li , Xuzhao Li , Jing Zhang , Yipei Wang , Xin Zhao , Kang Hao Cheong

Video-based visual relation detection tasks, such as video scene graph generation, play important roles in fine-grained video understanding. However, current video visual relation detection datasets have two main limitations that hinder the…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Tao Wu , Runyu He , Gangshan Wu , Limin Wang

In this paper, we propose H-MoRe, a novel pipeline for learning precise human-centric motion representation. Our approach dynamically preserves relevant human motion while filtering out background movement. Notably, unlike previous methods…

计算机视觉与模式识别 · 计算机科学 2025-04-16 Zhanbo Huang , Xiaoming Liu , Yu Kong

Events describe happenings in our world that are of importance. Naturally, understanding events mentioned in multimedia content and how they are related forms an important way of comprehending our world. Existing literature can infer if…

计算机视觉与模式识别 · 计算机科学 2023-12-21 Hammad A. Ayyubi , Christopher Thomas , Lovish Chum , Rahul Lokesh , Long Chen , Yulei Niu , Xudong Lin , Xuande Feng , Jaywon Koo , Sounak Ray , Shih-Fu Chang

For a complete comprehension of multi-person scenes, it is essential to go beyond basic tasks like detection and tracking. Higher-level tasks, such as understanding the interactions and social activities among individuals, are also crucial.…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Mahsa Ehsanpour , Ian Reid , Hamid Rezatofighi

Egocentric human pose estimation (HPE) using a head-mounted device is crucial for various VR and AR applications, but it faces significant challenges due to keypoint invisibility. Nevertheless, none of the existing egocentric HPE datasets…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Peng Dai , Yu Zhang , Yiqiang Feng , Zhen Fan , Yang Zhang

In human imitation learning, the imitator typically take the egocentric view as a benchmark, naturally transferring behaviors observed from an exocentric view to their owns, which provides inspiration for researching how robots can more…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Heqian Qiu , Zhaofeng Shi , Lanxiao Wang , Huiyu Xiong , Xiang Li , Hongliang Li

We introduce Human-like Video Models (HVM-1), large-scale video models pretrained with nearly 5000 hours of curated human-like video data (mostly egocentric, temporally extended, continuous video recordings), using the spatiotemporal masked…

计算机视觉与模式识别 · 计算机科学 2024-07-26 A. Emin Orhan

We present an approach to labeling short video clips with English verbs as event descriptions. A key distinguishing aspect of this work is that it labels videos with verbs that describe the spatiotemporal interaction between event…

We present a new deep learning approach for real-time 3D human action recognition from skeletal data and apply it to develop a vision-based intelligent surveillance system. Given a skeleton sequence, we propose to encode skeleton poses and…

计算机视觉与模式识别 · 计算机科学 2022-08-11 Huy Hieu Pham , Houssam Salmane , Louahdi Khoudour , Alain Crouzil , Pablo Zegers , Sergio A Velastin