English
Related papers

Related papers: Human-centric Behavior Description in Videos: New …

200 papers

In this paper we address the problem of motion event detection in athlete recordings from individual sports. In contrast to recent end-to-end approaches, we propose to use 2D human pose sequences as an intermediate representation that…

Computer Vision and Pattern Recognition · Computer Science 2020-04-23 Moritz Einfalt , Rainer Lienhart

When we say a person is texting, can you tell the person is walking or sitting? Emphatically, no. In order to solve this incomplete representation problem, this paper presents a sub-action descriptor for detailed action detection. The…

Computer Vision and Pattern Recognition · Computer Science 2017-10-11 Cheng-Bin Jin , Shengzhe Li , Hakil Kim

This paper presents our solution to ACM MM challenge: Large-scale Human-centric Video Analysis in Complex Events\cite{lin2020human}; specifically, here we focus on Track3: Crowd Pose Tracking in Complex Events. Remarkable progress has been…

Computer Vision and Pattern Recognition · Computer Science 2020-10-22 Li Yuan , Shuning Chang , Ziyuan Huang , Yichen Zhou , Yunpeng Chen , Xuecheng Nie , Francis E. H. Tay , Jiashi Feng , Shuicheng Yan

Every moment counts in action recognition. A comprehensive understanding of human activity in video requires labeling every frame according to the actions occurring, placing multiple labels densely over a video sequence. To study this…

Computer Vision and Pattern Recognition · Computer Science 2017-06-12 Serena Yeung , Olga Russakovsky , Ning Jin , Mykhaylo Andriluka , Greg Mori , Li Fei-Fei

Marine videos present significant challenges for video understanding due to the dynamics of marine objects and the surrounding environment, camera motion, and the complexity of underwater scenes. Existing video captioning datasets,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Quang-Trung Truong , Yuk-Kwan Wong , Vo Hoang Kim Tuyen Dang , Rinaldi Gotama , Duc Thanh Nguyen , Sai-Kit Yeung

We present a novel LLM-based pipeline for creating contextual descriptions of human body poses in images using only auxiliary attributes. This approach facilitates the creation of the MPII Pose Descriptions dataset, which includes natural…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Muhammad Saif Ullah Khan , Muhammad Ferjad Naeem , Federico Tombari , Luc Van Gool , Didier Stricker , Muhammad Zeshan Afzal

We consider the problem of estimating frame-level full human body meshes given a video of a person with natural motion dynamics. While much progress in this field has been in single image-based mesh estimation, there has been a recent…

Computer Vision and Pattern Recognition · Computer Science 2021-07-28 Runze Li , Srikrishna Karanam , Ren Li , Terrence Chen , Bir Bhanu , Ziyan Wu

Video captioning is the task of automatically generating a textual description of the actions in a video. Although previous work (e.g. sequence-to-sequence model) has shown promising results in abstracting a coarse description of a short…

Computer Vision and Pattern Recognition · Computer Science 2018-03-30 Xin Wang , Wenhu Chen , Jiawei Wu , Yuan-Fang Wang , William Yang Wang

Computer vision has a great potential to help our daily lives by searching for lost keys, watering flowers or reminding us to take a pill. To succeed with such tasks, computer vision methods need to be trained from real and diverse examples…

Computer Vision and Pattern Recognition · Computer Science 2016-07-28 Gunnar A. Sigurdsson , Gül Varol , Xiaolong Wang , Ali Farhadi , Ivan Laptev , Abhinav Gupta

We present a novel large-scale dataset and comprehensive baselines for end-to-end pedestrian detection and person recognition in raw video frames. Our baselines address three issues: the performance of various combinations of detectors and…

Computer Vision and Pattern Recognition · Computer Science 2017-04-07 Liang Zheng , Hengheng Zhang , Shaoyan Sun , Manmohan Chandraker , Yi Yang , Qi Tian

Tracking a target person from robot-egocentric views is crucial for developing autonomous robots that provide continuous personalized assistance or collaboration in Human-Robot Interaction (HRI) and Embodied AI. However, most existing…

Robotics · Computer Science 2025-07-10 Hanjing Ye , Yu Zhan , Weixi Situ , Guangcheng Chen , Jingwen Yu , Ziqi Zhao , Kuanqi Cai , Arash Ajoudani , Hong Zhang

Deep learning for human action recognition in videos is making significant progress, but is slowed down by its dependency on expensive manual labeling of large video collections. In this work, we investigate the generation of synthetic…

Computer Vision and Pattern Recognition · Computer Science 2017-07-20 César Roberto de Souza , Adrien Gaidon , Yohann Cabon , Antonio Manuel López Peña

Human-centric Video Anomaly Detection (VAD) aims to identify human behaviors that deviate from normal. At its core, human-centric VAD faces substantial challenges, such as the complexity of diverse human behaviors, the rarity of anomalies,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Armin Danesh Pazho , Shanle Yao , Ghazal Alinezhad Noghre , Babak Rahimi Ardabili , Vinit Katariya , Hamed Tabkhi

Accurate prediction of future person location and movement trajectory from an egocentric wearable camera can benefit a wide range of applications, such as assisting visually impaired people in navigation, and the development of mobility…

Computer Vision and Pattern Recognition · Computer Science 2023-01-02 Jianing Qiu , Frank P. -W. Lo , Xiao Gu , Yingnan Sun , Shuo Jiang , Benny Lo

Video surveillance can be significantly enhanced by using both top-view data, e.g., those from drone-mounted cameras in the air, and horizontal-view data, e.g., those from wearable cameras on the ground. Collaborative analysis of…

Computer Vision and Pattern Recognition · Computer Science 2019-07-29 Ruize Han , Yujun Zhang , Wei Feng , Chenxing Gong , Xiaoyu Zhang , Jiewen Zhao , Liang Wan , Song Wang

Human activity, which usually consists of several actions, generally covers interactions among persons and or objects. In particular, human actions involve certain spatial and temporal relationships, are the components of more complicated…

Computer Vision and Pattern Recognition · Computer Science 2020-06-30 Zhenyu Liu , Yaqiang Yao , Yan Liu , Yuening Zhu , Zhenchao Tao , Lei Wang , Yuhong Feng

Recent advancements in human video synthesis have enabled the generation of high-quality videos through the application of stable diffusion models. However, existing methods predominantly concentrate on animating solely the human element…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Jinlin Liu , Kai Yu , Mengyang Feng , Xiefan Guo , Miaomiao Cui

Dense video captioning (DVC) aims to generate multi-sentence descriptions to elucidate the multiple events in the video, which is challenging and demands visual consistency, discoursal coherence, and linguistic diversity. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2021-11-22 Xu Yan , Zhengcong Fei , Shuhui Wang , Qingming Huang , Qi Tian

Videos are more informative than images because they capture the dynamics of the scene. By representing motion in videos, we can capture dynamic activities. In this work, we introduce GPT-4 generated motion descriptions that capture…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Chinmaya Devaraj , Cornelia Fermuller , Yiannis Aloimonos

We introduce CameraBench, a large-scale dataset and benchmark designed to assess and improve camera motion understanding. CameraBench consists of ~3,000 diverse internet videos, annotated by experts through a rigorous multi-stage quality…

Computer Vision and Pattern Recognition · Computer Science 2025-09-01 Zhiqiu Lin , Siyuan Cen , Daniel Jiang , Jay Karhade , Hewei Wang , Chancharik Mitra , Tiffany Ling , Yuhan Huang , Sifan Liu , Mingyu Chen , Rushikesh Zawar , Xue Bai , Yilun Du , Chuang Gan , Deva Ramanan