中文
相关论文

相关论文: M3Act: Learning from Synthetic Human Group Activit…

200 篇论文

With the continuously thriving popularity around the world, fitness activity analytic has become an emerging research topic in computer vision. While a variety of new tasks and algorithms have been proposed recently, there are growing…

计算机视觉与模式识别 · 计算机科学 2023-04-20 Yansong Tang , Jinpeng Liu , Aoyang Liu , Bin Yang , Wenxun Dai , Yongming Rao , Jiwen Lu , Jie Zhou , Xiu Li

We introduce the Unity Perception package which aims to simplify and accelerate the process of generating synthetic datasets for computer vision tasks by offering an easy-to-use and highly customizable toolset. This open-source package…

Human behavior understanding is arguably one of the most important mid-level components in artificial intelligence. In order to efficiently make use of data, multi-task learning has been studied in diverse computer vision tasks including…

计算机视觉与模式识别 · 计算机科学 2018-02-15 Dong-Jin Kim , Jinsoo Choi , Tae-Hyun Oh , Youngjin Yoon , In So Kweon

Human head detection, keypoint estimation, and 3D head model fitting are essential tasks with many applications. However, traditional real-world datasets often suffer from bias, privacy, and ethical concerns, and they have been recorded in…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Orest Kupyn , Eugene Khvedchenia , Christian Rupprecht

Language-guided human motion synthesis has been a challenging task due to the inherent complexity and diversity of human behaviors. Previous methods face limitations in generalization to novel actions, often resulting in unrealistic or…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Yuanhao Zhai , Mingzhen Huang , Tianyu Luan , Lu Dong , Ifeoma Nwogu , Siwei Lyu , David Doermann , Junsong Yuan

While recent multimodal models have shown progress in vision-language tasks, small-scale variants still struggle with the fine-grained temporal reasoning required for video understanding. We introduce ReasonAct, a method that enhances video…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Jiaxin Liu , Zhaolu Kang

Synthesizing interaction-involved human motions has been challenging due to the high complexity of 3D environments and the diversity of possible human behaviors within. We present LAMA, Locomotion-Action-MAnipulation, to synthesize natural…

计算机视觉与模式识别 · 计算机科学 2023-09-11 Jiye Lee , Hanbyul Joo

For decades, human-computer interaction has fundamentally been manual. Even today, almost all productive work done on the computer necessitates human input at every step. Autonomous virtual agents represent an exciting step in automating…

Humans perform a variety of interactive motions, among which duet dance is one of the most challenging interactions. However, in terms of human motion generative models, existing works are still unable to generate high-quality interactive…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Ronghui Li , Youliang Zhang , Yachao Zhang , Yuxiang Zhang , Mingyang Su , Jie Guo , Ziwei Liu , Yebin Liu , Xiu Li

Relational object rearrangement (ROR) tasks (e.g., insert flower to vase) require a robot to manipulate objects with precise semantic and geometric reasoning. Existing approaches either rely on pre-collected demonstrations that struggle to…

机器人学 · 计算机科学 2025-09-23 Liang Heng , Jiadong Xu , Yiwen Wang , Xiaoqi Li , Muhe Cai , Yan Shen , Juan Zhu , Guanghui Ren , Hao Dong

Recent progress on physics-based character animation has shown impressive breakthroughs on human motion synthesis, through imitating motion capture data via deep reinforcement learning. However, results have mostly been demonstrated on…

计算机视觉与模式识别 · 计算机科学 2020-12-17 Yu-Wei Chao , Jimei Yang , Weifeng Chen , Jia Deng

Modeling crowd behavior relies on accurate data of pedestrian movements at a high level of detail. Imaging sensors such as cameras provide a good basis for capturing such detailed pedestrian motion data. However, currently available…

计算机视觉与模式识别 · 计算机科学 2012-10-11 Stefan Seer , Norbert Brändle , Carlo Ratti

The domain of Embodied AI, in which agents learn to complete tasks through interaction with their environment from egocentric observations, has experienced substantial growth with the advent of deep reinforcement learning and increased…

计算机视觉与模式识别 · 计算机科学 2020-08-31 Luca Weihs , Jordi Salvador , Klemen Kotar , Unnat Jain , Kuo-Hao Zeng , Roozbeh Mottaghi , Aniruddha Kembhavi

Synthesis of long-term human motion skeleton sequences is essential to aid human-centric video generation with potential applications in Augmented Reality, 3D character animations, pedestrian trajectory prediction, etc. Long-term human…

计算机视觉与模式识别 · 计算机科学 2020-12-22 Neeraj Battan , Yudhik Agrawal , Veeravalli Saisooryarao , Aman Goel , Avinash Sharma

The problem of task planning for artificial agents remains largely unsolved. While there has been increasing interest in data-driven approaches for the study of task planning for artificial agents, a significant remaining bottleneck is the…

计算机视觉与模式识别 · 计算机科学 2021-08-12 Jiafei Duan , Samson Yu , Hui Li Tan , Cheston Tan

Action recognition and human pose estimation are closely related but both problems are generally handled as distinct tasks in the literature. In this work, we propose a multitask framework for jointly 2D and 3D pose estimation from still…

计算机视觉与模式识别 · 计算机科学 2018-03-22 Diogo C. Luvizon , David Picard , Hedi Tabia

As two intimate reciprocal tasks, scene-aware human motion synthesis and analysis require a joint understanding between multiple modalities, including 3D body motions, 3D scenes, and textual descriptions. In this paper, we integrate these…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Xuehao Gao , Yang Yang , Shaoyi Du , Guo-Jun Qi , Junwei Han

Human activity understanding is crucial for building automatic intelligent system. With the help of deep learning, activity understanding has made huge progress recently. But some challenges such as imbalanced data distribution, action…

计算机视觉与模式识别 · 计算机科学 2019-08-07 Yong-Lu Li , Liang Xu , Xinpeng Liu , Xijie Huang , Yue Xu , Mingyang Chen , Ze Ma , Shiyi Wang , Hao-Shu Fang , Cewu Lu

Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input modalities, such as RGB or skeletal data, limiting their…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Tuyen Tran , Thao Minh Le , Hung Tran , Truyen Tran

Action recognition models have achieved impressive results by incorporating scene-level annotations, such as objects, their relations, 3D structure, and more. However, obtaining annotations of scene structure for videos requires a…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Roei Herzig , Ofir Abramovich , Elad Ben-Avraham , Assaf Arbelle , Leonid Karlinsky , Ariel Shamir , Trevor Darrell , Amir Globerson