中文
相关论文

相关论文: M3Act: Learning from Synthetic Human Group Activit…

200 篇论文

Advancements in deep neural networks have contributed to near perfect results for many computer vision problems such as object recognition, face recognition and pose estimation. However, human action recognition is still far from…

计算机视觉与模式识别 · 计算机科学 2021-10-11 Asanka G. Perera , Yee Wei Law , Titilayo T. Ogunwa , Javaan Chahl

Over the past few years there has been major progress in the field of synthetic data generation using simulation based techniques. These methods use high-end graphics engines and physics-based ray-tracing rendering in order to represent the…

计算机视觉与模式识别 · 计算机科学 2022-06-02 Paul Yudkin , Eli Friedman , Orly Zvitia , Gil Elbaz

3D functionality segmentation aims to identify the interactive element in a 3D scene required to perform an action described in free-form language (e.g., the handle to ``Open the second drawer of the cabinet near the bed''). Progress has…

3D animation of humans in action is quite challenging as it involves using a huge setup with several motion trackers all over the person's body to track the movements of every limb. This is time-consuming and may cause the person discomfort…

图形学 · 计算机科学 2020-02-10 Laxman Kumarapu , Prerana Mukherjee

Modeling and generating human reactions poses a significant challenge with broad applications for computer vision and human-computer interaction. Existing methods either treat multiple individuals as a single entity, directly generating…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Xiyan Xu , Sirui Xu , Yu-Xiong Wang , Liang-Yan Gui

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person vision, it is still…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Francesco Ragusa , Antonino Furnari , Giovanni Maria Farinella

Complex activities in real-world audio unfold over extended durations and exhibit hierarchical structure, yet most prior work focuses on short clips and isolated events. To bridge this gap, we introduce MultiAct, a new dataset and benchmark…

声音 · 计算机科学 2026-02-09 Peng Zhang , Qingyu Luo , Philip J. B. Jackson , Wenwu Wang

We present a method for teaching machines to understand and model the underlying spatial common sense of diverse human-object interactions in 3D in a self-supervised way. This is a challenging task, as there exist specific manifolds of the…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Sookwan Han , Hanbyul Joo

Controllable cooperative humanoid manipulation is a fundamental yet challenging problem for embodied intelligence, due to severe data scarcity, complexities in multi-agent coordination, and limited generalization across objects. In this…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Wei Yao , Haohan Ma , Hongwen Zhang , Yunlian Sun , Liangjun Xing , Zhile Yang , Yuanjun Guo , Yebin Liu , Jinhui Tang

Video-to-video synthesis is a challenging problem aiming at learning a translation function between a sequence of semantic maps and a photo-realistic video depicting the characteristics of a driving video. We propose a head-to-head system…

计算机视觉与模式识别 · 计算机科学 2020-06-19 Mohammad Rami Koujan , Michail Christos Doukas , Anastasios Roussos , Stefanos Zafeiriou

We address the computational problem of novel human pose synthesis. Given an image of a person and a desired pose, we produce a depiction of that person in that pose, retaining the appearance of both the person and background. We present a…

计算机视觉与模式识别 · 计算机科学 2018-04-23 Guha Balakrishnan , Amy Zhao , Adrian V. Dalca , Fredo Durand , John Guttag

Interactive humanoid video generation aims to synthesize lifelike visual agents that can engage with humans through continuous and responsive video. Despite recent advances in video synthesis, existing methods often grapple with the…

The ability to detect manipulated visual content is becoming increasingly important in many application fields, given the rapid advances in image synthesis methods. Of particular concern is the possibility of modifying the content of…

图像与视频处理 · 电气工程与系统科学 2024-02-02 Giada Zingarini , Davide Cozzolino , Riccardo Corvi , Giovanni Poggi , Luisa Verdoliva

Data-driven methods have great advantages in modeling complicated human behavioral dynamics and dealing with many human-robot interaction applications. However, collecting massive and annotated real-world human datasets has been a laborious…

机器人学 · 计算机科学 2024-03-12 Zhuo Xu , Rui Zhou , Yida Yin , Huidong Gao , Masayoshi Tomizuka , Jiachen Li

Human beings always engage in a vast range of activities and tasks that demonstrate their ability to adapt to different scenarios. Any human activity can be represented as a temporal sequence of actions performed to achieve a certain goal.…

计算机视觉与模式识别 · 计算机科学 2023-07-21 Vinayak Gupta , Srikanta Bedathur

Video-driven human reaction generation aims to synthesize 3D human motions that directly react to observed video sequences, which is crucial for building human-like interactive AI systems. However, existing methods often fail to effectively…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Yuan Zhou , Yongzhi Li , Yanqi Dai , Xingyu Zhu , Yi Tan , Qingshan Xu , Beier Zhu , Richang Hong , Hanwang Zhang

Generating useful network summaries is a challenging and important problem with several applications like sensemaking, visualization, and compression. However, most of the current work in this space do not take human feedback into account…

Visual-based human action recognition can be found in various application fields, e.g., surveillance systems, sports analytics, medical assistive technologies, or human-robot interaction frameworks, and it concerns the identification and…

计算机视觉与模式识别 · 计算机科学 2024-12-23 Antonios Gasteratos , Stavros N. Moutsis , Konstantinos A. Tsintotas , Yiannis Aloimonos

We present LARNet, a novel end-to-end approach for generating human action videos. A joint generative modeling of appearance and dynamics to synthesize a video is very challenging and therefore recent works in video synthesis have proposed…

计算机视觉与模式识别 · 计算机科学 2021-10-28 Naman Biyani , Aayush J Rana , Shruti Vyas , Yogesh S Rawat

Visuomotor policies have shown great promise in robotic manipulation but often require substantial amounts of human-collected data for effective performance. A key reason underlying the data demands is their limited spatial generalization…

机器人学 · 计算机科学 2025-02-25 Zhengrong Xue , Shuying Deng , Zhenyang Chen , Yixuan Wang , Zhecheng Yuan , Huazhe Xu