中文
相关论文

相关论文: HERO: Human Reaction Generation from Videos

200 篇论文

Human motion generation has shown great advances thanks to the recent diffusion models trained on large-scale motion capture data. Most of existing works, however, currently target animation of isolated people in empty scenes. Meanwhile,…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Yangsong Zhang , Abdul Ahad Butt , Gül Varol , Ivan Laptev

Spatio-temporal coherency is a major challenge in synthesizing high quality videos, particularly in synthesizing human videos that contain rich global and local deformations. To resolve this challenge, previous approaches have resorted to…

计算机视觉与模式识别 · 计算机科学 2024-11-13 Yaohui Wang , Xin Ma , Xinyuan Chen , Cunjian Chen , Antitza Dantcheva , Bo Dai , Yu Qiao

Appearance editing according to user needs is a pivotal task in video editing. Existing text-guided methods often lead to ambiguities regarding user intentions and restrict fine-grained control over editing specific aspects of objects. To…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Tongtong Su , Chengyu Wang , Jun Huang , Dongming Lu

Many application areas ranging from serious games for health to learning by demonstration in robotics, could benefit from large body movement datasets extracted from textual instructions accompanied by images. The interpretation of…

人机交互 · 计算机科学 2020-06-09 Himangshu Sarma , Robert Porzel , Jan Smeddinck , Rainer Malaka

Reconstructing human-object interaction in 3D from a single RGB image is a challenging task and existing data driven methods do not generalize beyond the objects present in the carefully curated 3D interaction datasets. Capturing…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Xianghui Xie , Bharat Lal Bhatnagar , Jan Eric Lenssen , Gerard Pons-Moll

Generating human-object interactions (HOIs) is critical with the tremendous advances of digital avatars. Existing datasets are typically limited to humans interacting with a single object while neglecting the ubiquitous manipulation of…

计算机视觉与模式识别 · 计算机科学 2024-09-12 Xintao Lv , Liang Xu , Yichao Yan , Xin Jin , Congsheng Xu , Shuwen Wu , Yifan Liu , Lincheng Li , Mengxiao Bi , Wenjun Zeng , Xiaokang Yang

Generating human motions from textual descriptions has gained growing research interest due to its wide range of applications. However, only a few works consider human-scene interactions together with text conditions, which is crucial for…

计算机视觉与模式识别 · 计算机科学 2024-05-14 Zhi Cen , Huaijin Pi , Sida Peng , Zehong Shen , Minghui Yang , Shuai Zhu , Hujun Bao , Xiaowei Zhou

Human action recognition is a well-known computer vision and pattern recognition task of identifying which action a man is actually doing. Extracting the keypoint information of a single human with both spatial and temporal features of…

计算机视觉与模式识别 · 计算机科学 2021-08-13 Yuke , Yang

Recent advances in model architectures, compute, and data scale have driven rapid progress in video generation, producing increasingly realistic content. Yet, no prior method systematically measures how faithfully these systems render human…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Yusu Fang , Tiange Xiang , Tian Tan , Narayan Schuetz , Scott Delp , Li Fei-Fei , Ehsan Adeli

Reconstructing 3D human motion and human-object interactions (HOI) from Internet videos is a fundamental step toward building large-scale datasets of human behavior. Existing methods struggle to recover globally consistent 3D motion under…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Hongjie Li , Heng Yu , Jiaman Li , Hong-Xing Yu , Ehsan Adeli , C. Karen Liu , Jiajun Wu

Human emotion synthesis is a crucial aspect of affective computing. It involves using computational methods to mimic and convey human emotions through various modalities, with the goal of enabling more natural and effective human-computer…

机器学习 · 计算机科学 2024-12-11 Fei Ma , Yukan Li , Yifan Xie , Ying He , Yi Zhang , Hongwei Ren , Zhou Liu , Wei Yao , Fuji Ren , Fei Richard Yu , Shiguang Ni

Human videos contain rich manipulation priors, but using them for robot learning remains difficult because raw observations entangle scene understanding, human motion, and embodiment-specific action. We introduce MoT-HRA, a hierarchical…

机器人学 · 计算机科学 2026-05-22 Yifan Xie , YuAn Wang , Guangyu Chen , Jinkun Liu , Yu Sun , Wenbo Ding

Human motion synthesis in complex scenes presents a fundamental challenge, extending beyond conventional Text-to-Motion tasks by requiring the integration of diverse modalities such as static environments, movable objects, natural language…

图形学 · 计算机科学 2025-05-20 Zichen Geng , Zeeshan Hayder , Wei Liu , Ajmal Mian

This paper proposes a simple yet effective method for human action recognition in video. The proposed method separately extracts local appearance and motion features using state-of-the-art three-dimensional convolutional neural networks…

计算机视觉与模式识别 · 计算机科学 2020-02-24 David Torpey , Turgay Celik

This paper describes the development of a real-time Human-Robot Interaction (HRI) system for a service robot based on 3D human activity recognition and human-like decision mechanism. The Human-Robot Interactive (HRI) system, which allows…

人机交互 · 计算机科学 2019-01-14 Kang Li , Jinting Wu , Xiaoguang Zhao , Min Tan

Video Object Grounding (VOG) is the problem of associating spatial object regions in the video to a descriptive natural language query. This is a challenging vision-language task that necessitates constructing the correct cross-modal…

多媒体 · 计算机科学 2022-08-12 Mengze Li , Tianbao Wang , Haoyu Zhang , Shengyu Zhang , Zhou Zhao , Wenqiao Zhang , Jiaxu Miao , Shiliang Pu , Fei Wu

This paper proposes a human activity recognition method which is based on features learned from 3D video data without incorporating domain knowledge. The experiments on data collected by RGBD cameras produce results outperforming other…

计算机视觉与模式识别 · 计算机科学 2015-08-11 Ngu Nguyen

The goal of this study is to develop and analyze multimodal models for predicting experienced affective responses of viewers watching movie clips. We develop hybrid multimodal prediction models based on both the video and audio of the…

计算机视觉与模式识别 · 计算机科学 2019-09-18 Ha Thi Phuong Thao , Dorien Herremans , Gemma Roig

The growing prevalence of realistic AI-generated videos on media platforms increasingly blurs the line between fact and fiction, eroding public trust. Understanding how people watch AI-generated videos offers a human-centered perspective…

人机交互 · 计算机科学 2026-05-12 Danqing Shi , Lan Jiang , Katherine M. Collins , Shangzhe Wu , Ayush Tewari , Miri Zilka

Efficiently detecting human intent to interact with ubiquitous robots is crucial for effective human-robot interaction (HRI) and collaboration. Over the past decade, deep learning has gained traction in this field, with most existing…

机器人学 · 计算机科学 2025-09-29 Farida Mohsen , Ali Safa