English
Related papers

Related papers: ReMoGen: Real-time Human Interaction-to-Reaction G…

200 papers

Egocentric human motion generation and forecasting with scene-context is crucial for enhancing AR/VR experiences, improving human-robot interaction, advancing assistive technologies, and enabling adaptive healthcare solutions by accurately…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Chaitanya Patel , Hiroki Nakamura , Yuta Kyuragi , Kazuki Kozuka , Juan Carlos Niebles , Ehsan Adeli

In recent advances of deep generative models, face reenactment -manipulating and controlling human face, including their head movement-has drawn much attention for its wide range of applicability. Despite its strong expressiveness, it is…

Computer Vision and Pattern Recognition · Computer Science 2022-02-23 Takuya Yashima , Takuya Narihira , Tamaki Kojima

GANs provide a framework for training generative models which mimic a data distribution. However, in many cases we wish to train these generative models to optimize some auxiliary objective function within the data it generates, such as…

Computer Vision and Pattern Recognition · Computer Science 2017-10-02 Andrew Kyle Lampinen , David So , Douglas Eck , Fred Bertsch

Responsing with image has been recognized as an important capability for an intelligent conversational agent. Yet existing works only focus on exploring the multimodal dialogue models which depend on retrieval-based methods, but neglecting…

Computation and Language · Computer Science 2022-03-30 Qingfeng Sun , Yujing Wang , Can Xu , Kai Zheng , Yaming Yang , Huang Hu , Fei Xu , Jessica Zhang , Xiubo Geng , Daxin Jiang

Human-like multimodal reaction generation is essential for natural group interactions between humans and embodied AI. However, existing approaches are limited to single-modality or speaking-only responses in dyadic interactions, making them…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Zhi-Yi Lin , Thomas Markhorst , Jouh Yeong Chew , Xucong Zhang

Extracting human motion from large-scale web videos offers a scalable solution to the data scarcity issue in character animation. However, some human parts in many video frames cannot be seen due to off-screen captures or occlusions. It…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Boyuan Li , Sipeng Zheng , Bin Cao , Ruihua Song , Zongqing Lu

We present ScaleMoGen, a scale-wise autoregressive framework for text-driven human motion generation. Unlike conventional autoregressive approaches that rely on standard next-token prediction, ScaleMoGen frames motion generation as a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Inwoo Hwang , Hojun Jang , Bing Zhou , Jian Wang , Young Min Kim , Chuan Guo

Current datasets for action recognition tasks face limitations stemming from traditional collection and generation methods, including the constrained range of action classes, absence of multi-viewpoint recordings, limited diversity, poor…

Computer Vision and Pattern Recognition · Computer Science 2024-01-25 Xingyu Song , Zhan Li , Shi Chen , Kazuyuki Demachi

Image-to-video (I2V) generation seeks to produce realistic motion sequences from a single reference image. Although recent methods exhibit strong temporal consistency, they often struggle when dealing with complex, non-repetitive human…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Ashkan Taghipour , Morteza Ghahremani , Mohammed Bennamoun , Farid Boussaid , Aref Miri Rekavandi , Zinuo Li , Qiuhong Ke , Hamid Laga

Face-to-face communication, as a common human activity, motivates the research on interactive head generation. A virtual agent can generate motion responses with both listening and speaking capabilities based on the audio or motion signals…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Ying Guo , Xi Liu , Cheng Zhen , Pengfei Yan , Xiaoming Wei

We propose a novel framework, On-Demand MOtion Generation (ODMO), for generating realistic and diverse long-term 3D human motion sequences conditioned only on action types with an additional capability of customization. ODMO shows…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Qiujing Lu , Yipeng Zhang , Mingjian Lu , Vwani Roychowdhury

Accurately reconstructing complex full multi-object scenes from sparse observations remains a core challenge in computer vision and a key step toward scalable and reliable simulation for robotics. In this work, we introduce RecGen, a…

Advances in video generation have significantly improved the realism and quality of created scenes. This has fueled interest in developing intuitive tools that let users leverage video generation as world simulators. Text-to-video (T2V)…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Zuhao Liu , Aleksandar Yanev , Ahmad Mahmood , Ivan Nikolov , Saman Motamed , Wei-Shi Zheng , Xi Wang , Lei Sun , Luc Van Gool , Danda Pani Paudel

Given the audio-visual clip of the speaker, facial reaction generation aims to predict the listener's facial reactions. The challenge lies in capturing the relevance between video and audio while balancing appropriateness, realism, and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Jiaming Li , Sheng Wang , Xin Wang , Yitao Zhu , Honglin Xiong , Zixu Zhuang , Qian Wang

The objective of the Multiple Appropriate Facial Reaction Generation (MAFRG) task is to produce contextually appropriate and diverse listener facial behavioural responses based on the multimodal behavioural data of the conversational…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Guanyu Hu , Jie Wei , Siyang Song , Dimitrios Kollias , Xinyu Yang , Zhonglin Sun , Odysseus Kaloidas

Humanoid whole-body loco-manipulation promises transformative capabilities for daily service and warehouse tasks. While recent advances in general motion tracking (GMT) have enabled humanoids to reproduce diverse human motions, these…

Robotics · Computer Science 2025-10-09 Siheng Zhao , Yanjie Ze , Yue Wang , C. Karen Liu , Pieter Abbeel , Guanya Shi , Rocky Duan

Imitation learning from large-scale, diverse human demonstrations has been shown to be effective for training robots, but collecting such data is costly and time-consuming. This challenge intensifies for multi-step bimanual mobile…

We present ReCoM, an efficient framework for generating high-fidelity and generalizable human body motions synchronized with speech. The core innovation lies in the Recurrent Embedded Transformer (RET), which integrates Dynamic Embedding…

Graphics · Computer Science 2025-03-31 Yong Xie , Yunlian Sun , Hongwen Zhang , Yebin Liu , Jinhui Tang

When designing robots to assist in everyday human activities, it is crucial to enhance user requests with visual cues from their surroundings for improved intent understanding. This process is defined as a multimodal classification task.…

Computation and Language · Computer Science 2025-06-18 Shang-Chi Tsai , Seiya Kawano , Angel Garcia Contreras , Koichiro Yoshino , Yun-Nung Chen

Human-robot interaction is increasingly moving toward multi-robot, socially grounded environments. Existing systems struggle to integrate multimodal perception, embodied expression, and coordinated decision-making in a unified framework.…

Robotics · Computer Science 2026-03-25 Shaid Hasan , Breenice Lee , Sujan Sarker , Tariq Iqbal
‹ Prev 1 4 5 6 7 8 10 Next ›