中文
相关论文

相关论文: GazeMoDiff: Gaze-guided Diffusion Model for Stocha…

200 篇论文

Probabilistic human motion prediction aims to forecast multiple possible future movements from past observations. While current approaches report high diversity and realism, they often generate motions with undetected limb stretching and…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Cecilia Curreli , Dominik Muhle , Abhishek Saroha , Zhenzhang Ye , Riccardo Marin , Daniel Cremers

We have recently seen tremendous progress in diffusion advances for generating realistic human motions. Yet, they largely disregard the multi-human interactions. In this paper, we present InterGen, an effective diffusion-based approach that…

计算机视觉与模式识别 · 计算机科学 2024-03-29 Han Liang , Wenqian Zhang , Wenxuan Li , Jingyi Yu , Lan Xu

Vision Transformers (ViT) have advanced computer vision, yet their efficacy in complex tasks like driving remains less explored. This study enhances ViT by integrating human eye gaze, captured via eye-tracking, to increase prediction…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Sharath Koorathota , Nikolas Papadopoulos , Jia Li Ma , Shruti Kumar , Xiaoxiao Sun , Arunesh Mittal , Patrick Adelman , Paul Sajda

We present GazeGen, a user interaction system that generates visual content (images and videos) for locations indicated by the user's eye gaze. GazeGen allows intuitive manipulation of visual content by targeting regions of interest with…

计算机视觉与模式识别 · 计算机科学 2024-11-19 He-Yen Hsieh , Ziyun Li , Sai Qian Zhang , Wei-Te Mark Ting , Kao-Den Chang , Barbara De Salvo , Chiao Liu , H. T. Kung

Objective. Motion artifacts in brain MRI, mainly from rigid head motion, degrade image quality and hinder downstream applications. Conventional methods to mitigate these artifacts, including repeated acquisitions or motion tracking, impose…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Mojtaba Safari , Shansong Wang , Qiang Li , Zach Eidex , Richard L. J. Qiu , Chih-Wei Chang , Hui Mao , Xiaofeng Yang

Human motion modeling traditionally separates motion generation and estimation into distinct tasks with specialized models. Motion generation models focus on creating diverse, realistic motions from inputs like text, audio, or keyframes,…

图形学 · 计算机科学 2025-05-05 Jiefeng Li , Jinkun Cao , Haotian Zhang , Davis Rempe , Jan Kautz , Umar Iqbal , Ye Yuan

Text-driven human motion generation in computer vision is both significant and challenging. However, current methods are limited to producing either deterministic or imprecise motion sequences, failing to effectively control the temporal…

计算机视觉与模式识别 · 计算机科学 2023-09-13 Yin Wang , Zhiying Leng , Frederick W. B. Li , Shun-Cheng Wu , Xiaohui Liang

This article suggests a reasoning-guided vision-language-motion diffusion framework (RG-VLMD) for generating instruction-aware co-speech gestures for humanoid robots in educational scenarios. The system integrates multi-modal affective…

机器人学 · 计算机科学 2026-03-20 Fuze Sun , Lingyu Li , Lekan Dai , Xinyu Fan

Due to the recent outbreak of COVID-19, many classes, exams, and meetings have been conducted non-face-to-face. However, the foundation for video conferencing solutions is still insufficient. So this technology has become an important…

计算机视觉与模式识别 · 计算机科学 2021-07-15 Suneung-Kim , Seong-Whan Lee

This paper tackles the problem of passive gaze estimation using both event and frame data. Considering the inherently different physiological structures, it is intractable to accurately estimate gaze purely based on a given state. Thus, we…

计算机视觉与模式识别 · 计算机科学 2024-07-01 Jiading Li , Zhiyu Zhu , Jinhui Hou , Junhui Hou , Jinjian Wu

We propose DemoDiffusion, a simple method for enabling robots to perform manipulation tasks by imitating a single human demonstration, without requiring task-specific training or paired human-robot data. Our approach is based on two…

机器人学 · 计算机科学 2026-03-10 Sungjae Park , Homanga Bharadhwaj , Shubham Tulsiani

Predicting human motion plays a crucial role in ensuring a safe and effective human-robot close collaboration in intelligent remanufacturing systems of the future. Existing works can be categorized into two groups: those focusing on…

机器人学 · 计算机科学 2023-08-01 Sibo Tian , Minghui Zheng , Xiao Liang

In this paper, we tackle the problem of scene-aware 3D human motion forecasting. A key challenge of this task is to predict future human motions that are consistent with the scene by modeling the human-scene interactions. While recent works…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Chaoyue Xing , Wei Mao , Miaomiao Liu

Forecasting dynamic scenes remains a fundamental challenge in computer vision, as limited observations make it difficult to capture coherent object-level motion and long-term temporal evolution. We present Motion Group-aware Gaussian…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Junmyeong Lee , Hoseung Choi , Minsu Cho

Recently, human motion analysis has experienced great improvement due to inspiring generative models such as the denoising diffusion model and large language model. While the existing approaches mainly focus on generating motions with…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Yiming Wu , Wei Ji , Kecheng Zheng , Zicheng Wang , Dong Xu

Online egocentric gaze estimation predicts where a camera wearer is looking from first-person video using only past and current frames, a task essential for augmented reality and assistive technologies. Unlike third-person gaze estimation,…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Jia Li , Wenjie Zhao , Shijian Deng , Bolin Lai , Yuheng Wu , RUijia Chen , Jon E. Froehlich , Yuhang Zhao , Yapeng Tian

Imitation learning by behavioral cloning is a prevalent method that has achieved some success in vision-based autonomous driving. The basic idea behind behavioral cloning is to have the neural network learn from observing a human expert's…

计算机视觉与模式识别 · 计算机科学 2019-08-19 Yuying Chen , Congcong Liu , Lei Tai , Ming Liu , Bertram E. Shi

Volumetric video emerges as a new attractive video paradigm in recent years since it provides an immersive and interactive 3D viewing experience with six degree-of-freedom (DoF). Unlike traditional 2D or panoramic videos, volumetric videos…

多媒体 · 计算机科学 2023-08-17 Kaiyuan Hu , Haowen Yang , Yili Jin , Junhua Liu , Yongting Chen , Miao Zhang , Fangxin Wang

The task of text2motion is to generate human motion sequences from given textual descriptions, where the model explores diverse mappings from natural language instructions to human body movements. While most existing works are confined to…

人工智能 · 计算机科学 2024-03-27 Kunhang Li , Yansong Feng

Gaze correction aims to redirect the person's gaze into the camera by manipulating the eye region, and it can be considered as a specific image resynthesis problem. Gaze correction has a wide range of applications in real life, such as…

计算机视觉与模式识别 · 计算机科学 2019-06-04 Jichao Zhang , Meng Sun , Jingjing Chen , Hao Tang , Yan Yan , Xueying Qin , Nicu Sebe