中文
相关论文

相关论文: GazeMoDiff: Gaze-guided Diffusion Model for Stocha…

200 篇论文

Following the gaze of other people and analyzing the target they are looking at can help us understand what they are thinking, and doing, and predict the actions that may follow. Existing methods for gaze following struggle to perform well…

计算机视觉与模式识别 · 计算机科学 2025-02-05 Feiyang Liu , Dan Guo , Jingyuan Xu , Zihao He , Shengeng Tang , Kun Li , Meng Wang

Many density estimation techniques for 3D human motion prediction require a significant amount of inference time, often exceeding the duration of the predicted time horizon. To address the need for faster density estimation for 3D human…

计算机视觉与模式识别 · 计算机科学 2026-02-02 Takahiro Maeda , Jinkun Cao , Norimichi Ukita , Kris Kitani

An increasing number of works explore collaborative human-computer systems in which human gaze is used to enhance computer vision systems. For object detection these efforts were so far restricted to late integration approaches that have…

计算机视觉与模式识别 · 计算机科学 2015-05-22 Iaroslav Shcherbatyi , Andreas Bulling , Mario Fritz

Recent advancements utilizing large-scale video data for learning video generation models demonstrate significant potential in understanding complex physical dynamics. It suggests the feasibility of leveraging diverse robot trajectory data…

计算机视觉与模式识别 · 计算机科学 2024-11-15 Youpeng Wen , Junfan Lin , Yi Zhu , Jianhua Han , Hang Xu , Shen Zhao , Xiaodan Liang

Gaze estimation, the task of predicting where an individual is looking, is a critical task with direct applications in areas such as human-computer interaction and virtual reality. Estimating the direction of looking in unconstrained…

计算机视觉与模式识别 · 计算机科学 2024-02-14 Andy Cătrună , Adrian Cosma , Emilian Rădoi

By borrowing the wisdom of human in gaze following, we propose a two-stage solution for gaze point prediction of the target persons in a scene. Specifically, in the first stage, both head image and its position are fed into a gaze direction…

计算机视觉与模式识别 · 计算机科学 2019-07-05 Dongze Lian , Zehao Yu , Shenghua Gao

We present PoseDiff, a conditional diffusion model that unifies robot state estimation and control within a single framework. At its core, PoseDiff maps raw visual observations into structured robot states-such as 3D keypoints or joint…

机器人学 · 计算机科学 2025-11-03 Haozhuo Zhang , Michele Caprio , Jing Shao , Qiang Zhang , Jian Tang , Shanghang Zhang , Wei Pan

In imitation learning for robotic manipulation, decomposing object manipulation tasks into sub-tasks enables the reuse of learned skills and the combination of learned behaviors to perform novel tasks, rather than simply replicating…

机器人学 · 计算机科学 2025-02-28 Ryo Takizawa , Yoshiyuki Ohmura , Yasuo Kuniyoshi

The existing state-of-the-art method for audio-visual conditioned video prediction uses the latent codes of the audio-visual frames from a multimodal stochastic network and a frame encoder to predict the next visual frame. However, a direct…

计算机视觉与模式识别 · 计算机科学 2023-09-21 Yating Xu , Conghui Hu , Gim Hee Lee

Gaze estimation involves predicting where the person is looking at within an image or video. Technically, the gaze information can be inferred from two different magnification levels: face orientation and eye orientation. The inference is…

计算机视觉与模式识别 · 计算机科学 2021-10-27 Ashesh , Chu-Song Chen , Hsuan-Tien Lin

Animating stylized avatars with dynamic poses and expressions has attracted increasing attention for its broad range of applications. Previous research has made significant progress by training controllable generative models to synthesize…

Gaze stabilization is critical for enabling fluid, accurate, and efficient interaction in immersive augmented reality (AR) environments, particularly during task-oriented visual behaviors. However, fixation sequences captured in active gaze…

人机交互 · 计算机科学 2025-10-03 Yaozheng Xia , Zaiping Zhu , Bo Pang , Shaorong Wang , Sheng Li

This paper addresses the gaze target detection problem in single images captured from the third-person perspective. We present a multimodal deep architecture to infer where a person in a scene is looking. This spatial model is trained on…

计算机视觉与模式识别 · 计算机科学 2022-08-24 Francesco Tonini , Cigdem Beyan , Elisa Ricci

A long-standing goal in computer vision is to capture, model, and realistically synthesize human behavior. Specifically, by learning from data, our goal is to enable virtual humans to navigate within cluttered indoor scenes and naturally…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Mohamed Hassan , Duygu Ceylan , Ruben Villegas , Jun Saito , Jimei Yang , Yi Zhou , Michael Black

This report reviews recent advancements in human motion prediction, reconstruction, and generation. Human motion prediction focuses on forecasting future poses and movements from historical data, addressing challenges like nonlinear…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Canxuan Gang , Yiran Wang

Controllable generation of 3D human motions becomes an important topic as the world embraces digital transformation. Existing works, though making promising progress with the advent of diffusion models, heavily rely on meticulously captured…

计算机视觉与模式识别 · 计算机科学 2024-01-25 Nhat M. Hoang , Kehong Gong , Chuan Guo , Michael Bi Mi

Estimating full-body human motion via sparse tracking signals from head-mounted displays and hand controllers in 3D scenes is crucial to applications in AR/VR. One of the biggest challenges to this task is the one-to-many mapping from…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Jiangnan Tang , Jingya Wang , Kaiyang Ji , Lan Xu , Jingyi Yu , Ye Shi

Eye gaze can provide rich information on human psychological activities, and has garnered significant attention in the field of Human-Robot Interaction (HRI). However, existing gaze estimation methods merely predict either the gaze…

计算机视觉与模式识别 · 计算机科学 2025-05-23 Haoming Huang , Musen Zhang , Jianxin Yang , Zhen Li , Jinkai Li , Yao Guo

Existing diffusion-based methods have achieved impressive results in human motion editing. However, these methods often exhibit significant ghosting and body distortion in unseen in-the-wild cases. In this paper, we introduce…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Yi Zuo , Lingling Li , Licheng Jiao , Fang Liu , Xu Liu , Wenping Ma , Shuyuan Yang , Yuwei Guo

This paper presents a motion data augmentation scheme incorporating motion synthesis encouraging diversity and motion correction imposing physical plausibility. This motion synthesis consists of our modified Variational AutoEncoder (VAE)…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Takahiro Maeda , Norimichi Ukita