中文
相关论文

相关论文: PoseTraj: Pose-Aware Trajectory Control in Video D…

200 篇论文

We introduce a Transformer based 6D Object Pose Estimation framework VideoPose, comprising an end-to-end attention based modelling architecture, that attends to previous frames in order to estimate accurate 6D Object Poses in videos. Our…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Apoorva Beedu , Huda Alamri , Irfan Essa

Estimating 3D poses of multiple humans in real-time is a classic but still challenging task in computer vision. Its major difficulty lies in the ambiguity in cross-view association of 2D poses and the huge state space when there are…

计算机视觉与模式识别 · 计算机科学 2021-07-30 Long Chen , Haizhou Ai , Rui Chen , Zijie Zhuang , Shuang Liu

We introduce FaceCam, a system that generates video under customizable camera trajectories for monocular human portrait video input. Recent camera control approaches based on large video-generation models have shown promising progress but…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Weijie Lyu , Ming-Hsuan Yang , Zhixin Shu

Data-efficient training of robust robot policies is the key to unlocking automation in a wide array of novel tasks. Current systems require large volumes of demonstrations to achieve robustness, which is impractical in many applications.…

机器人学 · 计算机科学 2026-03-10 Adam Hung , Bardienus Pieter Duisterhof , Jeffrey Ichnowski

Reconstruction of 3D neural fields from posed images has emerged as a promising method for self-supervised representation learning. The key challenge preventing the deployment of these 3D scene learners on large-scale video data is their…

计算机视觉与模式识别 · 计算机科学 2023-06-02 Cameron Smith , Yilun Du , Ayush Tewari , Vincent Sitzmann

We present a novel approach for hand-object action recognition that leverages 2D point tracks as an additional motion cue. While most existing methods rely on RGB appearance, human pose estimation, or their combination, our work…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Dennis Holzmann , Sven Wachsmuth

In recent years, generative artificial intelligence has achieved significant advancements in the field of image generation, spawning a variety of applications. However, video generation still faces considerable challenges in various…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Yuang Zhang , Jiaxi Gu , Li-Wen Wang , Han Wang , Junqi Cheng , Yuefeng Zhu , Fangyuan Zou

Object pose tracking is a fundamental and essential task for robotics to perform tasks in the home and industrial settings. The most commonly used sensors to do so are RGB-D cameras, which can hit limitations in highly dynamic environments…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Zhichao Li , Chiara Bartolozzi , Lorenzo Natale , Arren Glover

Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Zhixue Fang , Xu He , Songlin Tang , Haoxian Zhang , Qingfeng Li , Xiaoqiang Liu , Pengfei Wan , Kun Gai

Online monocular 3D reconstruction enables dense scene recovery from streaming video but remains fundamentally limited by the stability-adaptation dilemma: the reconstruction model must rapidly incorporate novel viewpoints while preserving…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Lanbo Xu , Liang Guo , Caigui Jiang , Cheng Wang

Data scarcity continues to be a major challenge in the field of robotic manipulation. Although diffusion models provide a promising solution for generating robotic manipulation videos, existing methods largely depend on 2D trajectories,…

机器人学 · 计算机科学 2025-11-14 Ying Li , Xiaobao Wei , Xiaowei Chi , Yuming Li , Zhongyu Zhao , Hao Wang , Ningning Ma , Ming Lu , Sirui Han , Shanghang Zhang

Category-level 3D pose estimation is a fundamentally important problem in computer vision and robotics, e.g. for embodied agents or to train 3D generative models. However, so far methods that estimate the category-level object pose require…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Leonhard Sommer , Artur Jesslen , Eddy Ilg , Adam Kortylewski

We propose a training-free and robust solution to offer camera movement control for off-the-shelf video diffusion models. Unlike previous work, our method does not require any supervised finetuning on camera-annotated datasets or…

计算机视觉与模式识别 · 计算机科学 2025-02-26 Chen Hou , Zhibo Chen

Humans can effortlessly anticipate how objects might move or change through interaction--imagining a cup being lifted, a knife slicing, or a lid being closed. We aim to endow computational systems with a similar ability to predict plausible…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Rustin Soraki , Homanga Bharadhwaj , Ali Farhadi , Roozbeh Mottaghi

Single-view 3D shape retrieval is a fundamental yet challenging task that is increasingly important with the growth of available 3D data. Existing approaches largely fall into two categories: those using contrastive learning to map point…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Jiaxin Shi , Guofeng Zhang , Wufei Ma , Naifu Liang , Adam Kortylewski , Alan Yuille

Existing monocular 3D pose estimation methods primarily rely on joint positional features, while overlooking intrinsic directional and angular correlations within the skeleton. As a result, they often produce implausible poses under joint…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Ming Xu , Xu Zhang

Most 3d human pose estimation methods assume that input -- be it images of a scene collected from one or several viewpoints, or from a video -- is given. Consequently, they focus on estimates leveraging prior knowledge and measurement by…

计算机视觉与模式识别 · 计算机科学 2020-12-17 Erik Gärtner , Aleksis Pirinen , Cristian Sminchisescu

Camera localization in 3D LiDAR maps has gained increasing attention due to its promising ability to handle complex scenarios, surpassing the limitations of visual-only localization methods. However, existing methods mostly focus on…

机器人学 · 计算机科学 2024-10-28 Huai Yu , Kuangyi Chen , Wen Yang , Sebastian Scherer , Gui-Song Xia

Motion control is crucial for generating expressive and compelling video content; however, most existing video generation models rely mainly on text prompts for control, which struggle to capture the nuances of dynamic actions and temporal…

Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) approaches. Existing AR video diffusion models are limited to…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Kesen Zhao , Jiaxin Shi , Beier Zhu , Junbao Zhou , Xiaolong Shen , Yuan Zhou , Qianru Sun , Hanwang Zhang