中文
相关论文

相关论文: Diffusion Models Enable Zero-Shot Pose Estimation …

200 篇论文

Denoising diffusion models are a powerful type of generative models used to capture complex distributions of real-world signals. However, their applicability is limited to scenarios where training samples are readily available, which is not…

计算机视觉与模式识别 · 计算机科学 2023-11-20 Ayush Tewari , Tianwei Yin , George Cazenavette , Semon Rezchikov , Joshua B. Tenenbaum , Frédo Durand , William T. Freeman , Vincent Sitzmann

Nowadays, generative models are shaping various fields such as art, design, and human-computer interaction, yet accompanied by challenges related to copyright infringement and content management. In response, existing research seeks to…

计算机视觉与模式识别 · 计算机科学 2025-01-20 Tianyun Yang , Juan Cao , Danding Wang , Chang Xu

When it comes to deploying deep vision models, the behavior of these systems must be explicable to ensure confidence in their reliability and fairness. A common approach to evaluate deep learning models is to build a labeled test set with…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Jinqi Luo , Zhaoning Wang , Chen Henry Wu , Dong Huang , Fernando De la Torre

Zero-shot 6D object pose estimation involves the detection of novel objects with their 6D poses in cluttered scenes, presenting significant challenges for model generalizability. Fortunately, the recent Segment Anything Model (SAM) has…

计算机视觉与模式识别 · 计算机科学 2024-03-07 Jiehong Lin , Lihua Liu , Dekun Lu , Kui Jia

We introduce Motion2VecSets, a 4D diffusion model for dynamic surface reconstruction from point cloud sequences. While existing state-of-the-art methods have demonstrated success in reconstructing non-rigid objects using neural field…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Wei Cao , Chang Luo , Biao Zhang , Matthias Nießner , Jiapeng Tang

In zero-shot skeleton-based action recognition (ZSAR), aligning skeleton features with the text features of action labels is essential for accurately predicting unseen actions. ZSAR faces a fundamental challenge in bridging the modality gap…

计算机视觉与模式识别 · 计算机科学 2025-07-17 Jeonghyeok Do , Munchurl Kim

Trackers and video generators solve closely related problems: the former analyze motion, while the latter synthesize it. We show that this connection enables pretrained video diffusion models to perform zero-shot point tracking by simply…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Ayush Shrivastava , Sanyam Mehta , Daniel Geng , Andrew Owens

Lower extremity amputees face challenges in natural locomotion, which is partially compensated using powered assistive systems, e.g., micro-processor controlled prosthetic leg. In this paper, a radar-based perception system is proposed to…

信号处理 · 电气工程与系统科学 2021-07-09 Fady Aziz , Bassam Elmakhzangy , Christophe Maufroy , Urs Schneider , Marco F. Huber

Animation techniques bring digital 3D worlds and characters to life. However, manual animation is tedious and automated techniques are often specialized to narrow shape classes. In our work, we propose a technique for automatic re-animation…

计算机视觉与模式识别 · 计算机科学 2024-11-15 Lukas Uzolas , Elmar Eisemann , Petr Kellnhofer

Markerless estimation of 3D Kinematics has the great potential to clinically diagnose and monitor movement disorders without referrals to expensive motion capture labs; however, current approaches are limited by performing multiple…

计算机视觉与模式识别 · 计算机科学 2023-01-16 Marian Bittner , Wei-Tse Yang , Xucong Zhang , Ajay Seth , Jan van Gemert , Frans C. T. van der Helm

2D Gaussian Splatting (2DGS) is an emerging explicit scene representation method with significant potential for image compression due to high fidelity and high compression ratios. However, existing low-light enhancement algorithms operate…

计算机视觉与模式识别 · 计算机科学 2026-01-23 Yuhan Chen , Wenxuan Yu , Guofa Li , Yijun Xu , Ying Fang , Yicui Shi , Long Cao , Wenbo Chu , Keqiang Li

Recent work has explored a range of model families for human motion generation, including Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), and diffusion-based models. Despite their differences, many methods rely on…

计算机视觉与模式识别 · 计算机科学 2025-10-15 David Björkstrand , Tiesheng Wang , Lars Bretzner , Josephine Sullivan

Traditionally, monocular 3D human pose estimation employs a machine learning model to predict the most likely 3D pose for a given input image. However, a single image can be highly ambiguous and induces multiple plausible solutions for the…

计算机视觉与模式识别 · 计算机科学 2022-11-30 Karl Holmquist , Bastian Wandt

The lack of haptically aware upper-limb prostheses forces amputees to rely largely on visual cues to complete activities of daily living. In contrast, able-bodied individuals inherently rely on conscious haptic perception and automatic…

机器人学 · 计算机科学 2022-11-18 Neha Thomas , Farimah Fazlollahi , Jeremy D. Brown , Katherine J. Kuchenbecker

In emergency departments, rural hospitals, or clinics in less developed regions, clinicians often lack fast image analysis by trained radiologists, which can have a detrimental effect on patients' healthcare. Large Language Models (LLMs)…

人工智能 · 计算机科学 2024-09-11 David Bani-Harouni , Nassir Navab , Matthias Keicher

Robust in-bed human pose estimation under blanket occlusion remains challenging due to the scarcity of reliable labeled training data for heavily covered poses. Existing approaches rely on multi-modal sensing or image-to-image translation…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Navid Aslankhani Khameneh , Marco Carletti , Cigdem Beyan

Gait recognition offers a non-intrusive biometric solution by identifying individuals through their walking patterns. Although discriminative models have achieved notable success in this domain, the full potential of generative models…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Haijun Xiong , Bin Feng , Bang Wang , Xinggang Wang , Wenyu Liu

Diffusion Models have demonstrated remarkable capabilities in handling inverse problems, offering high-quality posterior-sampling-based solutions. Despite significant advances, a fundamental trade-off persists regarding the way the…

计算机视觉与模式识别 · 计算机科学 2025-11-20 Noam Elata , Hyungjin Chung , Jong Chul Ye , Tomer Michaeli , Michael Elad

Gait, as one of unique biometric features, has the advantage of being recognized from a long distance away, can be widely used in public security. Considering 3D pose estimation is more challenging than 2D pose estimation in practice , we…

计算机视觉与模式识别 · 计算机科学 2020-12-10 Na Li , Xinbo Zhao , Chong Ma

Single camera 3D pose estimation is an ill-defined problem due to inherent ambiguities from depth, occlusion or keypoint noise. Multi-hypothesis pose estimation accounts for this uncertainty by providing multiple 3D poses consistent with…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Paweł A. Pierzchlewicz , Caio O. da Silva , R. James Cotton , Fabian H. Sinz