中文
相关论文

相关论文: RemoCap: Disentangled Representation Learning for …

200 篇论文

Vision-language models like CLIP excel at recognizing the single, prominent object in a scene. However, they struggle in complex scenes containing multiple objects. We identify a fundamental reason for this limitation: VLM feature space…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Samyak Rawlekar , Yujun Cai , Yiwei Wang , Ming-Hsuan Yang , Narendra Ahuja

Recent work has shown that object-centric representations can greatly help improve the accuracy of learning dynamics while also bringing interpretability. In this work, we take this idea one step further, ask the following question: "can…

计算机视觉与模式识别 · 计算机科学 2024-07-04 Sanket Gandhi , Atul , Samanyu Mahajan , Vishal Sharma , Rushil Gupta , Arnab Kumar Mondal , Parag Singla

Reconstructing 3D human heads in low-view settings presents technical challenges, mainly due to the pronounced risk of overfitting with limited views and high-frequency signals. To address this, we propose geometry decomposition and adopt a…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Baixin Xu , Jiarui Zhang , Kwan-Yee Lin , Chen Qian , Ying He

A key challenge in self-supervised video representation learning is how to effectively capture motion information besides context bias. While most existing works implicitly achieve this with video-specific pretext tasks (e.g., predicting…

计算机视觉与模式识别 · 计算机科学 2021-04-05 Lianghua Huang , Yu Liu , Bin Wang , Pan Pan , Yinghui Xu , Rong Jin

Dynamic scene understanding is an essential capability in robotics and VR/AR. In this paper we propose Co-Section, an optimization-based approach to 3D dynamic scene reconstruction, which infers hidden shape information from intersection…

计算机视觉与模式识别 · 计算机科学 2021-12-13 Michael Strecke , Joerg Stueckler

We present the first method for real-time full body capture that estimates shape and motion of body and hands together with a dynamic 3D face model from a single color image. Our approach uses a new neural network architecture that exploits…

计算机视觉与模式识别 · 计算机科学 2021-04-16 Yuxiao Zhou , Marc Habermann , Ikhsanul Habibie , Ayush Tewari , Christian Theobalt , Feng Xu

In this paper, we present a novel strategy to design disentangled 3D face shape representation. Specifically, a given 3D face shape is decomposed into identity part and expression part, which are both encoded and decoded in a nonlinear way.…

计算机视觉与模式识别 · 计算机科学 2019-03-05 Zi-Hang Jiang , Qianyi Wu , Keyu Chen , Juyong Zhang

Monocular 3D object detection (M3OD) is intrinsically ill-posed, hence training a high-performance deep learning based M3OD model requires a humongous amount of labeled data with complicated visual variation from diverse scenes, variety of…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Zhaonian Kuang , Rui Ding , Meng Yang , Xinhu Zheng , Gang Hua

We address the challenging problem of dense dynamic scene reconstruction and camera pose estimation from multiple freely moving cameras -- a setting that arises naturally when multiple observers capture a shared event. Prior approaches…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Shuo Sun , Unal Artan , Malcolm Mielle , Achim J. Lilienthaland , Martin Magnusson

Human de-occlusion, which aims to infer the appearance of invisible human parts from an occluded image, has great value in many human-related tasks, such as person re-id, and intention inference. To address this task, this paper proposes a…

计算机视觉与模式识别 · 计算机科学 2024-02-08 Guoqiang Liang , Jiahao Hu , Qingyue Wang , Shizhou Zhang

Existing human Motion Capture (MoCap) methods mostly focus on the visual similarity while neglecting the physical plausibility. As a result, downstream tasks such as driving virtual human in 3D scene or humanoid robots in real world suffer…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Shenghao Ren , Yi Lu , Jiayi Huang , Jiayi Zhao , He Zhang , Tao Yu , Qiu Shen , Xun Cao

Multimodal representations that enable cross-modal retrieval are widely used. However, these often lack interpretability making it difficult to explain the retrieved results. Solutions such as learning sparse disentangled representations…

信息检索 · 计算机科学 2025-06-25 Prachi J , Sumit Bhatia , Srikanta Bedathur

Marker-less 3D human motion capture from a single colour camera has seen significant progress. However, it is a very challenging and severely ill-posed problem. In consequence, even the most accurate state-of-the-art approaches have…

计算机视觉与模式识别 · 计算机科学 2020-12-10 Soshi Shimada , Vladislav Golyanik , Weipeng Xu , Christian Theobalt

Capturing the interactions between humans and their environment in 3D is important for many applications in robotics, graphics, and vision. Recent works to reconstruct the 3D human and object from a single RGB image do not have consistent…

计算机视觉与模式识别 · 计算机科学 2023-11-01 Xianghui Xie , Bharat Lal Bhatnagar , Gerard Pons-Moll

Tracking body and hand motions in the 3D space is essential for social and self-presence in augmented and virtual environments. Unlike the popular 3D pose estimation setting, the problem is often formulated as inside-out tracking based on…

计算机视觉与模式识别 · 计算机科学 2020-12-08 Mathias Parger , Chengcheng Tang , Yuanlu Xu , Christopher Twigg , Lingling Tao , Yijing Li , Robert Wang , Markus Steinberger

We propose a new methodology to estimate the 3D displacement field of deformable objects from video sequences using standard monocular cameras. We solve in real time the complete (possibly visco-)hyperelasticity problem to properly describe…

计算机视觉与模式识别 · 计算机科学 2021-10-14 Alberto Badias , Iciar Alfaro , David Gonzalez , Francisco Chinesta , Elias Cueto

Disentangled representation learning offers useful properties such as dimension reduction and interpretability, which are essential to modern deep learning approaches. Although deep learning techniques have been widely applied to…

机器学习 · 计算机科学 2022-04-11 Sichen Zhao , Wei Shao , Jeffrey Chan , Flora D. Salim

Markerless human motion capture (mocap) from multiple RGB cameras is a widely studied problem. Existing methods either need calibrated cameras or calibrate them relative to a static camera, which acts as the reference frame for the mocap…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Nitin Saini , Chun-hao P. Huang , Michael J. Black , Aamir Ahmad

A long-standing challenge in scene analysis is the recovery of scene arrangements under moderate to heavy occlusion, directly from monocular video. While the problem remains a subject of active research, concurrent advances have been made…

图形学 · 计算机科学 2019-07-19 Aron Monszpart , Paul Guerrero , Duygu Ceylan , Ersin Yumer , Niloy J. Mitra

Reconstructing complete and animatable 3D human avatars from monocular videos remains challenging, particularly under severe occlusions. While 3D Gaussian Splatting has enabled photorealistic human rendering, existing methods struggle with…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Jinlong Fan , Shanshan Zhao , Liang Zheng , Jing Zhang , Yuxiang Yang , Mingming Gong