English
Related papers

Related papers: MonoArt: Progressive Structural Reasoning for Mono…

200 papers

Articulated 3D reconstruction has valuable applications in various domains, yet it remains costly and demands intensive work from domain experts. Recent advancements in template-free learning methods show promising results with monocular…

Computer Vision and Pattern Recognition · Computer Science 2023-12-11 Tao Tu , Ming-Feng Li , Chieh Hubert Lin , Yen-Chi Cheng , Min Sun , Ming-Hsuan Yang

In many robotic applications, especially for the autonomous driving, understanding the semantic information and the geometric structure of surroundings are both essential. Semantic 3D maps, as a carrier of the environmental knowledge, are…

Computer Vision and Pattern Recognition · Computer Science 2019-07-25 Yucai Bai , Lei Fan , Ziyu Pan , Long Chen

Creating interactive digital environments for gaming, robotics, and simulation relies on articulated 3D objects whose functionality emerges from their part geometry and kinematic structure. However, existing approaches remain fundamentally…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Penghao Wang , Siyuan Xie , Hongyu Yan , Xianghui Yang , Jingwei Huang , Chunchao Guo , Jiayuan Gu

3D Reconstruction of moving articulated objects without additional information about object structure is a challenging problem. Current methods overcome such challenges by employing category-specific skeletal models. Consequently, they do…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Hao Zhang , Fang Li , Samyak Rawlekar , Narendra Ahuja

Inspired by the recent success of methods that employ shape priors to achieve robust 3D reconstructions, we propose a novel recurrent neural network architecture that we call the 3D Recurrent Reconstruction Neural Network (3D-R2N2). The…

Computer Vision and Pattern Recognition · Computer Science 2016-04-05 Christopher B. Choy , Danfei Xu , JunYoung Gwak , Kevin Chen , Silvio Savarese

Existing inverse physics methods recover physical parameters from multi-view videos, where geometric constraints across views resolve scale and 3D structure. In monocular settings, however, such constraints are absent, leading to severe…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Daniel Rho , Jun Myeong Choi , Matthew Thornton , Biswadip Dey , Roni Sengupta

Articulated objects, as prevalent entities in human life, their 3D representations play crucial roles across various applications. However, achieving both high-fidelity textured surface reconstruction and dynamic generation for articulated…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Di Wu , Liu Liu , Zhou Linli , Anran Huang , Liangtu Song , Qiaojun Yu , Qi Wu , Cewu Lu

Articulated objects (e.g., doors and drawers) exist everywhere in our life. Different from rigid objects, articulated objects have higher degrees of freedom and are rich in geometries, semantics, and part functions. Modeling different kinds…

Computer Vision and Pattern Recognition · Computer Science 2023-11-22 Yushi Du , Ruihai Wu , Yan Shen , Hao Dong

Fine-grained visual reasoning in multimodal large language models (MLLMs) is bottlenecked by single-pass global image encoding: key evidence often lies in tiny objects, cluttered regions, subtle markings, or dense charts. We present…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Hao Ding , Zhichuan Yang , Weijie Ge , Ziqin Gao , Chaoyi Lu , Lei Zhao

In recent years, Neural Radiance Fields (NeRF) have achieved remarkable progress in dynamic human reconstruction and rendering. Part-based rendering paradigms, guided by human segmentation, allow for flexible parameter allocation based on…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Yao Lu , Jiawei Li , Ming Jiang

Understanding the 3D motion of articulated objects is essential in robotic scene understanding, mobile manipulation, and motion planning. Prior methods for articulation estimation have primarily focused on controlled settings, assuming…

Monocular and stereo depth estimation offer complementary strengths: monocular methods capture rich contextual priors but lack geometric precision, while stereo approaches leverage epipolar geometry yet struggle with ambiguities such as…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Tongfan Guan , Jiaxin Guo , Chen Wang , Yun-Hui Liu

We propose SelfRecon, a clothed human body reconstruction method that combines implicit and explicit representations to recover space-time coherent geometries from a monocular self-rotating human video. Explicit methods require a predefined…

Computer Vision and Pattern Recognition · Computer Science 2022-04-06 Boyi Jiang , Yang Hong , Hujun Bao , Juyong Zhang

Spatial reasoning (SR), the ability to infer 3D spatial information from 2D inputs, is essential for real-world applications such as embodied AI and autonomous driving. However, existing research primarily focuses on indoor environments and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Qirui Wang , Jingyi He , Yining Pan , Si Yong Yeo , Xulei Yang , Shijie Li

Mobile monocular 3D object detection (Mono3D) (e.g., on a vehicle, a drone, or a robot) is an important yet challenging task. Existing transformer-based offline Mono3D models adopt grid-based vision tokens, which is suboptimal when using…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Yunsong Zhou , Hongzi Zhu , Quan Liu , Shan Chang , Minyi Guo

Monocular depth estimation, enabled by self-supervised learning, is a key technique for 3D perception in computer vision. However, it faces significant challenges in real-world scenarios, which encompass adverse weather variations, motion…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Runze Chen , Haiyong Luo , Fang Zhao , Jingze Yu , Yupeng Jia , Juan Wang , Xuepeng Ma

Monocular 3D object detection is well-known to be a challenging vision task due to the loss of depth information; attempts to recover depth using separate image-only approaches lead to unstable and noisy depth estimates, harming 3D…

Computer Vision and Pattern Recognition · Computer Science 2019-05-15 Ivan Barabanau , Alexey Artemov , Evgeny Burnaev , Vyacheslav Murashkin

Three-dimensional (3D) ultrasound (US) aims to provide sonographers with the spatial relationships of anatomical structures, playing a crucial role in clinical diagnosis. Recently, deep-learning-based freehand 3D US has made significant…

Image and Video Processing · Electrical Eng. & Systems 2025-06-23 Mingyuan Luo , Xin Yang , Zhongnuo Yan , Yan Cao , Yuanji Zhang , Xindi Hu , Jin Wang , Haoxuan Ding , Wei Han , Litao Sun , Dong Ni

We present ArtMesh, a mesh-native method for reconstructing articulated objects explicitly as connected triangle meshes with per-part rigid motion from multi-view images in start and end states. Existing 3D Gaussian Splatting pipelines for…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Sylvia Yuan , Dan Wang , Ravi Ramamoorthi , Xinrui Cui

While DETR-like architectures have demonstrated significant potential for monocular 3D object detection, they are often hindered by a critical limitation: the exclusion of 3D attributes from the bipartite matching process. This exclusion…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Kiet Dang Vu , Trung Thai Tran , Kien Nguyen Do Trung , Duc Dung Nguyen