English
Related papers

Related papers: KaoLRM: Repurposing Pre-trained Large Reconstructi…

200 papers

In recent decades, 3D morphable model (3DMM) has been commonly used in image-based photorealistic 3D face reconstruction. However, face images are often corrupted by serious occlusion by non-face objects including eyeglasses, masks, and…

Computer Vision and Pattern Recognition · Computer Science 2019-09-09 Xiaowei Yuan , In Kyu Park

Reconstructing 3D objects from a single image is an intriguing but challenging problem. One promising solution is to utilize multi-view (MV) 3D reconstruction to fuse generated MV images into consistent 3D objects. However, the generated…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Yizheng Chen , Rengan Xie , Qi Ye , Sen Yang , Zixuan Xie , Tianxiao Chen , Rong Li , Yuchi Huo

We present an efficient and automatic approach for accurate reconstruction of instances of big 3D objects from multiple, unorganized and unstructured point clouds, in presence of dynamic clutter and occlusions. In contrast to conventional…

Computer Vision and Pattern Recognition · Computer Science 2017-08-18 Tolga Birdal , Slobodan Ilic

Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging implicit geometric priors from 3D reconstruction models,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jiaxin Zhang , Junjun Jiang , Haijie Li , Youyu Chen , Kui Jiang , Dave Zhenyu Chen

3D human reconstruction and animation are long-standing topics in computer graphics and vision. However, existing methods typically rely on sophisticated dense-view capture and/or time-consuming per-subject optimization procedures. To…

Graphics · Computer Science 2025-06-04 Zhiyuan Yu , Zhe Li , Hujun Bao , Can Yang , Xiaowei Zhou

The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has motivated extending these models to understand 3D scenes, aiming for human-like visual-spatial intelligence. Nevertheless, achieving deep spatial…

We propose tttLRM, a novel large 3D reconstruction model that leverages a Test-Time Training (TTT) layer to enable long-context, autoregressive 3D reconstruction with linear computational complexity, further scaling the model's capability.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Chen Wang , Hao Tan , Wang Yifan , Zhiqin Chen , Yuheng Liu , Kalyan Sunkavalli , Sai Bi , Lingjie Liu , Yiwei Hu

Large Multimodal Models (LMMs) have achieved remarkable success in vision-language tasks, yet their vast parameter counts are often underutilized during both training and inference. In this work, we embrace the idea of looping back to move…

Machine Learning · Computer Science 2026-02-11 Ruihan Xu , Yuting Gao , Lan Wang , Jianing Li , Weihao Chen , Qingpei Guo , Ming Yang , Shiliang Zhang

Human-like generalization in open-world remains a fundamental challenge for robotic manipulation. Existing learning-based methods, including reinforcement learning, imitation learning, and vision-language-action-models (VLAs), often…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Jingjing Wang , Zhengdong Hong , Chong Bao , Yuke Zhu , Junhan Sun , Guofeng Zhang

3D Morphable Model (3DMM) fitting has widely benefited face analysis due to its strong 3D priori. However, previous reconstructed 3D faces suffer from degraded visual verisimilitude due to the loss of fine-grained geometry, which is…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Xiangyu Zhu , Chang Yu , Di Huang , Zhen Lei , Hao Wang , Stan Z. Li

We introduce the Large Sparse Reconstruction Model to study how scaling transformer context windows impacts feed-forward 3D reconstruction. Although recent object-centric feed-forward methods deliver robust, high-quality reconstruction,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Zhengqin Li , Cheng Zhang , Jakob Engel , Zhao Dong

3D face reconstruction and face alignment are two fundamental and highly related topics in computer vision. Recently, some works start to use deep learning models to estimate the 3DMM coefficients to reconstruct 3D face geometry. However,…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Zihao Jian , Minshan Xie

Multi-modal Large Language Models (MLLMs) exhibit impressive capabilities in 2D tasks, yet encounter challenges in discerning the spatial positions, interrelations, and causal logic in scenes when transitioning from 2D to 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Haomiao Xiong , Yunzhi Zhuge , Jiawen Zhu , Lu Zhang , Huchuan Lu

The recovery of high-quality images from images corrupted by lens flare presents a significant challenge in low-level vision. Contemporary deep learning methods frequently entail training a lens flare removing model from scratch. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Tianwen Zhou , Qihao Duan , Zitong Yu

Complete reconstruction of surgical scenes is crucial for robot-assisted surgery (RAS). Deep depth estimation is promising but existing works struggle with depth discontinuities, resulting in noisy predictions at object boundaries and do…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Xu Wang , Shuai Zhang , Baoru Huang , Danail Stoyanov , Evangelos B. Mazomenos

While latent diffusion models (LDMs), such as Stable Diffusion, are designed for high-resolution (HR) image generation, they often struggle with significant structural distortions when generating images at resolutions higher than their…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Boyuan Cao , Jiaxin Ye , Yujie Wei , Hongming Shan

3D Morphable Models (3DMMs) enable controllable facial geometry and expression editing for reconstruction, animation, and AR/VR, but traditional PCA-based mesh models are limited in resolution, detail, and photorealism. Neural volumetric…

3D face reconstruction from a single 2D image is a very important topic in computer vision. However, the current reconstruction methods are usually non-sensitive to face identities and over-sensitive to facial poses, which may result in…

Computer Vision and Pattern Recognition · Computer Science 2019-05-17 Yao Luo , Xiaoguang Tu , Mei Xie

Recovering dense 3D geometry from unposed images remains a foundational challenge in computer vision. Current state-of-the-art models are predominantly trained on perspective datasets, which implicitly constrains them to a standard pinhole…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Namitha Guruprasad , Abhay Yadav , Cheng Peng , Rama Chellappa

Recent advances in diffusion-based Large Restoration Models (LRMs) have significantly improved photo-realistic image restoration by leveraging the internal knowledge embedded within model weights. However, existing LRMs often suffer from…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Hang Guo , Tao Dai , Zhihao Ouyang , Taolin Zhang , Yaohua Zha , Bin Chen , Shu-tao Xia