English
Related papers

Related papers: TORE: Token Reduction for Efficient Human Mesh Rec…

200 papers

We present THUNDR, a transformer-based deep neural network methodology to reconstruct the 3d pose and shape of people, given monocular RGB images. Key to our methodology is an intermediate 3d marker representation, where we aim to combine…

Computer Vision and Pattern Recognition · Computer Science 2021-06-18 Mihai Zanfir , Andrei Zanfir , Eduard Gabriel Bazavan , William T. Freeman , Rahul Sukthankar , Cristian Sminchisescu

Human imitation has become topical recently, driven by GAN's ability to disentangle human pose and body content. However, the latest methods hardly focus on 3D information, and to avoid self-occlusion, a massive amount of input images are…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Haoxi Ran , Guangfu Wang , Li Lu

Recent advances in 3D point cloud transformers have led to state-of-the-art results in tasks such as semantic segmentation and reconstruction. However, these models typically rely on dense token representations, incurring high computational…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Tuan Anh Tran , Duy M. H. Nguyen , Hoai-Chau Tran , Michael Barz , Khoa D. Doan , Roger Wattenhofer , Ngo Anh Vien , Mathias Niepert , Daniel Sonntag , Paul Swoboda

We present a novel method, called NeTO, for capturing 3D geometry of solid transparent objects from 2D images via volume rendering. Reconstructing transparent objects is a very challenging task, which is ill-suited for general-purpose…

Computer Vision and Pattern Recognition · Computer Science 2023-09-11 Zongcheng Li , Xiaoxiao Long , Yusen Wang , Tuo Cao , Wenping Wang , Fei Luo , Chunxia Xiao

Recently, model-based retrieval has emerged as a new paradigm in text retrieval that discards the index in the traditional retrieval model and instead memorizes the candidate corpora using model parameters. This design employs a…

Information Retrieval · Computer Science 2023-05-19 Ruiyang Ren , Wayne Xin Zhao , Jing Liu , Hua Wu , Ji-Rong Wen , Haifeng Wang

3D Human Body Reconstruction from a monocular image is an important problem in computer vision with applications in virtual and augmented reality platforms, animation industry, en-commerce domain, etc. While several of the existing works…

Computer Vision and Pattern Recognition · Computer Science 2019-08-20 Abbhinav Venkat , Chaitanya Patel , Yudhik Agrawal , Avinash Sharma

Multimodal Large Language Models have demonstrated remarkable capabilities in video understanding, yet face prohibitive computational costs and performance degradation from ''context rot'' due to massive visual token redundancy. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Shida Wang , YongXiang Hua , Zhou Tao , Haoyu Cao , Linli Xu

We present an approach that can reconstruct hands in 3D from monocular input. Our approach for Hand Mesh Recovery, HaMeR, follows a fully transformer-based architecture and can analyze hands with significantly increased accuracy and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-11 Georgios Pavlakos , Dandan Shan , Ilija Radosavovic , Angjoo Kanazawa , David Fouhey , Jitendra Malik

Token compression is crucial for mitigating the quadratic complexity of self-attention mechanisms in Vision Transformers (ViTs), which often involve numerous input tokens. Existing methods, such as ToMe, rely on GPU-inefficient operations…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Simin Huo , Ning Li

We present a novel method for reconstructing clothed humans from a sparse set of, e.g., 1 to 6 RGB images. Despite impressive results from recent works employing deep implicit representation, we revisit the volumetric approach and…

Computer Vision and Pattern Recognition · Computer Science 2023-07-26 Sicong Tang , Guangyuan Wang , Qing Ran , Lingzhi Li , Li Shen , Ping Tan

Reconstructing hand-held objects in 3D from monocular images remains a significant challenge in computer vision. Most existing approaches rely on implicit 3D representations, which produce overly smooth reconstructions and are…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Zerui Chen , Rolandos Alexandros Potamias , Shizhe Chen , Cordelia Schmid

We address the problem of regressing 3D human pose and shape from a single image, with a focus on 3D accuracy. The current best methods leverage large datasets of 3D pseudo-ground-truth (p-GT) and 2D keypoints, leading to robust…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Sai Kumar Dwivedi , Yu Sun , Priyanka Patel , Yao Feng , Michael J. Black

Recently, vision transformers have shown great success in a set of human reconstruction tasks such as 2D human pose estimation (2D HPE), 3D human pose estimation (3D HPE), and human mesh reconstruction (HMR) tasks. In these tasks, feature…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Ce Zheng , Matias Mendieta , Taojiannan Yang , Guo-Jun Qi , Chen Chen

We propose a Transformer-based framework for 3D human texture estimation from a single image. The proposed Transformer is able to effectively exploit the global information of the input image, overcoming the limitations of existing methods…

Computer Vision and Pattern Recognition · Computer Science 2021-09-07 Xiangyu Xu , Chen Change Loy

In autonomous driving applications, accurate and efficient road surface reconstruction is paramount. This paper introduces RoMe, a novel framework designed for the robust reconstruction of large-scale road surfaces. Leveraging a unique mesh…

Computer Vision and Pattern Recognition · Computer Science 2024-06-24 Ruohong Mei , Wei Sui , Jiaxin Zhang , Xue Qin , Gang Wang , Tao Peng , Cong Yang

Increasing the throughput of the Transformer architecture, a foundational component used in numerous state-of-the-art models for vision and language tasks (e.g., GPT, LLaVa), is an important problem in machine learning. One recent and…

Recent advancements in both transformer-based methods and spiral neighbor sampling techniques have greatly enhanced hand mesh reconstruction. Transformers excel in capturing complex vertex relationships, and spiral neighbor sampling is…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Huilong Xie , Wenwei Song , Wenxiong Kang , Yihong Lin

Existing deep learning-based human mesh reconstruction approaches have a tendency to build larger networks in order to achieve higher accuracy. Computational complexity and model size are often neglected, despite being key characteristics…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Ce Zheng , Matias Mendieta , Pu Wang , Aidong Lu , Chen Chen

Incremental scene reconstruction is essential to the navigation in robotics. Most of the conventional methods typically make use of either TSDF (truncated signed distance functions) volume or neural networks to implicitly represent the…

Robotics · Computer Science 2024-04-30 Shaofan Liu , Junbo Chen , Jianke Zhu

Tomography has had an important impact on the physical, biological, and medical sciences. To date, most tomographic applications have been focused on 3D scalar reconstructions. However, in some crucial applications, vector tomography is…

Medical Physics · Physics 2023-10-20 Minh Pham , Xingyuan Lu , Arjun Rana , Stanley Osher , Jianwei Miao