English
Related papers

Related papers: TORE: Token Reduction for Efficient Human Mesh Rec…

200 papers

Good pre-trained visual representations could enable robots to learn visuomotor policy efficiently. Still, existing representations take a one-size-fits-all-tasks approach that comes with two important drawbacks: (1) Being completely…

Robotics · Computer Science 2024-11-05 Jianing Qian , Yunshuang Li , Bernadette Bucher , Dinesh Jayaraman

Single-image human mesh recovery provides a compact 3D, person-centric representation that supports analysis, animation, AR and VR, rehabilitation, and human-computer interaction. However, prevailing systems impose an intact-limb prior and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Jiaying Ying , Heming Du , Kaihao Zhang , Sean M. Tweedy , Xin Yu

Image tokenization has enabled major advances in autoregressive image generation by providing compressed, discrete representations that are more efficient to process than raw pixels. While traditional approaches use 2D grid tokenization,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Roman Bachmann , Jesse Allardice , David Mizrahi , Enrico Fini , Oğuzhan Fatih Kar , Elmira Amirloo , Alaaeldin El-Nouby , Amir Zamir , Afshin Dehghan

Model-based learned iterative reconstruction methods have recently been shown to outperform classical reconstruction algorithms. Applicability of these methods to large scale inverse problems is however limited by the available memory for…

Image and Video Processing · Electrical Eng. & Systems 2020-04-21 Andreas Hauptmann , Jonas Adler , Simon Arridge , Ozan Öktem

Sparse-view Computed Tomography (CT) reconstructs images from a limited number of X-ray projections to reduce radiation and scanning time, which makes reconstruction an ill-posed inverse problem. Deep learning methods achieve high-fidelity…

Image and Video Processing · Electrical Eng. & Systems 2025-12-16 Aujasvit Datta , Jiayun Wang , Asad Aali , Armeet Singh Jatyani , Anima Anandkumar

We present a novel method to learn temporally consistent 3D reconstruction of clothed people from a monocular video. Recent methods for 3D human reconstruction from monocular video using volumetric, implicit or parametric human shape…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Akin Caliskan , Armin Mustafa , Adrian Hilton

Recent end-to-end automatic speech recognition (ASR) systems often utilize a Transformer-based acoustic encoder that generates embedding at a high frame rate. However, this design is inefficient, particularly for long speech signals due to…

Computation and Language · Computer Science 2023-06-29 Yuang Li , Yu Wu , Jinyu Li , Shujie Liu

Table structure recognition (TSR) aims at extracting tables in images into machine-understandable formats. Recent methods solve this problem by predicting the adjacency relations of detected cell boxes or learning to directly generate the…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Rujiao Long , Hangdi Xing , Zhibo Yang , Qi Zheng , Zhi Yu , Cong Yao , Fei Huang

Annotating lots of 3D medical images for training segmentation models is time-consuming. The goal of weakly supervised semantic segmentation is to train segmentation models without using any ground truth segmentation masks. Our work…

Image and Video Processing · Electrical Eng. & Systems 2024-04-23 Marius Schmidt-Mengin , Alexis Benichoux , Shibeshih Belachew , Nikos Komodakis , Nikos Paragios

Speeding up the data acquisition is one of the central aims to advance tomographic imaging. On the one hand, this reduces motion artifacts due to undesired movements, and on the other hand this decreases the examination time for the…

Numerical Analysis · Mathematics 2015-01-20 Michael Sandbichler , Felix Krahmer , Thomas Berer , Peter Burgholzer , Markus Haltmeier

Since its inception, Vision Transformer (ViT) has emerged as a prevalent model in the computer vision domain. Nonetheless, the multi-head self-attention (MHSA) mechanism in ViT is computationally expensive due to its calculation of…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Zhe Bian , Zhe Wang , Wenqiang Han , Kangping Wang

We propose Token Turing Machines (TTM), a sequential, autoregressive Transformer model with memory for real-world sequential visual understanding. Our model is inspired by the seminal Neural Turing Machine, and has an external memory…

The misaligned human texture across different human parts is one of the main limitations of existing 3D human reconstruction methods. Each human part, such as a jacket or pants, should maintain a distinct texture without blending into…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Hyeongjin Nam , Donghwan Kim , Gyeongsik Moon , Kyoung Mu Lee

Three-dimensional medical image segmentation is a fundamental yet computationally demanding task due to the cubic growth of voxel processing and the redundant computation on homogeneous regions. To address these limitations, we propose…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Sen Zeng , Hong Zhou , Zheng Zhu , Yang Liu

Achieving high-quality High Dynamic Range (HDR) imaging on resource-constrained edge devices is a critical challenge in computer vision, as its performance directly impacts downstream tasks such as intelligent surveillance and autonomous…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Yu-Shen Huang , Tzu-Han Chen , Cheng-Yen Hsiao , Shaou-Gang Miaou

In 3D human shape and pose estimation from a monocular video, models trained with limited labeled data cannot generalize well to videos with occlusion, which is common in the wild videos. The recent human neural rendering approaches…

Computer Vision and Pattern Recognition · Computer Science 2023-09-22 Yu Cheng , Bo Wang , Robby T. Tan

This paper presents a simple yet powerful method for 3D human mesh reconstruction from a single RGB image. Most recently, the non-local interactions of the whole mesh vertices have been effectively estimated in the transformer while the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-20 Jeonghwan Kim , Mi-Gyeong Gwon , Hyunwoo Park , Hyukmin Kwon , Gi-Mun Um , Wonjun Kim

Human Mesh Recovery (HMR) is fundamentally ambiguous: under occlusion or weak depth cues, multiple 3D bodies can explain the same image evidence. This ambiguity is not uniform across the body, as torso pose and root structure are often…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Patrick Kwon , Chen Chen

In this paper, we aim at the problem of tensor data completion. Tensor-train decomposition is adopted because of its powerful representation ability and linear scalability to tensor order. We propose an algorithm named Sparse Tensor-train…

Numerical Analysis · Computer Science 2018-03-23 Longhao Yuan , Qibin Zhao , Jianting Cao

Token filtering to reduce irrelevant tokens prior to self-attention is a straightforward way to enable efficient vision Transformer. This is the first work to view token filtering from a feature selection perspective, where we weigh the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-25 Hong Wang , Su Yang , Xiaoke Huang , Weishan Zhang
‹ Prev 1 8 9 10 Next ›