English
Related papers

Related papers: HAP: Structure-Aware Masked Image Modeling for Hum…

200 papers

Deep image prior (DIP) is a recently proposed technique for solving imaging inverse problems by fitting the reconstructed images to the output of an untrained convolutional neural network. Unlike pretrained feedforward neural networks, the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 Kevin Zhang , Mingyang Xie , Maharshi Gor , Yi-Ting Chen , Yvonne Zhou , Christopher A. Metzler

Multimodal pretraining is effective for building general-purpose representations, but in many practical deployments, only one modality is heavily used during downstream fine-tuning. Standard pretraining strategies treat all modalities…

Machine Learning · Computer Science 2026-01-30 Atik Faysal , Mohammad Rostami , Reihaneh Gh. Roshan , Nikhil Muralidhar , Huaxia Wang

We present FITE, a First-Implicit-Then-Explicit framework for modeling human avatars in clothing. Our framework first learns implicit surface templates representing the coarse clothing topology, and then employs the templates to guide the…

Computer Vision and Pattern Recognition · Computer Science 2022-07-15 Siyou Lin , Hongwen Zhang , Zerong Zheng , Ruizhi Shao , Yebin Liu

Human parsing is a key topic in image processing with many applications, such as surveillance analysis, human-robot interaction, person search, and clothing category classification, among many others. Recently, due to the success of deep…

Computer Vision and Pattern Recognition · Computer Science 2023-01-31 Xiaomei Zhang , Xiangyu Zhu , Ming Tang , Zhen Lei

Human capabilities in understanding visual relations are far superior to those of AI systems, especially for previously unseen objects. For example, while AI systems struggle to determine whether two such objects are visually the same or…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Oleh Kolner , Thomas Ortner , Stanisław Woźniak , Angeliki Pantazi

State-of-the-art approaches toward image restoration can be classified into model-based and learning-based. The former - best represented by sparse coding techniques - strive to exploit intrinsic prior knowledge about the unknown…

Image and Video Processing · Electrical Eng. & Systems 2018-11-29 Fangfang Wu , Weisheng Dong , Guangming Shi , Xin Li

Brain structural MRI has been widely used to assess the future progression of cognitive impairment (CI). Previous learning-based studies usually suffer from the issue of small-sized labeled training data, while there exist a huge amount of…

Image and Video Processing · Electrical Eng. & Systems 2023-06-27 Lintao Zhang , Jinjian Wu , Lihong Wang , Li Wang , David C. Steffens , Shijun Qiu , Guy G. Potter , Mingxia Liu

This paper introduces a new architecture for human pose estimation using a multi- layer convolutional network architecture and a modified learning technique that learns low-level features and higher-level weak spatial models. Unconstrained…

Computer Vision and Pattern Recognition · Computer Science 2014-04-24 Arjun Jain , Jonathan Tompson , Mykhaylo Andriluka , Graham W. Taylor , Christoph Bregler

Many vision and language models suffer from poor visual grounding - often falling back on easy-to-learn language priors rather than basing their decisions on visual concepts in the image. In this work, we propose a generic approach called…

Computer Vision and Pattern Recognition · Computer Science 2019-10-29 Ramprasaath R. Selvaraju , Stefan Lee , Yilin Shen , Hongxia Jin , Shalini Ghosh , Larry Heck , Dhruv Batra , Devi Parikh

Parsing human poses in images is fundamental in extracting critical visual information for artificial intelligent agents. Our goal is to learn self-contained body part representations from images, which we call visual symbols, and their…

Computer Vision and Pattern Recognition · Computer Science 2013-04-24 Fang Wang , Yi Li

Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for a wide range of applications. In this paper, we continually pre-train prevailing VFMs in a multimodal manner such that they can effortlessly process…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Yitong Chen , Lingchen Meng , Wujian Peng , Zuxuan Wu , Yu-Gang Jiang

Reconstructing multi-human body mesh from a single monocular image is an important but challenging computer vision problem. In addition to the individual body mesh models, we need to estimate relative 3D positions among subjects to generate…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Chenyan Wu , Yandong Li , Xianfeng Tang , James Wang

We introduce Gaussian masking for Language-Image Pre-Training (GLIP) a novel, straightforward, and effective technique for masking image patches during pre-training of a vision-language model. GLIP builds on Fast Language-Image Pre-Training…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Mingliang Liang , Martha Larson

In this paper, we propose a structured feature learning framework to reason the correlations among body joints at the feature level in human pose estimation. Different from existing approaches of modelling structures on score maps or…

Computer Vision and Pattern Recognition · Computer Science 2016-03-31 Xiao Chu , Wanli Ouyang , Hongsheng Li , Xiaogang Wang

Recently, human pose estimation mainly focuses on how to design a more effective and better deep network structure as human features extractor, and most designed feature extraction networks only introduce the position of each anatomical…

Computer Vision and Pattern Recognition · Computer Science 2022-12-06 Zhangjian Ji , Zilong Wang , Ming Zhang , Yapeng Chen , Yuhua Qian

Decoding visual images from brain activity has significant potential for advancing brain-computer interaction and enhancing the understanding of human perception. Recent approaches align the representation spaces of images and brain…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Nona Rajabi , Antônio H. Ribeiro , Miguel Vasco , Farzaneh Taleb , Mårten Björkman , Danica Kragic

From an image of a person, we can easily infer the natural 3D pose and shape of the person even if ambiguity exists. This is because we have a mental model that allows us to imagine a person's appearance at different viewing directions from…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Hanbyel Cho , Yooshin Cho , Jaesung Ahn , Junmo Kim

Despite the noticeable progress in perceptual tasks like detection, instance segmentation and human parsing, computers still perform unsatisfactorily on visually understanding humans in crowded scenes, such as group behavior analysis,…

Computer Vision and Pattern Recognition · Computer Science 2018-07-09 Jian Zhao , Jianshu Li , Yu Cheng , Li Zhou , Terence Sim , Shuicheng Yan , Jiashi Feng

We present Sapiens, a family of models for four fundamental human-centric vision tasks -- 2D pose estimation, body-part segmentation, depth estimation, and surface normal prediction. Our models natively support 1K high-resolution inference…

Computer Vision and Pattern Recognition · Computer Science 2024-08-28 Rawal Khirodkar , Timur Bagautdinov , Julieta Martinez , Su Zhaoen , Austin James , Peter Selednik , Stuart Anderson , Shunsuke Saito

Masked image modeling (MIM) has emerged as a promising approach for pre-training Vision Transformers (ViTs). MIMs predict masked tokens token-wise to recover target signals that are tokenized from images or generated by pre-trained models…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Taekyung Kim , Byeongho Heo , Dongyoon Han