English
Related papers

Related papers: SkeleGuide: Explicit Skeleton Reasoning for Contex…

200 papers

Multimodal pretraining has revolutionized visual understanding, but its impact on video-based person re-identification (ReID) remains underexplored. Existing approaches often rely on video-text pairs, yet suffer from two fundamental…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Rifen Lin , Alex Jinpeng Wang , Jiawei Mo , Min Li

Structure-guided image completion aims to inpaint a local region of an image according to an input guidance map from users. While such a task enables many practical applications for interactive editing, existing methods often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Haitian Zheng , Zhe Lin , Jingwan Lu , Scott Cohen , Eli Shechtman , Connelly Barnes , Jianming Zhang , Qing Liu , Yuqian Zhou , Sohrab Amirghodsi , Jiebo Luo

Multimodal reasoning is a process of understanding, integrating and inferring information across different data modalities. It has recently attracted surging academic attention as a benchmark for Artificial Intelligence (AI). Although there…

Computation and Language · Computer Science 2025-09-16 Fenghua Cheng , Jinxiang Wang , Sen Wang , Zi Huang , Xue Li

Pedestrian intention prediction is crucial for autonomous driving. In particular, knowing if pedestrians are going to cross in front of the ego-vehicle is core to performing safe and comfortable maneuvers. Creating accurate and fast models…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Muhammad Naveed Riaz , Maciej Wielgosz , Abel Garcia Romera , Antonio M. Lopez

Object skeletons offer a concise representation of structural information, capturing essential aspects of posture and orientation that are crucial for autonomous driving applications. However, a unified architecture that simultaneously…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Yasamin Borhani , Taylor Mordan , Yihan Wang , Reyhaneh Hosseininejad , Javad Khoramdel , Alexandre Alahi

Recent advances in tracking sensors and pose estimation software enable smart systems to use trajectories of skeleton joint locations for supervised learning. We study the problem of accurately recognizing sign language words, which is key…

Computer Vision and Pattern Recognition · Computer Science 2022-02-04 Joachim Gudmundsson , Martin P. Seybold , John Pfeifer

We present 3DHumanGAN, a 3D-aware generative adversarial network that synthesizes photorealistic images of full-body humans with consistent appearances under different view-angles and body-poses. To tackle the representational and…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Zhuoqian Yang , Shikai Li , Wayne Wu , Bo Dai

Can we directly visualize what we imagine in our brain together with what we describe? The inherent nature of human perception reveals that, when we think, our body can combine language description and build a vivid picture in our brain.…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Ling Wang , Chen Wu , Lin Wang

This study investigates identity-preserving image synthesis, an intriguing task in image generation that seeks to maintain a subject's identity while adding a personalized, stylistic touch. Traditional methods, such as Textual Inversion and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-07 Yuxuan Yan , Chi Zhang , Rui Wang , Yichao Zhou , Gege Zhang , Pei Cheng , Gang Yu , Bin Fu

Our work addresses the problem of egocentric human pose estimation from downwards-facing cameras on head-mounted devices (HMD). This presents a challenging scenario, as parts of the body often fall outside of the image or are occluded.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-29 Hanz Cuevas-Velasquez , Charlie Hewitt , Sadegh Aliakbarian , Tadas Baltrušaitis

We present a new approach for synthesizing novel views of people in new poses. Our novel differentiable renderer enables the synthesis of highly realistic images from any viewpoint. Rather than operating over mesh-based structures, our…

Computer Vision and Pattern Recognition · Computer Science 2023-03-29 Guillaume Rochette , Chris Russell , Richard Bowden

Detecting object skeletons in natural images presents challenging, due to varied object scales, the complexity of backgrounds and various noises. The skeleton is a highly compressing shape representation, which can bring some essential…

Computer Vision and Pattern Recognition · Computer Science 2020-12-08 Xiuxiu Bai , Lele Ye , Zhe Liu

Action recognition from well-segmented 3D skeleton video has been intensively studied. However, due to the difficulty in representing the 3D skeleton video and the lack of training data, action detection from streaming 3D skeleton video…

Computer Vision and Pattern Recognition · Computer Science 2017-04-20 Bo Li , Huahui Chen , Yucheng Chen , Yuchao Dai , Mingyi He

In the realm of skeleton-based action recognition, the traditional methods which rely on coarse body keypoints fall short of capturing subtle human actions. In this work, we propose Expressive Keypoints that incorporates hand and foot…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Yijie Yang , Jinlu Zhang , Jiaxu Zhang , Zhigang Tu

3D human pose estimation from sketches has broad applications in computer animation and film production. Unlike traditional human pose estimation, this task presents unique challenges due to the abstract and disproportionate nature of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Li Wang , Yiyu Zhuang , Yanwen Wang , Xun Cao , Chuan Guo , Xinxin Zuo , Hao Zhu

Generative modeling of anatomical structures plays a crucial role in virtual imaging trials, which allow researchers to perform studies without the costs and constraints inherent to in vivo and phantom studies. For clinical relevance,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Bram de Wilde , Max T. Rietberg , Guillaume Lajoinie , Jelmer M. Wolterink

While head-mounted devices are becoming more compact, they provide egocentric views with significant self-occlusions of the device user. Hence, existing methods often fail to accurately estimate complex 3D poses from egocentric views. In…

Computer Vision and Pattern Recognition · Computer Science 2024-05-16 Hiroyasu Akada , Jian Wang , Vladislav Golyanik , Christian Theobalt

Multi-image spatial reasoning remains challenging for current multimodal large language models (MLLMs). While single-view perception is inherently 2D, reasoning over multiple views requires building a coherent scene understanding across…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Xuejun Zhang , Aditi Tiwari , Zhenhailong Wang , Heng Ji

Human action recognition (HAR) has achieved impressive results with deep learning models, but their decision-making process remains opaque due to their black-box nature. Ensuring interpretability is crucial, especially for real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Jongseo Lee , Wooil Lee , Gyeong-Moon Park , Seong Tae Kim , Jinwoo Choi

We propose a decoupled 3D scene generation framework called SceneMaker in this work. Due to the lack of sufficient open-set de-occlusion and pose estimation priors, existing methods struggle to simultaneously produce high-quality geometry…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Yukai Shi , Weiyu Li , Zihao Wang , Hongyang Li , Xingyu Chen , Ping Tan , Lei Zhang
‹ Prev 1 8 9 10 Next ›