English
Related papers

Related papers: Ask, Pose, Unite: Scaling Data Acquisition for Clo…

200 papers

Language is often used to describe physical interaction, yet most 3D human pose estimation methods overlook this rich source of information. We bridge this gap by leveraging large multimodal models (LMMs) as priors for reconstructing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Sanjay Subramanian , Evonne Ng , Lea Müller , Dan Klein , Shiry Ginosar , Trevor Darrell

Vision Language Models (VLMs) have achieved impressive performance on spatial reasoning benchmarks, yet these evaluations mask critical weaknesses in understanding object interactions. Current benchmarks test high level relationships ('left…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Vineet Bhat , Sungsu Kim , Valts Blukis , Greg Heinrich , Prashanth Krishnamurthy , Ramesh Karri , Stan Birchfield , Farshad Khorrami , Jonathan Tremblay

Current vision-language models (VLMs) are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Dewen Zhang , Tahir Hussain , Wangpeng An , Hayaru Shouno

The ability to anticipate human-object interactions is highly desirable in an intelligent assistive system in order to guide users during daily life activities and understand their short and long-term goals. Creating systems with such…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Daniele Materia , Francesco Ragusa , Giovanni Maria Farinella

Human motion generation has shown great advances thanks to the recent diffusion models trained on large-scale motion capture data. Most of existing works, however, currently target animation of isolated people in empty scenes. Meanwhile,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Yangsong Zhang , Abdul Ahad Butt , Gül Varol , Ivan Laptev

Human parsing and pose estimation have recently received considerable interest due to their substantial application potentials. However, the existing datasets have limited numbers of images and annotations and lack a variety of human…

Computer Vision and Pattern Recognition · Computer Science 2018-04-09 Xiaodan Liang , Ke Gong , Xiaohui Shen , Liang Lin

Bootstrapping from pre-trained language models has been proven to be an efficient approach for building vision-language models (VLM) for tasks such as image captioning or visual question answering. However, outputs of these models rarely…

Machine Learning · Computer Science 2023-06-01 Manuel Brack , Patrick Schramowski , Björn Deiseroth , Kristian Kersting

Parse graphs boost human pose estimation (HPE) by integrating context and hierarchies, yet prior work mostly focuses on single modality modeling, ignoring the potential of multimodal fusion. Notably, language offers rich HPE priors like…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Shibang Liu , Xuemei Xie , Guangming Shi

Multimodal Vision-Language Models (VLMs) enable powerful applications from their fused understanding of images and language, but many perform poorly on UI tasks due to the lack of UI training data. In this paper, we adapt a recipe for…

Human-Computer Interaction · Computer Science 2023-10-10 Yue Jiang , Eldon Schoop , Amanda Swearngin , Jeffrey Nichols

Obtaining data in the medical field is challenging, making the adoption of AI technology within the space slow and high-risk. We evaluate whether we can overcome this obstacle with synthetic data generated by large language models (LLMs).…

Machine Learning · Computer Science 2024-09-04 Paulo Soares , Sean McCurdy , Andrew J. Gerber , Peter Fonagy

We introduce ChatPose, a framework employing Large Language Models (LLMs) to understand and reason about 3D human poses from images or textual descriptions. Our work is motivated by the human ability to intuitively understand postures from…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Yao Feng , Jing Lin , Sai Kumar Dwivedi , Yu Sun , Priyanka Patel , Michael J. Black

Human-centric visual understanding is an important desideratum for effective human-robot interaction. In order to navigate crowded public places, social robots must be able to interpret the activity of the surrounding humans. This paper…

Computer Vision and Pattern Recognition · Computer Science 2023-07-28 Shengnan Hu , Ce Zheng , Zixiang Zhou , Chen Chen , Gita Sukthankar

Modelling interactions between humans and objects in natural environments is central to many applications including gaming, virtual and mixed reality, as well as human behavior analysis and human-robot collaboration. This challenging…

Computer Vision and Pattern Recognition · Computer Science 2022-04-15 Bharat Lal Bhatnagar , Xianghui Xie , Ilya A. Petrov , Cristian Sminchisescu , Christian Theobalt , Gerard Pons-Moll

The data scarcity problem is a crucial factor that hampers the model performance of IMU-based human motion capture. However, effective data augmentation for IMU-based motion capture is challenging, since it has to capture the physical…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Zhuojun Li , Chun Yu , Chen Liang , Yuanchun Shi

Natural Language Explanation (NLE) aims to elucidate the decision-making process by providing detailed, human-friendly explanations in natural language. It helps demystify the decision-making processes of large vision-language models…

Computation and Language · Computer Science 2024-12-10 Patrick Amadeus Irawan , Genta Indra Winata , Samuel Cahyawijaya , Ayu Purwarianti

Due to visual ambiguities and inter-person occlusions, existing human pose estimation methods cannot recover plausible close interactions from in-the-wild videos. Even state-of-the-art large foundation models~(\eg, SAM) cannot accurately…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Buzhen Huang , Chen Li , Chongyang Xu , Dongyue Lu , Jinnan Chen , Yangang Wang , Gim Hee Lee

Cross-view person matching and 3D human pose estimation in multi-camera networks are particularly difficult when the cameras are extrinsically uncalibrated. Existing efforts generally require large amounts of 3D data for training neural…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Yan Xu , Kris Kitani

Vision-language models (VLMs) work well in tasks ranging from image captioning to visual question answering (VQA), yet they struggle with spatial reasoning, a key skill for understanding our physical world that humans excel at. We find that…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Michael Ogezi , Freda Shi

Humans live within a 3D space and constantly interact with it to perform tasks. Such interactions involve physical contact between surfaces that is semantically meaningful. Our goal is to learn how humans interact with scenes and leverage…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Mohamed Hassan , Partha Ghosh , Joachim Tesch , Dimitrios Tzionas , Michael J. Black

Training a Multimodal Large Language Model (MLLM) from scratch, like GPT-4, is resource-intensive. Regarding Large Language Models (LLMs) as the core processor for multimodal information, our paper introduces LMEye, a human-like eye with a…

Computer Vision and Pattern Recognition · Computer Science 2023-09-29 Yunxin Li , Baotian Hu , Xinyu Chen , Lin Ma , Yong Xu , Min Zhang
‹ Prev 1 2 3 10 Next ›