English
Related papers

Related papers: Ask, Pose, Unite: Scaling Data Acquisition for Clo…

200 papers

In this paper, we present the Intra- and Inter-Human Relation Networks (I^2R-Net) for Multi-Person Pose Estimation. It involves two basic modules. First, the Intra-Human Relation Module operates on a single person and aims to capture…

Computer Vision and Pattern Recognition · Computer Science 2022-06-28 Yiwei Ding , Wenjin Deng , Yinglin Zheng , Pengfei Liu , Meihong Wang , Xuan Cheng , Jianmin Bao , Dong Chen , Ming Zeng

Human Pose Estimation (HPE) involves detecting and localizing keypoints on the human body from visual data. In 3D HPE, occlusions, where parts of the body are not visible in the image, pose a significant challenge for accurate pose…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Filipa Lino , Carlos Santiago , Manuel Marques

Multiview diffusion models have shown considerable success in image-to-3D generation for general objects. However, when applied to human data, existing methods have yet to deliver promising results, largely due to the challenges of scaling…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Yuhan Wang , Fangzhou Hong , Shuai Yang , Liming Jiang , Wayne Wu , Chen Change Loy

Social intelligence, the ability to interpret emotions, intentions, and behaviors, is essential for effective communication and adaptive responses. As robots and AI systems become more prevalent in caregiving, healthcare, and education, the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Erika Mori , Yue Qiu , Hirokatsu Kataoka , Yoshimitsu Aoki

Expressive human pose and shape estimation (EHPS) unifies body, hands, and face motion capture with numerous applications. Despite encouraging progress, current state-of-the-art methods still depend largely on a confined set of training…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Zhongang Cai , Wanqi Yin , Ailing Zeng , Chen Wei , Qingping Sun , Yanjun Wang , Hui En Pang , Haiyi Mei , Mingyuan Zhang , Lei Zhang , Chen Change Loy , Lei Yang , Ziwei Liu

Natural language plays a critical role in many computer vision applications, such as image captioning, visual question answering, and cross-modal retrieval, to provide fine-grained semantic information. Unfortunately, while human pose is…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Ginger Delmas , Philippe Weinzaepfel , Thomas Lucas , Francesc Moreno-Noguer , Grégory Rogez

3D pose estimation from sparse multi-views is a critical task for numerous applications, including action recognition, sports analysis, and human-robot interaction. Optimization-based methods typically follow a two-stage pipeline, first…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Tony Danjun Wang , Tolga Birdal , Nassir Navab , Lennart Bastian

In human-robot interaction (HRI), the beginning of an interaction is often complex. Whether the robot should communicate with the human is dependent on several situational factors (e.g., the current human's activity, urgency of the…

Human-Computer Interaction · Computer Science 2025-03-21 Kazuhiro Sasabuchi , Naoki Wake , Atsushi Kanehira , Jun Takamatsu , Katsushi Ikeuchi

Photorealistic avatars of human faces have come a long way in recent years, yet research along this area is limited by a lack of publicly available, high-quality datasets covering both, dense multi-view camera captures, and rich facial…

Large Language Models (LLMs) are known to have limited extrapolation ability beyond their pre-trained context window, constraining their application in downstream tasks with lengthy inputs. Recent studies have sought to extend LLMs' context…

Computation and Language · Computer Science 2024-01-17 Yikai Zhang , Junlong Li , Pengfei Liu

Following on recent advances in large language models (LLMs) and subsequent chat models, a new wave of large vision-language models (LVLMs) has emerged. Such models can incorporate images as input in addition to text, and perform tasks such…

Computers and Society · Computer Science 2024-02-09 Kathleen C. Fraser , Svetlana Kiritchenko

Computer vision (CV) has achieved great success in interpreting semantic meanings from images, yet CV algorithms can be brittle for tasks with adverse vision conditions and the ones suffering from data/label pair limitation. One of this…

Computer Vision and Pattern Recognition · Computer Science 2020-08-21 Shuangjun Liu , Xiaofei Huang , Nihang Fu , Cheng Li , Zhongnan Su , Sarah Ostadabbas

Generating accurate and concise textual summaries from multimodal documents is challenging, especially when dealing with visually complex content like scientific posters. We introduce PosterSum, a novel benchmark to advance the development…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Rohit Saxena , Pasquale Minervini , Frank Keller

Visual Question Answering (VQA) in remote sensing (RS) is pivotal for interpreting Earth observation data. However, existing RS VQA datasets are constrained by limitations in annotation richness, question diversity, and the assessment of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xing Zi , Jinghao Xiao , Yunxiao Shi , Xian Tao , Jun Li , Ali Braytee , Mukesh Prasad

The availability of a large labeled dataset is a key requirement for applying deep learning methods to solve various computer vision tasks. In the context of understanding human activities, existing public datasets, while large in size, are…

Computer Vision and Pattern Recognition · Computer Science 2023-05-18 Yizhak Ben-Shabat , Xin Yu , Fatemeh Sadat Saleh , Dylan Campbell , Cristian Rodriguez-Opazo , Hongdong Li , Stephen Gould

Multi-modal large language models (MLLMs) have demonstrated remarkable vision-language capabilities, primarily due to the exceptional in-context understanding and multi-task learning strengths of large language models (LLMs). The advent of…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Jianing Li , Xi Nan , Ming Lu , Li Du , Shanghang Zhang

The difficulty and expense of obtaining large-scale human responses make Large Language Models (LLMs) an attractive alternative and a promising proxy for human behavior. However, prior work shows that LLMs often produce homogeneous outputs…

Artificial Intelligence · Computer Science 2025-10-09 Manh Hung Nguyen , Sebastian Tschiatschek , Adish Singla

Facial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stakes environment. In…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Xinqi Fan , Jingting Li , John See , Moi Hoon Yap , Wen-Huang Cheng , Xiaobai Li , Xiaopeng Hong , Su-Jing Wang , Adrian K. Davision

Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. To develop socially intelligent AI technologies, it is crucial to develop models that can…

Recent multimodal large language models (MLLMs) such as GPT-4o and Qwen3-Omni show strong perception but struggle in multi-speaker, dialogue-centric settings that demand agentic reasoning tracking who speaks, maintaining roles, and…