English
Related papers

Related papers: Ask, Pose, Unite: Scaling Data Acquisition for Clo…

200 papers

Vision-Language Models (VLMs) are trained on vast amounts of data captured by humans emulating our understanding of the world. However, known as visual illusions, human's perception of reality isn't always faithful to the physical world.…

Artificial Intelligence · Computer Science 2023-11-02 Yichi Zhang , Jiayi Pan , Yuchen Zhou , Rui Pan , Joyce Chai

We propose a novel representation of virtual humans for highly realistic real-time animation and rendering in 3D applications. We learn pose dependent appearance and geometry from highly accurate dynamic mesh sequences obtained from…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Wieland Morgenstern , Milena T. Bagdasarian , Anna Hilsmann , Peter Eisert

Large Vision-Language Models (LVLMs) have shown promising capabilities in understanding and generating information by integrating both visual and textual data. However, current models are still prone to hallucinations, which degrade the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Robert Wijaya , Ngoc-Bao Nguyen , Ngai-Man Cheung

Human Pose Estimation (HPE) is one of the fundamental problems in computer vision. It has applications ranging from virtual reality, human behavior analysis, video surveillance, anomaly detection, self-driving to medical assistance. The…

Computer Vision and Pattern Recognition · Computer Science 2021-12-23 Milan Kresović , Thong Duy Nguyen

Recent advancements in Computer Assisted Diagnosis have shown promising performance in medical imaging tasks, particularly in chest X-ray analysis. However, the interaction between these models and radiologists has been primarily limited to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Yunsoo Kim , Jinge Wu , Yusuf Abdulle , Yue Gao , Honghan Wu

Speech Emotion Recognition models typically use single categorical labels, overlooking the inherent ambiguity of human emotions. Ambiguous Emotion Recognition addresses this by representing emotions as probability distributions, but…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-22 Wenda Zhang , Hongyu Jin , Siyi Wang , Zhiqiang Wei , Ting Dang

Human-object interaction (HOI) detection aims to comprehend the intricate relationships between humans and objects, predicting $<human, action, object>$ triplets, and serving as the foundation for numerous computer vision tasks. The…

Computer Vision and Pattern Recognition · Computer Science 2023-11-08 Yichao Cao , Qingfei Tang , Xiu Su , Chen Song , Shan You , Xiaobo Lu , Chang Xu

We introduce MERGE, a system for situational grounding of actors, objects, and events in dynamic human-robot group interactions. Effective collaboration in such settings requires consistent situational awareness, built on persistent…

In this paper, we concern on the bottom-up paradigm in multi-person pose estimation (MPPE). Most previous bottom-up methods try to consider the relation of instances to identify different body parts during the post processing, while…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Ruoqi Yin , Jianqin Yin

Comprehensive evaluation of Multimodal Large Language Models (MLLMs) has recently garnered widespread attention in the research community. However, we observe that existing benchmarks present several common barriers that make it difficult…

Computer Vision and Pattern Recognition · Computer Science 2025-02-06 Yi-Fan Zhang , Huanyu Zhang , Haochen Tian , Chaoyou Fu , Shuangqing Zhang , Junfei Wu , Feng Li , Kun Wang , Qingsong Wen , Zhang Zhang , Liang Wang , Rong Jin , Tieniu Tan

Human mesh recovery can be approached using either regression-based or optimization-based methods. Regression models achieve high pose accuracy but struggle with model-to-image alignment due to the lack of explicit 2D-3D correspondences. In…

Computer Vision and Pattern Recognition · Computer Science 2025-02-07 Chongyang Xu , Buzhen Huang , Chengfang Zhang , Ziliang Feng , Yangang Wang

Understanding group-level social interactions in public spaces is crucial for urban planning, informing the design of socially vibrant and inclusive environments. Detecting such interactions from images involves interpreting subtle visual…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Liu Liu , Alexandra Kudaeva , Marco Cipriano , Fatimeh Al Ghannam , Freya Tan , Gerard de Melo , Andres Sevtsuk

Recognition of Handwritten Mathematical Expressions (HMEs) is a challenging problem because of the ambiguity and complexity of two-dimensional handwriting. Moreover, the lack of large training data is a serious issue, especially for…

Computer Vision and Pattern Recognition · Computer Science 2019-01-23 Anh Duc Le , Bipin Indurkhya , Masaki Nakagawa

A well-designed strong-weak augmentation strategy and the stable teacher to generate reliable pseudo labels are essential in the teacher-student framework of semi-supervised learning (SSL). Considering these in mind, to suit the…

Computer Vision and Pattern Recognition · Computer Science 2022-03-16 JongMok Kim , Hwijun Lee , Jaeseung Lim , Jongkeun Na , Nojun Kwak , Jin Young Choi

Generating realistic and physically plausible 3D Human-Object Interactions (HOI) remains a key challenge in motion generation. One primary reason is that describing these physical constraints with words alone is difficult. To address this…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Songjin Cai , Linjie Zhong , Ling Guo , Changxing Ding

Text-guided human pose editing has gained significant traction in AIGC applications. However,it remains plagued by structural anomalies and generative artifacts. Existing evaluation metrics often isolate authenticity detection from quality…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Ningyu Sun , Zhaolin Cai , Zitong Xu , Peihang Chen , Huiyu Duan , Yichao Yan , Xiongkuo Min , Xiaokang Yang

Humans constantly interact with objects in daily life tasks. Capturing such processes and subsequently conducting visual inferences from a fixed viewpoint suffers from occlusions, shape and texture ambiguities, motions, etc. To mitigate the…

Computer Vision and Pattern Recognition · Computer Science 2022-12-16 Juze Zhang , Haimin Luo , Hongdi Yang , Xinru Xu , Qianyang Wu , Ye Shi , Jingyi Yu , Lan Xu , Jingya Wang

Multimodal Large Language Models (MLLMs) struggle with accurately capturing camera-object relations, especially for object orientation, camera viewpoint, and camera shots. This stems from the fact that existing MLLMs are trained on images…

Text-to-image diffusion has attracted vast attention due to its impressive image-generation capabilities. However, when it comes to human-centric text-to-image generation, particularly in the context of faces and hands, the results often…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Jie Zhu , Yixiong Chen , Mingyu Ding , Ping Luo , Leye Wang , Jingdong Wang

Accurately estimating hand pose and hand-object contact events is essential for robot data-collection, immersive virtual environments, and biomechanical analysis, yet remains challenging due to visual occlusion, subtle contact cues,…

Human-Computer Interaction · Computer Science 2025-08-25 Yuemin Mao , Uksang Yoo , Yunchao Yao , Shahram Najam Syed , Luca Bondi , Jonathan Francis , Jean Oh , Jeffrey Ichnowski