English
Related papers

Related papers: Ask, Pose, Unite: Scaling Data Acquisition for Clo…

200 papers

Network visualization has traditionally relied on heuristic metrics, such as stress, under the assumption that optimizing them leads to aesthetic and informative layouts. However, no single metric consistently produces the most effective…

Machine Learning · Computer Science 2026-04-07 Peng Zhang , Xuefeng Li , Xiaoqi Wang , Han-Wei Shen , Yifan Hu

Accurate 6-DoF object pose estimation and tracking are critical for reliable robotic manipulation. However, zero-shot methods often fail under viewpoint-induced ambiguities and fixed-camera setups struggle when objects move or become…

Robotics · Computer Science 2026-03-10 Sheng Liu , Zhe Li , Weiheng Wang , Han Sun , Heng Zhang , Hongpeng Chen , Yusen Qin , Arash Ajoudani , Yizhao Wang

Recent advances in 3D human-aware generation have made significant progress. However, existing methods still struggle with generating novel Human Object Interaction (HOI) from text, particularly for open-set objects. We identify three main…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Jinlu Zhang , Yixin Chen , Zan Wang , Jie Yang , Yizhou Wang , Siyuan Huang

Pose estimation of the human body and hands is a fundamental problem in computer vision, and learning-based solutions require a large amount of annotated data. In this work, we improve the efficiency of the data annotation process for 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-01-19 Qi Feng , Kun He , He Wen , Cem Keskin , Yuting Ye

Previous research in human gesture recognition has largely overlooked multi-person interactions, which are crucial for understanding the social context of naturally occurring gestures. This limitation in existing datasets presents a…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Xu Cao , Pranav Virupaksha , Wenqi Jia , Bolin Lai , Fiona Ryan , Sangmin Lee , James M. Rehg

Current vision-language multimodal models are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Dewen Zhang , Wangpeng An , Hayaru Shouno

The raise of collaborative robotics has led to wide range of sensor technologies to detect human-machine interactions: at short distances, proximity sensors detect nontactile gestures virtually occlusion-free, while at medium distances,…

Computer Vision and Pattern Recognition · Computer Science 2019-10-17 Christoph Heindl , Markus Ikeda , Gernot Stübl , Andreas Pichler , Josef Scharinger

Egocentric human pose estimation (HPE) using a head-mounted device is crucial for various VR and AR applications, but it faces significant challenges due to keypoint invisibility. Nevertheless, none of the existing egocentric HPE datasets…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Peng Dai , Yu Zhang , Yiqiang Feng , Zhen Fan , Yang Zhang

Internet-scaled datasets are a luxury for human-robot interaction (HRI) researchers, as collecting natural interaction data in the wild is time-consuming and logistically challenging. The problem is exacerbated by robots' different form…

Robotics · Computer Science 2024-12-31 Fanjun Bu , Wendy Ju

In recent years, we have seen an emergence of data-driven approaches in robotics. However, most existing efforts and datasets are either in simulation or focus on a single task in isolation such as grasping, pushing or poking. In order to…

Robotics · Computer Science 2018-10-17 Pratyusha Sharma , Lekha Mohan , Lerrel Pinto , Abhinav Gupta

Analyzing individual emotions during group conversation is crucial in developing intelligent agents capable of natural human-machine interaction. While reliable emotion recognition techniques depend on different modalities (text, audio,…

The Large Vision Language Model (VLM) has recently addressed remarkable progress in bridging two fundamental modalities. VLM, trained by a sufficiently large dataset, exhibits a comprehensive understanding of both visual and linguistic to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Donggoo Kang , Dasol Jeong , Hyunmin Lee , Sangwoo Park , Hasil Park , Sunkyu Kwon , Yeongjoon Kim , Joonki Paik

Egocentric human pose estimation (HPE) using wearable sensors is essential for VR/AR applications. Most methods rely solely on either egocentric-view images or sparse Inertial Measurement Unit (IMU) signals, leading to inaccuracies due to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-07 Zhen Fan , Peng Dai , Zhuo Su , Xu Gao , Zheng Lv , Jiarui Zhang , Tianyuan Du , Guidong Wang , Yang Zhang

In this paper, we present the solution of our team HFUT-VUT for the MultiMediate Grand Challenge 2023 at ACM Multimedia 2023. The solution covers three sub-challenges: bodily behavior recognition, eye contact detection, and next speaker…

Computer Vision and Pattern Recognition · Computer Science 2023-08-04 Kun Li , Dan Guo , Guoliang Chen , Feiyang Liu , Meng Wang

In this paper, we introduce ILLUME, a unified multimodal large language model (MLLM) that seamlessly integrates multimodal understanding and generation capabilities within a single large language model through a unified next-token…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Chunwei Wang , Guansong Lu , Junwei Yang , Runhui Huang , Jianhua Han , Lu Hou , Wei Zhang , Hang Xu

Recent advancements in computer vision have seen a rise in the prominence of applications using neural networks to understand human poses. However, while accuracy has been steadily increasing on State-of-the-Art datasets, these datasets…

Computer Vision and Pattern Recognition · Computer Science 2022-06-10 Ghazal Alinezhad Noghre , Armin Danesh Pazho , Justin Sanchez , Nathan Hewitt , Christopher Neff , Hamed Tabkhi

Recent text-to-image models excel at generating high-quality object-centric images from instructions. However, images should also encapsulate rich interactions between objects, where existing models often fall short, likely due to limited…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Xinyi Gu , Jiayuan Mao

Over the last decade, Autonomous Delivery Robots (ADRs) have transformed conventional delivery methods, responding to the growing e-commerce demand. However, the readiness of ADRs to navigate safely among pedestrians in shared urban areas…

Robotics · Computer Science 2024-02-16 E. Sherafat , B. Farooq

Vision Language Models (VLMs) have demonstrated strong capabilities in understanding visual content, yet their ability to predict where humans look on user interfaces remains unexplored. We present UIGaze, a study investigating how closely…

Human-Computer Interaction · Computer Science 2026-04-30 Min Song , Yoonseong Lee , Yeonhu Seo

The analysis of the ubiquitous human-human interactions is pivotal for understanding humans as social beings. Existing human-human interaction datasets typically suffer from inaccurate body motions, lack of hand gestures and fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Liang Xu , Xintao Lv , Yichao Yan , Xin Jin , Shuwen Wu , Congsheng Xu , Yifan Liu , Yizhou Zhou , Fengyun Rao , Xingdong Sheng , Yunhui Liu , Wenjun Zeng , Xiaokang Yang