English
Related papers

Related papers: TokenHSI: Unified Synthesis of Physical Human-Scen…

200 papers

Recently, the vision transformer and its variants have played an increasingly important role in both monocular and multi-view human pose estimation. Considering image patches as tokens, transformers can model the global dependencies within…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 Haoyu Ma , Zhe Wang , Yifei Chen , Deying Kong , Liangjian Chen , Xingwei Liu , Xiangyi Yan , Hao Tang , Xiaohui Xie

Recognizing interactive action plays an important role in human-robot interaction and collaboration. Previous methods use late fusion and co-attention mechanism to capture interactive relations, which have limited learning capability or…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Yuhang Wen , Zixuan Tang , Yunsheng Pang , Beichen Ding , Mengyuan Liu

Human-machine interfaces (HMI) play a pivotal role in the rehabilitation and daily assistance of lower-limb amputees. The brain of such interfaces is a control model that detects the user's intention using sensor input and generates…

Computational Engineering, Finance, and Science · Computer Science 2021-10-08 Sharmita Dey , Takashi Yoshida , Robert H. Foerster , Michael Ernst , Thomas Schmalz , Rodrigo M. Carnier , Arndt F. Schilling

Successfully addressing a wide variety of tasks is a core ability of autonomous agents, requiring flexibly adapting the underlying decision-making strategies and, as we argue in this work, also adapting the perception modules. An analogical…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Pierre Marza , Laetitia Matignon , Olivier Simonin , Christian Wolf

Recent advancements in multimodal large language models (MLLMs) have demonstrated remarkable capabilities in processing diverse data types, yet significant disparities persist between human cognitive processes and computational approaches…

Computation and Language · Computer Science 2025-05-09 Dongxing Yu

Computational models of how users perceive and act within a virtual or physical environment offer enormous potential for the understanding and design of user interactions. Cognition models have been used to understand the role of attention…

Wearable accelerometers enable large-scale health monitoring, yet learning robust human-activity representations has been constrained by scarce labeled data. While self-supervised learning offers a remedy, existing methods treat sensor…

Machine Learning · Computer Science 2026-05-28 Prithviraj Tarale , Kiet Chu , Abhishek Varghese , Kai-Chun Liu , Maxwell A. Xu , Mohit Iyyer , Sunghoon I. Lee

Skeleton-based human action recognition has received widespread attention in recent years due to its diverse range of application scenarios. Due to the different sources of human skeletons, skeleton data naturally exhibit heterogeneity. The…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Hongsong Wang , Xiaoyan Ma , Jidong Kuang , Jie Gui

Existing GUI agent models relying on coordinate-based one-step visual grounding struggle with generalizing to varying input resolutions and aspect ratios. Alternatives introduce coordinate-free strategies yet suffer from learning under…

Machine Learning · Computer Science 2026-02-04 Xiaoce Wang , Guibin Zhang , Junzhe Li , Jinzhe Tu , Chun Li , Ming Li

Long-horizon manipulation has been a long-standing challenge in the robotics community. We propose ReinforceGen, a system that combines task decomposition, data generation, imitation learning, and motion planning to form an initial…

Robotics · Computer Science 2025-12-19 Zihan Zhou , Animesh Garg , Ajay Mandlekar , Caelan Garrett

Long-range Human-Robot Interaction (HRI) remains underexplored. Within it, Command Source Identification (CSI) - determining who issued a command - is especially challenging due to multi-user and distance-induced sensor ambiguity. We…

Human-Computer Interaction · Computer Science 2026-03-26 Chengwen Zhang , Chun Yu , Borong Zhuang , Haopeng Jin , Qingyang Wan , Zhuojun Li , Zhe He , Zhoutong Ye , Yu Mei , Chang Liu , Weinan Shi , Yuanchun Shi

Scene graph generation (SGG) and human-object interaction (HOI) detection are two important visual tasks aiming at localising and recognising relationships between objects, and interactions between humans and objects, respectively.…

Computer Vision and Pattern Recognition · Computer Science 2023-11-06 Tao He , Lianli Gao , Jingkuan Song , Yuan-Fang Li

Enabling humanoid robots to perform agile and adaptive interactive tasks has long been a core challenge in robotics. Current approaches are bottlenecked by either the scarcity of realistic interaction data or the need for meticulous,…

Robotics · Computer Science 2026-02-03 Yinhuai Wang , Qihan Zhao , Yuen Fui Lau , Runyi Yu , Hok Wai Tsui , Qifeng Chen , Jingbo Wang , Jiangmiao Pang , Ping Tan

Imitation learning is a promising approach for training humanoid robots to both walk and manipulate, but it requires a large number of demonstrations, which are time-intensive and difficult to collect via teleoperation. Existing…

We tackle the challenges of synthesizing versatile, physically simulated human motions for full-body object manipulation. Unlike prior methods that are focused on detailed motion tracking, trajectory following, or teleoperation, our…

Robotics · Computer Science 2025-12-12 Chen Tessler , Yifeng Jiang , Erwin Coumans , Zhengyi Luo , Gal Chechik , Xue Bin Peng

Remote telepresence via next-generation mixed reality platforms can provide higher levels of immersion for computer-mediated communications, allowing participants to engage in a wide spectrum of activities, previously not possible in 2D…

Human-Computer Interaction · Computer Science 2022-04-04 Mohammad Keshavarzi , Michael Zollhoefer , Allen Y. Yang , Patrick Peluse , Luisa Caldas

This paper presents a novel human-robot interaction (HRI) framework that enables intuitive gesture-driven control through a capacitance-based woven tactile skin. Unlike conventional interfaces that rely on panels or handheld devices, the…

We introduce Toyteller, an AI-powered storytelling system where users generate a mix of story text and visuals by directly manipulating character symbols like they are toy-playing. Anthropomorphized symbol motions can convey rich and…

Human-Computer Interaction · Computer Science 2025-01-24 John Joon Young Chung , Melissa Roemmele , Max Kreminski

Generating and representing human behavior are of major importance for various computer vision applications. Commonly, human video synthesis represents behavior as sequences of postures while directly predicting their likely progressions or…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Andreas Blattmann , Timo Milbich , Michael Dorkenwald , Björn Ommer

In this work we propose a multi-task spatio-temporal network, called SUSiNet, that can jointly tackle the spatio-temporal problems of saliency estimation, action recognition and video summarization. Our approach employs a single network…

Computer Vision and Pattern Recognition · Computer Science 2019-04-16 Petros Koutras , Petros Maragos