English
Related papers

Related papers: PGformer: Proxy-Bridged Game Transformer for Multi…

200 papers

Natural human interactions for Mixed Reality Applications are overwhelmingly multimodal: humans communicate intent and instructions via a combination of visual, aural and gestural cues. However, supporting low-latency and accurate…

Human-Computer Interaction · Computer Science 2020-12-21 Darshana Rathnayake , Ashen de Silva , Dasun Puwakdandawa , Lakmal Meegahapola , Archan Misra , Indika Perera

Transformers are powerful visual learners, in large part due to their conspicuous lack of manually-specified priors. This flexibility can be problematic in tasks that involve multiple-view geometry, due to the near-infinite possible…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Yash Bhalgat , Joao F. Henriques , Andrew Zisserman

Human pose estimation based on Channel State Information (CSI) has emerged as a promising approach for non-intrusive and precise human activity monitoring, yet faces challenges including accurate multi-person pose recognition and effective…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Yanyi Qu , Haoyang Ma , Wenhui Xiong

3D human motion forecasting aims to enable autonomous applications. Estimating uncertainty for each prediction (i.e., confidence based on probability density or quantile) is essential for safety-critical contexts like human-robot…

Robotics · Computer Science 2025-07-22 Yue Ma , Kanglei Zhou , Fuyang Yu , Frederick W. B. Li , Xiaohui Liang

Motion in-betweening, a fundamental task in character animation, consists of generating motion sequences that plausibly interpolate user-provided keyframe constraints. It has long been recognized as a labor-intensive and challenging…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Setareh Cohan , Guy Tevet , Daniele Reda , Xue Bin Peng , Michiel van de Panne

Aggregating multi-modality data to obtain reliable data representation attracts more and more attention. Recent studies demonstrate that Transformer models usually work well for multi-modality tasks. Existing Transformers generally either…

Computer Vision and Pattern Recognition · Computer Science 2023-03-17 Xixi Wang , Xiao Wang , Bo Jiang , Jin Tang , Bin Luo

Social interaction is a fundamental aspect of human behavior and communication. The way individuals position themselves in relation to others, also known as proxemics, conveys social cues and affects the dynamics of social interaction.…

Computer Vision and Pattern Recognition · Computer Science 2023-12-14 Lea Müller , Vickie Ye , Georgios Pavlakos , Michael Black , Angjoo Kanazawa

Real-time recognition and prediction of surgical activities are fundamental to advancing safety and autonomy in robot-assisted surgery. This paper presents a multimodal transformer architecture for real-time recognition and prediction of…

Robotics · Computer Science 2024-10-27 Keshara Weerasinghe , Seyed Hamid Reza Roodabeh , Kay Hutchinson , Homa Alemzadeh

Forecasting the future states of surrounding traffic participants is a crucial capability for autonomous vehicles. The recently proposed occupancy flow field prediction introduces a scalable and effective representation to jointly predict…

Computer Vision and Pattern Recognition · Computer Science 2023-07-07 Haochen Liu , Zhiyu Huang , Chen Lv

Interactive decision-making is essential in applications such as autonomous driving, where the agent must infer the behavior of nearby human drivers while planning in real-time. Traditional predict-then-act frameworks are often insufficient…

Predicting the future motion of dynamic agents is of paramount importance to ensuring safety and assessing risks in motion planning for autonomous robots. In this study, we propose a two-stage motion prediction method, called R-Pred,…

Computer Vision and Pattern Recognition · Computer Science 2023-07-17 Sehwan Choi , Jungho Kim , Junyong Yun , Jun Won Choi

Gaze object prediction aims to predict the location and category of the object that is watched by a human. Previous gaze object prediction works use CNN-based object detectors to predict the object's location. However, we find that…

Computer Vision and Pattern Recognition · Computer Science 2024-02-22 Binglu Wang , Chenxi Guo , Yang Jin , Haisheng Xia , Nian Liu

Many visual scenes contain text that carries crucial information, and it is thus essential to understand text in images for downstream reasoning tasks. For example, a deep water label on a warning sign warns people about the danger in the…

Computer Vision and Pattern Recognition · Computer Science 2020-03-26 Ronghang Hu , Amanpreet Singh , Trevor Darrell , Marcus Rohrbach

Motion prediction is a classic problem in computer vision, which aims at forecasting future motion given the observed pose sequence. Various deep learning models have been proposed, achieving state-of-the-art performance on motion…

Computer Vision and Pattern Recognition · Computer Science 2022-01-10 Pengxiang Su , Zhenguang Liu , Shuang Wu , Lei Zhu , Yifang Yin , Xuanjing Shen

Prior work on human motion forecasting has mostly focused on predicting the future motion of single subjects in isolation from their past pose sequence. In the presence of closely interacting people, however, this strategy fails to account…

Computer Vision and Pattern Recognition · Computer Science 2022-01-14 Isinsu Katircioglu , Costa Georgantas , Mathieu Salzmann , Pascal Fua

We address the problem of accurate capture and expressive modelling of interactive behaviors happening between two persons in daily scenarios. Different from previous works which either only consider one person or focus on conversational…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Yinghao Huang , Leo Ho , Dafei Qin , Mingyi Shi , Taku Komura

Effective pair programming depends on coordination of attention, cognitive effort, and joint regulation over time, yet most adaptive learning systems remain individual-centric and reactive. This paper introduces ProPACT, a proactive…

Human-Computer Interaction · Computer Science 2026-05-05 Anahita Golrang , Kshitij Sharma , olga viberg

Recently, occluded person re-identification(Re-ID) remains a challenging task that people are frequently obscured by other people or obstacles, especially in a crowd massing situation. In this paper, we propose a self-supervised deep…

Computer Vision and Pattern Recognition · Computer Science 2022-02-11 Mi Zhou , Hongye Liu , Zhekun Lv , Wei Hong , Xiai Chen

Transformer models rely on self-attention to capture token dependencies but face challenges in effectively integrating positional information while allowing multi-head attention (MHA) flexibility. Prior methods often model semantic and…

Machine Learning · Computer Science 2025-05-28 Jintian Shao , Hongyi Huang , Jiayi Wu , Beiwen Zhang , ZhiYu Wu , You Shan , MingKai Zheng

Object handover is a common form of interaction that is widely present in collaborative tasks. However, achieving it efficiently remains a challenge. We address the problem of ensuring resilient robotic actions that can adapt to complex…

Robotics · Computer Science 2026-04-30 Omar Faris , Sławomir Tadeja , Fulvio Forni