English
Related papers

Related papers: A Modular Multimodal Architecture for Gaze Target …

200 papers

Turn-taking prediction is crucial for seamless interactions. This study introduces a novel, lightweight framework for accurate turn-taking prediction in triadic conversations without relying on computationally intensive methods. Unlike…

Human-Computer Interaction · Computer Science 2025-05-30 Seongsil Heo , Calvin Murdock , Michael Proulx , Christi Miller

Gaze target detection (GTD) is the task of predicting where a person in an image is looking. This is a challenging task, as it requires the ability to understand the relationship between the person's head, body, and eyes, as well as the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Athul M. Mathew , Arshad Ali Khan , Thariq Khalid , Faroq AL-Tam , Riad Souissi

Understanding the human-object interactions (HOIs) from a video is essential to fully comprehend a visual scene. This line of research has been addressed by detecting HOIs from images and lately from videos. However, the video-based HOI…

Computer Vision and Pattern Recognition · Computer Science 2023-06-07 Zhifan Ni , Esteve Valls Mascaró , Hyemin Ahn , Dongheui Lee

Effective assisted living environments must be able to perform inferences on how their occupants interact with one another as well as with surrounding objects. To accomplish this goal using a vision-based automated approach, multiple tasks…

Computer Vision and Pattern Recognition · Computer Science 2019-09-23 Philipe A. Dias , Damiano Malafronte , Henry Medeiros , Francesca Odone

In autonomous driving and robotics, there is a growing interest in utilizing short-term historical data to enhance multi-camera 3D object detection, leveraging the continuous and correlated nature of input video streams. Recent work has…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Seokha Moon , Hongbeen Park , Jungphil Kwon , Jaekoo Lee , Jinkyu Kim

Human affordance learning investigates contextually relevant novel pose prediction such that the estimated pose represents a valid human action within the scene. While the task is fundamental to machine perception and automated interactive…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Prasun Roy , Saumik Bhattacharya , Subhankar Ghosh , Umapada Pal , Michael Blumenstein

We present GazeMotion, a novel method for human motion forecasting that combines information on past human poses with human eye gaze. Inspired by evidence from behavioural sciences showing that human eye and body movements are closely…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Zhiming Hu , Syn Schmitt , Daniel Haeufle , Andreas Bulling

Close human-robot cooperation is a key enabler for new developments in advanced manufacturing and assistive applications. Close cooperation require robots that can predict human actions and intent, and understand human non-verbal cues.…

Human-Computer Interaction · Computer Science 2019-02-19 Paul Schydlo , Mirko Rakovic , Lorenzo Jamone , José Santos-Victor

Predicting pedestrian behavior is a crucial task for intelligent driving systems. Accurate predictions require a deep understanding of various contextual elements that potentially impact the way pedestrians behave. To address this…

Computer Vision and Pattern Recognition · Computer Science 2022-10-17 Amir Rasouli , Iuliia Kotseruba

Training a multimodal network is challenging and it requires complex architectures to achieve reasonable performance. We show that one reason for this phenomena is the difference between the convergence rate of various modalities. We…

Artificial Intelligence · Computer Science 2020-11-13 Aya Abdelsalam Ismail , Mahmudul Hasan , Faisal Ishtiaq

By borrowing the wisdom of human in gaze following, we propose a two-stage solution for gaze point prediction of the target persons in a scene. Specifically, in the first stage, both head image and its position are fed into a gaze direction…

Computer Vision and Pattern Recognition · Computer Science 2019-07-05 Dongze Lian , Zehao Yu , Shenghua Gao

Trajectory prediction is crucial for autonomous vehicles. The planning system not only needs to know the current state of the surrounding objects but also their possible states in the future. As for vehicles, their trajectories are…

Robotics · Computer Science 2020-07-07 Chenxu Luo , Lin Sun , Dariush Dabiri , Alan Yuille

The potential of multimodal generative artificial intelligence (mAI) to replicate human grounded language understanding, including the pragmatic, context-rich aspects of communication, remains to be clarified. Humans are known to use…

We address the problem of detecting attention targets in video. Our goal is to identify where each person in each frame of a video is looking, and correctly handle the case where the gaze target is out-of-frame. Our novel architecture…

Computer Vision and Pattern Recognition · Computer Science 2020-04-01 Eunji Chong , Yongxin Wang , Nataniel Ruiz , James M. Rehg

Joint visual attention is characterized by two or more individuals looking at a common target at the same time. The ability to identify joint attention in scenes, the people involved, and their common target, is fundamental to the…

Computer Vision and Pattern Recognition · Computer Science 2018-04-13 Daniel Harari , Joshua B. Tenenbaum , Shimon Ullman

Pedestrian crossing intention prediction is essential for the deployment of autonomous vehicles (AVs) in urban environments. Ideal prediction provides AVs with critical environmental cues, thereby reducing the risk of pedestrian-related…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Yuanzhe Li , Steffen Müller

Multi-object tracking is an important ability for an autonomous vehicle to safely navigate a traffic scene. Current state-of-the-art follows the tracking-by-detection paradigm where existing tracks are associated with detected objects…

Computer Vision and Pattern Recognition · Computer Science 2021-10-12 Hsu-kuang Chiu , Jie Li , Rares Ambrus , Jeannette Bohg

Gaze following aims to interpret human-scene interactions by predicting the person's focal point of gaze. Prevailing approaches often adopt a two-stage framework, whereby multi-modality information is extracted in the initial stage for gaze…

Computer Vision and Pattern Recognition · Computer Science 2024-11-15 Yuehao Song , Xinggang Wang , Jingfeng Yao , Wenyu Liu , Jinglin Zhang , Xiangmin Xu

Due to the foveated nature of the human vision system, people can focus their visual attention on a small region of their visual field at a time, which usually contains only a single object. Estimating this object of attention in…

Computer Vision and Pattern Recognition · Computer Science 2019-12-17 Zehua Zhang , Chen Yu , David Crandall

Most modern multi-object tracking (MOT) systems follow the tracking-by-detection paradigm. It first localizes the objects of interest, then extracting their individual appearance features to make data association. The individual features,…

Computer Vision and Pattern Recognition · Computer Science 2020-07-02 Tianyi Liang , Long Lan , Zhigang Luo