中文
相关论文

相关论文: A Modular Multimodal Architecture for Gaze Target …

200 篇论文

Turn-taking prediction is crucial for seamless interactions. This study introduces a novel, lightweight framework for accurate turn-taking prediction in triadic conversations without relying on computationally intensive methods. Unlike…

人机交互 · 计算机科学 2025-05-30 Seongsil Heo , Calvin Murdock , Michael Proulx , Christi Miller

Gaze target detection (GTD) is the task of predicting where a person in an image is looking. This is a challenging task, as it requires the ability to understand the relationship between the person's head, body, and eyes, as well as the…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Athul M. Mathew , Arshad Ali Khan , Thariq Khalid , Faroq AL-Tam , Riad Souissi

Understanding the human-object interactions (HOIs) from a video is essential to fully comprehend a visual scene. This line of research has been addressed by detecting HOIs from images and lately from videos. However, the video-based HOI…

计算机视觉与模式识别 · 计算机科学 2023-06-07 Zhifan Ni , Esteve Valls Mascaró , Hyemin Ahn , Dongheui Lee

Effective assisted living environments must be able to perform inferences on how their occupants interact with one another as well as with surrounding objects. To accomplish this goal using a vision-based automated approach, multiple tasks…

计算机视觉与模式识别 · 计算机科学 2019-09-23 Philipe A. Dias , Damiano Malafronte , Henry Medeiros , Francesca Odone

In autonomous driving and robotics, there is a growing interest in utilizing short-term historical data to enhance multi-camera 3D object detection, leveraging the continuous and correlated nature of input video streams. Recent work has…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Seokha Moon , Hongbeen Park , Jungphil Kwon , Jaekoo Lee , Jinkyu Kim

Human affordance learning investigates contextually relevant novel pose prediction such that the estimated pose represents a valid human action within the scene. While the task is fundamental to machine perception and automated interactive…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Prasun Roy , Saumik Bhattacharya , Subhankar Ghosh , Umapada Pal , Michael Blumenstein

We present GazeMotion, a novel method for human motion forecasting that combines information on past human poses with human eye gaze. Inspired by evidence from behavioural sciences showing that human eye and body movements are closely…

计算机视觉与模式识别 · 计算机科学 2024-07-12 Zhiming Hu , Syn Schmitt , Daniel Haeufle , Andreas Bulling

Close human-robot cooperation is a key enabler for new developments in advanced manufacturing and assistive applications. Close cooperation require robots that can predict human actions and intent, and understand human non-verbal cues.…

人机交互 · 计算机科学 2019-02-19 Paul Schydlo , Mirko Rakovic , Lorenzo Jamone , José Santos-Victor

Predicting pedestrian behavior is a crucial task for intelligent driving systems. Accurate predictions require a deep understanding of various contextual elements that potentially impact the way pedestrians behave. To address this…

计算机视觉与模式识别 · 计算机科学 2022-10-17 Amir Rasouli , Iuliia Kotseruba

Training a multimodal network is challenging and it requires complex architectures to achieve reasonable performance. We show that one reason for this phenomena is the difference between the convergence rate of various modalities. We…

人工智能 · 计算机科学 2020-11-13 Aya Abdelsalam Ismail , Mahmudul Hasan , Faisal Ishtiaq

By borrowing the wisdom of human in gaze following, we propose a two-stage solution for gaze point prediction of the target persons in a scene. Specifically, in the first stage, both head image and its position are fed into a gaze direction…

计算机视觉与模式识别 · 计算机科学 2019-07-05 Dongze Lian , Zehao Yu , Shenghua Gao

Trajectory prediction is crucial for autonomous vehicles. The planning system not only needs to know the current state of the surrounding objects but also their possible states in the future. As for vehicles, their trajectories are…

机器人学 · 计算机科学 2020-07-07 Chenxu Luo , Lin Sun , Dariush Dabiri , Alan Yuille

The potential of multimodal generative artificial intelligence (mAI) to replicate human grounded language understanding, including the pragmatic, context-rich aspects of communication, remains to be clarified. Humans are known to use…

We address the problem of detecting attention targets in video. Our goal is to identify where each person in each frame of a video is looking, and correctly handle the case where the gaze target is out-of-frame. Our novel architecture…

计算机视觉与模式识别 · 计算机科学 2020-04-01 Eunji Chong , Yongxin Wang , Nataniel Ruiz , James M. Rehg

Joint visual attention is characterized by two or more individuals looking at a common target at the same time. The ability to identify joint attention in scenes, the people involved, and their common target, is fundamental to the…

计算机视觉与模式识别 · 计算机科学 2018-04-13 Daniel Harari , Joshua B. Tenenbaum , Shimon Ullman

Pedestrian crossing intention prediction is essential for the deployment of autonomous vehicles (AVs) in urban environments. Ideal prediction provides AVs with critical environmental cues, thereby reducing the risk of pedestrian-related…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Yuanzhe Li , Steffen Müller

Multi-object tracking is an important ability for an autonomous vehicle to safely navigate a traffic scene. Current state-of-the-art follows the tracking-by-detection paradigm where existing tracks are associated with detected objects…

计算机视觉与模式识别 · 计算机科学 2021-10-12 Hsu-kuang Chiu , Jie Li , Rares Ambrus , Jeannette Bohg

Gaze following aims to interpret human-scene interactions by predicting the person's focal point of gaze. Prevailing approaches often adopt a two-stage framework, whereby multi-modality information is extracted in the initial stage for gaze…

计算机视觉与模式识别 · 计算机科学 2024-11-15 Yuehao Song , Xinggang Wang , Jingfeng Yao , Wenyu Liu , Jinglin Zhang , Xiangmin Xu

Due to the foveated nature of the human vision system, people can focus their visual attention on a small region of their visual field at a time, which usually contains only a single object. Estimating this object of attention in…

计算机视觉与模式识别 · 计算机科学 2019-12-17 Zehua Zhang , Chen Yu , David Crandall

Most modern multi-object tracking (MOT) systems follow the tracking-by-detection paradigm. It first localizes the objects of interest, then extracting their individual appearance features to make data association. The individual features,…

计算机视觉与模式识别 · 计算机科学 2020-07-02 Tianyi Liang , Long Lan , Zhigang Luo