中文
相关论文

相关论文: A Modular Multimodal Architecture for Gaze Target …

200 篇论文

Camera, LiDAR and radar are common perception sensors for autonomous driving tasks. Robust prediction of 3D object detection is optimally based on the fusion of these sensors. To exploit their abilities wisely remains a challenge because…

计算机视觉与模式识别 · 计算机科学 2024-05-21 Ziang Guo , Zakhar Yagudin , Selamawit Asfaw , Artem Lykov , Dzmitry Tsetserukou

The visual focus of attention (VFOA) has been recognized as a prominent conversational cue. We are interested in estimating and tracking the VFOAs associated with multi-party social interactions. We note that in this type of situations the…

计算机视觉与模式识别 · 计算机科学 2018-12-21 Benoît Massé , Silèye Ba , Radu Horaud

Multi-object tracking (MOT) has profound applications in a variety of fields, including surveillance, sports analytics, self-driving, and cooperative robotics. Despite considerable advancements, existing MOT methodologies tend to falter…

计算机视觉与模式识别 · 计算机科学 2023-12-20 Hamza Mukhtar , Muhammad Usman Ghani Khan

For successful deployment of robots in multifaceted situations, an understanding of the robot for its environment is indispensable. With advancing performance of state-of-the-art object detectors, the capability of robots to detect objects…

人机交互 · 计算机科学 2023-03-02 Daniel Weber , Wolfgang Fuhl , Enkelejda Kasneci , Andreas Zell

Its numerous applications make multi-human 3D pose estimation a remarkably impactful area of research. Nevertheless, assuming a multiple-view system composed of several regular RGB cameras, 3D multi-pose estimation presents several…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Daniel Rodriguez-Criado , Pilar Bachiller , George Vogiatzis , Luis J. Manso

Making accurate inferences about other individuals' locus of attention is essential for human social interactions and will be important for AI to effectively interact with humans. In this study, we compare how a CNN (convolutional neural…

计算机视觉与模式识别 · 计算机科学 2021-04-20 Nicole X. Han , William Yang Wang , Miguel P. Eckstein

Most human behaviors consist of multiple parts, steps, or subtasks. These structures guide our action planning and execution, but when we observe others, the latent structure of their actions is typically unobservable, and must be inferred…

人工智能 · 计算机科学 2018-09-28 Ryo Nakahashi , Chris L. Baker , Joshua B. Tenenbaum

This article describes a multi-modal method using simulated Lidar data via ray tracing and image pixel loss with differentiable rendering to optimize an object's position with respect to an observer or some referential objects in a computer…

系统与控制 · 电气工程与系统科学 2023-09-07 Sean Zanyk-McLean , Krishna Kumar , Paul Navratil

Gaze estimation involves predicting where the person is looking at within an image or video. Technically, the gaze information can be inferred from two different magnification levels: face orientation and eye orientation. The inference is…

计算机视觉与模式识别 · 计算机科学 2021-10-27 Ashesh , Chu-Song Chen , Hsuan-Tien Lin

An increasing number of works explore collaborative human-computer systems in which human gaze is used to enhance computer vision systems. For object detection these efforts were so far restricted to late integration approaches that have…

计算机视觉与模式识别 · 计算机科学 2015-05-22 Iaroslav Shcherbatyi , Andreas Bulling , Mario Fritz

Multiple clustering has gained significant attention in recent years due to its potential to reveal multiple hidden structures of data from different perspectives. The advent of deep multiple clustering techniques has notably advanced the…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Jiawei Yao , Qi Qian , Juhua Hu

Human pose forecasting garners attention for its diverse applications. However, challenges in modeling the multi-modal nature of human motion and intricate interactions among agents persist, particularly with longer timescales and more…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Jaewoo Jeong , Daehee Park , Kuk-Jin Yoon

Predicting the behavior of road users accurately is crucial to enable the safe operation of autonomous vehicles in urban or densely populated areas. Therefore, there has been a growing interest in time series motion prediction research,…

机器学习 · 计算机科学 2024-10-22 Camiel Oerlemans , Bram Grooten , Michiel Braat , Alaa Alassi , Emilia Silvas , Decebal Constantin Mocanu

We introduce a framework for navigating through cluttered environments by connecting multiple cameras together while simultaneously preserving privacy. Occlusions and obstacles in large environments are often challenging situations for…

机器学习 · 计算机科学 2022-12-05 Hui Lu , Mia Chiquier , Carl Vondrick

Anticipating human motion in crowded scenarios is essential for developing intelligent transportation systems, social-aware robots and advanced video surveillance applications. A key component of this task is represented by the inherently…

计算机视觉与模式识别 · 计算机科学 2021-07-09 Alessia Bertugli , Simone Calderara , Pasquale Coscia , Lamberto Ballan , Rita Cucchiara

Advancements in multimodal foundation models have enabled the development of Computer Use Agents (CUAs) capable of autonomously interacting with GUI environments. As CUAs are not restricted to certain tools, they allow to automate more…

机器学习 · 计算机科学 2026-04-10 Dominik Seip , Matthias Hein

Perceiving humans in the context of Intelligent Transportation Systems (ITS) often relies on multiple cameras or expensive LiDAR sensors. In this work, we present a new cost-effective vision-based method that perceives humans' locations in…

计算机视觉与模式识别 · 计算机科学 2021-05-04 Lorenzo Bertoni , Sven Kreiss , Alexandre Alahi

Humans express their emotions via facial expressions, voice intonation and word choices. To infer the nature of the underlying emotion, recognition models may use a single modality, such as vision, audio, and text, or a combination of…

机器学习 · 计算机科学 2022-02-21 Vandana Rajan , Alessio Brutti , Andrea Cavallaro

Gaze behaviors such as eye-contact or shared attention are important markers for diagnosing developmental disorders in children. While previous studies have looked at some of these elements, the analysis is usually performed on private…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Samy Tafasca , Anshul Gupta , Jean-Marc Odobez

Precisely detecting which object a person is paying attention to is critical for human-robot interaction since it provides important cues for the next action from the human user. We propose an end-to-end approach for gaze target detection:…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Zhi-Yi Lin , Jouh Yeong Chew , Jan van Gemert , Xucong Zhang