中文
相关论文

相关论文: Seeing and Hearing Egocentric Actions: How Much Ca…

200 篇论文

We introduce the task of spatially localizing narrated interactions in videos. Key to our approach is the ability to learn to spatially localize interactions with self-supervision on a large corpus of videos with accompanying transcribed…

计算机视觉与模式识别 · 计算机科学 2021-12-03 Reuben Tan , Bryan A. Plummer , Kate Saenko , Hailin Jin , Bryan Russell

Learning an agent model that behaves like humans-capable of jointly perceiving the environment, predicting the future, and taking actions from a first-person perspective-is a fundamental challenge in computer vision. Existing methods…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Lu Chen , Yizhou Wang , Shixiang Tang , Qianhong Ma , Tong He , Wanli Ouyang , Xiaowei Zhou , Hujun Bao , Sida Peng

Egocentric assistants often rely on first-person view data to capture user behavior and context for personalized services. Since different users exhibit distinct habits, preferences, and routines, such personalization is essential for truly…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Yanshuo Wang , Yuan Xu , Xuesong Li , Jie Hong , Yizhou Wang , Chang Wen Chen , Wentao Zhu

Communicating in noisy, multi-talker environments is challenging, especially for people with hearing impairments. Egocentric video data can potentially be used to identify a user's conversation partners, which could be used to inform…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Tobias Dorszewski , Søren A. Fuglsang , Jens Hjortkjær

Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment. This early fusion of audio and visual cues, demonstrated through cognitive psychology and…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Shentong Mo , Pedro Morgado

Research in multi-modal interfaces aims to provide solutions to immersion and increase overall human performance. A promising direction is combining auditory, visual and haptic interaction between the user and the simulated environment.…

人机交互 · 计算机科学 2020-04-01 Eleftherios Triantafyllidis , Christopher McGreavy , Jiacheng Gu , Zhibin Li

In a wearable camera video, we see what the camera wearer sees. While this makes it easy to know roughly what he chose to look at, it does not immediately reveal when he was engaged with the environment. Specifically, at what moments did…

计算机视觉与模式识别 · 计算机科学 2016-04-05 Yu-Chuan Su , Kristen Grauman

Contact-rich manipulation tasks in unstructured environments often require both haptic and visual feedback. It is non-trivial to manually design a robot controller that combines these modalities which have very different characteristics.…

Recent advances in egocentric video understanding models are promising, but their heavy computational expense is a barrier for many real-world applications. To address this challenge, we propose EgoDistill, a distillation-based approach…

计算机视觉与模式识别 · 计算机科学 2023-01-06 Shuhan Tan , Tushar Nagarajan , Kristen Grauman

Egocentric videos offer fine-grained information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding a viewer's behaviors and intentions. We provide a labeled dataset…

计算机视觉与模式识别 · 计算机科学 2022-08-09 Lingzhi Zhang , Shenghao Zhou , Simon Stent , Jianbo Shi

Personality computing and affective computing have gained recent interest in many research areas. The datasets for the task generally have multiple modalities like video, audio, language and bio-signals. In this paper, we propose a flexible…

计算机视觉与模式识别 · 计算机科学 2023-01-13 Tanay Agrawal , Dhruv Agarwal , Michal Balazia , Neelabh Sinha , Francois Bremond

We present the Human And Robot Multimodal Observations of Natural Interactive Collaboration (HARMONIC) data set. This is a large multimodal data set of human interactions with a robotic arm in a shared autonomy setting designed to imitate…

机器人学 · 计算机科学 2020-08-03 Benjamin A. Newman , Reuben M. Aronson , Siddartha S. Srinivasa , Kris Kitani , Henny Admoni

Speech emotion recognition (SER) has received a great deal of attention in recent years in the context of spontaneous conversations. While there have been notable results on datasets like the well known corpus of naturalistic dyadic…

计算与语言 · 计算机科学 2024-01-02 Alex-Răzvan Ispas , Théo Deschamps-Berger , Laurence Devillers

Understanding affect is central to anticipating human behavior, yet current egocentric vision benchmarks largely ignore the person's emotional states that shape their decisions and actions. Existing tasks in egocentric perception focus on…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Matthias Jammot , Björn Braun , Paul Streli , Rafael Wampfler , Christian Holz

Human comprehension of a video stream is naturally broad: in a few instants, we are able to understand what is happening, the relevance and relationship of objects, and forecast what will follow in the near future, everything all at once.…

计算机视觉与模式识别 · 计算机科学 2024-03-06 Simone Alberto Peirone , Francesca Pistilli , Antonio Alliegro , Giuseppe Averta

Automatically describing video, or video captioning, has been widely studied in the multimedia field. This paper proposes a new task of sensor-augmented egocentric-video captioning, a newly constructed dataset for it called MMAC Captions,…

计算机视觉与模式识别 · 计算机科学 2021-09-08 Katsuyuki Nakamura , Hiroki Ohashi , Mitsuhiro Okada

We humans are good at translating third-person observations of hand-object interactions (HOI) into an egocentric view. However, current methods struggle to replicate this ability of view adaptation from third-person to first-person.…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Boshen Xu , Sipeng Zheng , Qin Jin

Understanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Dingkang Yang , Mingcheng Li , Linhao Qu , Kun Yang , Peng Zhai , Song Wang , Lihua Zhang

As the demand for analyzing egocentric videos grows, egocentric visual attention prediction, anticipating where a camera wearer will attend, has garnered increasing attention. However, it remains challenging due to the inherent complexity…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Sungjune Park , Hongda Mao , Qingshuang Chen , Yong Man Ro , Yelin Kim

Multimodal egocentric activity recognition integrates visual and inertial cues for robust first-person behavior understanding. However, deploying such systems in open-world environments requires detecting novel activities while continuously…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Wonseon Lim , Hyejeong Im , Dae-Won Kim