中文
相关论文

相关论文: How Much Does Audio Matter to Recognize Egocentric…

200 篇论文

Emotions play a critical role in our everyday lives by altering how we perceive, process and respond to our environment. Affective computing aims to instill in computers the ability to detect and act on the emotions of human actors. A core…

计算与语言 · 计算机科学 2020-08-31 Connor T. Heaton , David M. Schwartz

This paper presents a framework for recognition of human activity from egocentric video and eye tracking data obtained from a head-mounted eye tracker. Three channels of information such as eye movement, ego-motion, and visual features are…

计算机视觉与模式识别 · 计算机科学 2018-05-21 Anjith George , Aurobinda Routray

Advancements in egocentric video datasets like Ego4D, EPIC-Kitchens, and Ego-Exo4D have enriched the study of first-person human interactions, which is crucial for applications in augmented reality and assisted living. Despite these…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Joungbin An , Yunsu Park , Hyolim Kang , Seon Joo Kim

This work proposes to use passive acoustic perception as an additional sensing modality for intelligent vehicles. We demonstrate that approaching vehicles behind blind corners can be detected by sound before such vehicles enter in…

机器人学 · 计算机科学 2021-02-26 Yannick Schulz , Avinash Kini Mattar , Thomas M. Hehn , Julian F. P. Kooij

Egocentric gaze anticipation serves as a key building block for the emerging capability of Augmented Reality. Notably, gaze behavior is driven by both visual cues and audio signals during daily activities. Motivated by this observation, we…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Bolin Lai , Fiona Ryan , Wenqi Jia , Miao Liu , James M. Rehg

Wearable cameras capture a first-person view of the daily activities of the camera wearer, offering a visual diary of the user behaviour. Detection of the appearance of people the camera user interacts with for social interactions analysis…

计算机视觉与模式识别 · 计算机科学 2019-05-13 Estefania Talavera , Alexandre Cola , Nicolai Petkov , Petia Radeva

Having knowledge of the environmental context of the user i.e. the knowledge of the users' indoor location and the semantics of their environment, can facilitate the development of many of location-aware applications. In this paper, we…

声音 · 计算机科学 2018-04-03 Muhammad A. Shah , Bhiksha Raj , Khaled A. Harras

Environmental Sound Classification is an important problem of sound recognition and is more complicated than speech recognition problems as environmental sounds are not well structured with respect to time and frequency. Researchers have…

声音 · 计算机科学 2024-08-27 Aditya Dawn , Wazib Ansar

In the past few years, several case studies have illustrated that the use of occupancy information in buildings leads to energy-efficient and low-cost HVAC operation. The widely presented techniques for occupancy estimation include…

声音 · 计算机科学 2016-03-01 Qian Huang , Zhenhao Ge , Chao Lu

Sound can exert forces on objects of any material and shape. This has made the contactless manipulation of objects by intense ultrasound a fascinating area of research with wide-ranging applications. While much is understood for acoustic…

软凝聚态物质 · 物理学 2024-01-17 Melody X. Lim , Bryan VanSaders , Heinrich M. Jaeger

Egocentric vision captures the scene from the point of view of the camera wearer, while exocentric vision captures the overall scene context. Jointly modeling ego and exo views is crucial to developing next-generation AI agents. The…

计算机视觉与模式识别 · 计算机科学 2025-05-12 Anirudh Thatipelli , Shao-Yuan Lo , Amit K. Roy-Chowdhury

Recent progress in network-based audio event classification has shown the benefit of pre-training models on visual data such as ImageNet. While this process allows knowledge transfer across different domains, training a model on large-scale…

声音 · 计算机科学 2021-05-21 Sascha Hornauer , Ke Li , Stella X. Yu , Shabnam Ghaffarzadegan , Liu Ren

A key function of auditory cognition is the association of characteristic sounds with their corresponding semantics over time. Humans attempting to discriminate between fine-grained audio categories, often replay the same discriminative…

声音 · 计算机科学 2023-03-14 Alexandros Stergiou , Dima Damen

The audio data is increasing day by day throughout the globe with the increase of telephonic conversations, video conferences and voice messages. This research provides a mechanism for identifying a speaker in an audio file, based on the…

声音 · 计算机科学 2022-05-31 Syeda Rabia Arshad , Syed Mujtaba Haider , Abdul Basit Mughal

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the…

音频与语音处理 · 电气工程与系统科学 2022-05-13 Otavio Braga , Olivier Siohan

One of the many tasks facing the typically-developing child language learner is learning to discriminate between the distinctive sounds that make up words in their native language. Here we investigate whether multimodal…

计算与语言 · 计算机科学 2024-07-24 Sophia Zhi , Roger P. Levy , Stephan C. Meylan

Speech emotion recognition is a challenging problem because human convey emotions in subtle and complex ways. For emotion recognition on human speech, one can either extract emotion related features from audio signals or employ speech…

计算与语言 · 计算机科学 2020-04-06 Haiyang Xu , Hui Zhang , Kun Han , Yun Wang , Yiping Peng , Xiangang Li

Localizing visual sounds consists on locating the position of objects that emit sound within an image. It is a growing research area with potential applications in monitoring natural and urban environments, such as wildlife migration and…

声音 · 计算机科学 2022-04-12 Ho-Hsiang Wu , Magdalena Fuentes , Prem Seetharaman , Juan Pablo Bello

The increasing availability of wearable XR devices opens new perspectives for Egocentric Action Recognition (EAR) systems, which can provide deeper human understanding and situation awareness. However, deploying real-time algorithms on…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Marco Calzavara , Ard Kastrati , Matteo Macchini , Dushan Vasilevski , Roger Wattenhofer

Acoustic emotion recognition aims to categorize the affective state of the speaker and is still a difficult task for machine learning models. The difficulties come from the scarcity of training data, general subjectivity in emotion…

计算与语言 · 计算机科学 2018-04-02 Egor Lakomkin , Cornelius Weber , Sven Magg , Stefan Wermter