中文
相关论文

相关论文: SonifyAR: Context-Aware Sound Generation in Augmen…

200 篇论文

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

声音 · 计算机科学 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

Robotic perception is becoming a key technology for navigation aids, especially helping individuals with visual impairments through spatial sonification. This paper introduces a mapping representation that accurately captures scene geometry…

机器人学 · 计算机科学 2025-04-18 Lan Wu , Craig Jin , Monisha Mushtary Uttsha , Teresa Vidal-Calleja

Physical environment understanding is vital in delivering immersive and interactive mobile augmented reality (AR) user experiences. Recently, we have witnessed a transition in the design of environment understanding systems, from visual…

分布式、并行与集群计算 · 计算机科学 2023-10-18 Yiqin Zhao , Ashkan Ganj , Tian Guo

The growing adoption of augmented and virtual reality (AR and VR) technologies in industrial training and on-the-job assistance has created new opportunities for intelligent, context-aware support systems. As workers perform complex tasks…

人机交互 · 计算机科学 2025-11-18 Mahya Qorbani , Kamran Paynabar , Mohsen Moghaddam

In multimedia applications such as films and video games, spatial audio techniques are widely employed to enhance user experiences by simulating 3D sound: transforming mono audio into binaural formats. However, this process is often complex…

多媒体 · 计算机科学 2025-02-14 Xiaojing Liu , Ogulcan Gurelli , Yan Wang , Joshua Reiss

Spatial audio is an essential medium to audiences for 3D visual and auditory experience. However, the recording devices and techniques are expensive or inaccessible to the general public. In this work, we propose a self-supervised audio…

声音 · 计算机科学 2019-05-15 Yu-Ding Lu , Hsin-Ying Lee , Hung-Yu Tseng , Ming-Hsuan Yang

Social robots are required not only to understand human intentions but also to effectively communicate their intentions or own internal states to users. This study explores the use of sonification to provide explicit auditory feedback,…

机器人学 · 计算机科学 2024-11-15 Simone Arreghini , Antonio Paolillo , Gabriele Abbate , Alessandro Giusti

In Extended Reality (XR), rendering sound that accurately simulates real-world acoustics is pivotal in creating lifelike and believable virtual experiences. However, existing XR spatial audio rendering methods often struggle with real-time…

Developing embodied agents in simulation has been a key research topic in recent years. Exciting new tasks, algorithms, and benchmarks have been developed in various simulators. However, most of them assume deaf agents in silent…

机器人学 · 计算机科学 2023-09-19 Ruohan Gao , Hao Li , Gokul Dharan , Zhuzhu Wang , Chengshu Li , Fei Xia , Silvio Savarese , Li Fei-Fei , Jiajun Wu

In this position paper, we propose researching the combination of Augmented Reality (AR) and Artificial Intelligence (AI) to support conversations, inspired by the interfaces of dialogue systems commonly found in videogames. AR-capable…

人机交互 · 计算机科学 2025-03-10 Julián Méndez , Marc Satkowski

Public speaking anxiety affects many individuals, yet opportunities for real-world practice remain limited. This study explores how augmented reality (AR) can provide an accessible training environment for public speaking. Drawing from…

人机交互 · 计算机科学 2025-04-16 Mark Edison Jim , Jan Benjamin Yap , Gian Chill Laolao , Andrei Zachary Lim , Jordan Aiko Deja

This paper presents Matrix, an advanced AI-powered framework designed for real-time 3D object generation in Augmented Reality (AR) environments. By integrating a cutting-edge text-to-3D generative AI model, multilingual speech-to-text…

人机交互 · 计算机科学 2025-03-24 Majid Behravan , Denis Gracanin

In video game design, audio (both environmental background music and object sound effects) play a critical role. Sounds are typically pre-created assets designed for specific locations or objects in a game. However, user-generated content…

人机交互 · 计算机科学 2024-04-29 Thomas Marrinan , Pakeeza Akram , Oli Gurmessa , Anthony Shishkin

Augmented Reality (AR) systems are increasingly integrating foundation models, such as Multimodal Large Language Models (MLLMs), to provide more context-aware and adaptive user experiences. This integration has led to the development of AR…

人工智能 · 计算机科学 2025-08-13 Dongwook Choi , Taeyoon Kwon , Dongil Yang , Hyojun Kim , Jinyoung Yeo

Sound designers search for sounds in large sound effects libraries using aspects such as sound class or visual context. However, the metadata needed for such search is often missing or incomplete, and requires significant manual effort to…

声音 · 计算机科学 2026-02-17 Sripathi Sridhar , Prem Seetharaman , Oriol Nieto , Mark Cartwright , Justin Salamon

Mixed-reality (MR) soundscapes blend real-world sound with virtual audio from hearing devices, presenting intricate auditory information that is hard to discern and differentiate. This is particularly challenging for blind or visually…

人机交互 · 计算机科学 2024-05-28 Ruei-Che Chang , Chia-Sheng Hung , Bing-Yu Chen , Dhruv Jain , Anhong Guo

In Augmented Reality (AR) environment, realistic interactions between the virtual and real objects play a crucial role in user experience. Much of recent advances in AR has been largely focused on developing geometry-aware environment, but…

计算机视觉与模式识别 · 计算机科学 2018-03-19 Long Chen , Karl Francis , Wen Tang

We present a virtual reality (VR) environment featuring conversational avatars powered by a locally-deployed LLM, integrated with automatic speech recognition (ASR), text-to-speech (TTS), and lip-syncing. Through a pilot study, we explored…

人机交互 · 计算机科学 2025-01-08 Mykola Maslych , Christian Pumarada , Amirpouya Ghasemaghaei , Joseph J. LaViola

Recently, a lot of works show promising directions for audio design in augmented reality (AR). These works are mainly focused on how to improve user experience and make AR more realistic. But even though these improvements seem promising,…

人机交互 · 计算机科学 2022-09-07 Esmée Henrieke Anne de Haas , Lik-Hang Lee

Training audio-to-image generative models requires an abundance of diverse audio-visual pairs that are semantically aligned. Such data is almost always curated from in-the-wild videos, given the cross-modal semantic correspondence that is…

声音 · 计算机科学 2025-01-10 Darius Petermann , Mahdi M. Kalayeh