English
Related papers

Related papers: MoXaRt: Audio-Visual Object-Guided Sound Interacti…

200 papers

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these…

Computer Vision and Pattern Recognition · Computer Science 2022-01-07 Hao Jiang , Calvin Murdock , Vamsi Krishna Ithapu

To demonstrate the value of machine learning based smart health technologies, researchers have to deploy their solutions into complex real-world environments with real participants. This gives rise to many, oftentimes unexpected, challenges…

Vision is often used as a complementary modality for audio speech recognition (ASR), especially in the noisy environment where performance of solo audio modality significantly deteriorates. After combining visual modality, ASR is upgraded…

Computer Vision and Pattern Recognition · Computer Science 2020-05-14 Bo Xu , Cheng Lu , Yandong Guo , Jacob Wang

Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, these tasks have been…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-16 Daniel Michelsanti , Zheng-Hua Tan , Shi-Xiong Zhang , Yong Xu , Meng Yu , Dong Yu , Jesper Jensen

Learning how objects sound from video is challenging, since they often heavily overlap in a single audio channel. Current methods for visually-guided audio source separation sidestep the issue by training with artificially mixed video…

Computer Vision and Pattern Recognition · Computer Science 2019-08-22 Ruohan Gao , Kristen Grauman

Target speech separation refers to isolating target speech from a multi-speaker mixture signal by conditioning on auxiliary information about the target speaker. Different from the mainstream audio-visual approaches which usually require…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Leyuan Qu , Cornelius Weber , Stefan Wermter

While automatic speech recognition (ASR) systems degrade significantly in noisy environments, audio-visual speech recognition (AVSR) systems aim to complement the audio stream with noise-invariant visual cues and improve the system's…

Sound · Computer Science 2024-04-09 He Wang , Pengcheng Guo , Pan Zhou , Lei Xie

Recently, significant progress has been made in audio source separation by the application of deep learning techniques. Current methods that combine both audio and visual information use 2D representations such as images to guide the…

Sound · Computer Science 2021-02-04 Francesc Lluís , Vasileios Chatziioannou , Alex Hofmann

Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR,…

This paper introduces IRIS, an Immersive Robot Interaction System leveraging Extended Reality (XR). Existing XR-based systems enable efficient data collection but are often challenging to reproduce and reuse due to their specificity to…

The prevalence of Internet of things (IoT) devices and abundance of sensor data has created an increase in real-time data processing such as recognition of speech, image, and video. While currently such processes are offloaded to the…

Computer Vision and Pattern Recognition · Computer Science 2018-03-22 Ramyad Hadidi , Jiashen Cao , Matthew Woodward , Michael S. Ryoo , Hyesoon Kim

The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, existing approaches…

Multimedia · Computer Science 2023-08-21 Sung Jin Um , Dongjin Kim , Jung Uk Kim

With the widespread adoption of Extended Reality (XR) headsets, spatial computing technologies are gaining increasing attention. Spatial computing enables interaction with virtual elements through natural input methods such as eye tracking,…

Human-Computer Interaction · Computer Science 2025-06-16 Zhimin Wang , Maohang Rao , Shanghua Ye , Weitao Song , Feng Lu

eXtended Reality (XR) autism research, ranging from Augmented Reality to Virtual Reality, focuses on socio-emotional abilities and high-functioning autism. However common autism interventions address the entire spectrum over social, sensory…

Human-Computer Interaction · Computer Science 2022-04-11 Valentin Bauer , Tifanie Bouchara , Patrick Bourdot

Conventional career guidance platforms rely on static, text-driven interfaces that struggle to engage users or deliver personalised, evidence-based insights. Although Computer-Assisted Career Guidance Systems have evolved since the 1960s,…

Computational Engineering, Finance, and Science · Computer Science 2026-04-09 N. D. Tantaroudas , A. J. McCracken , I. Karachalios , E. Papatheou , V. Pastrikakis

A virtual acoustic source inside a medium can be created by emitting a time-reversed point-source response from the enclosing boundary into the medium. However, in many practical situations the medium can be accessed from one side only. In…

Recent advances in generative models have enabled modern Text-to-Audio (TTA) systems to synthesize audio with high perceptual quality. However, TTA systems often struggle to maintain semantic consistency with the input text, leading to…

Sound · Computer Science 2026-01-13 Bochao Sun , Yang Xiao , Han Yin

Multimodal large language models (MLLMs) require a nuanced interpretation of complex image information, typically leveraging a vision encoder to perceive various visual scenarios. However, relying solely on a single vision encoder to handle…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Xin He , Xumeng Han , Longhui Wei , Lingxi Xie , Qi Tian

Integrating mixed reality (MR) with artificial intelligence (AI) technologies, including vision, language, audio, reasoning, and planning, enables the AI-powered MR assistant [1] to substantially elevate human efficiency. This enhancement…

Human-Computer Interaction · Computer Science 2024-05-10 Yan-Ming Chiou , Bob Price , Chien-Chung Shen , Syed Ali Asif

We propose a mesh-based neural network (MESH2IR) to generate acoustic impulse responses (IRs) for indoor 3D scenes represented using a mesh. The IRs are used to create a high-quality sound experience in interactive applications and audio…

Sound · Computer Science 2022-07-13 Anton Ratnarajah , Zhenyu Tang , Rohith Chandrashekar Aralikatti , Dinesh Manocha