中文
相关论文

相关论文: SoundShift: Exploring Sound Manipulations for Acce…

200 篇论文

We propose SoundBrush, a model that uses sound as a brush to edit and manipulate visual scenes. We extend the generative capabilities of the Latent Diffusion Model (LDM) to incorporate audio information for editing visual scenes. Inspired…

计算机视觉与模式识别 · 计算机科学 2025-01-03 Kim Sung-Bin , Kim Jun-Seong , Junseok Ko , Yewon Kim , Tae-Hyun Oh

Neurodivergent people frequently experience decreased sound tolerance, with estimates suggesting it affects 50-70% of this population. This heightened sensitivity can provoke reactions ranging from mild discomfort to severe distress,…

音频与语音处理 · 电气工程与系统科学 2025-09-15 Alexander Popescu , Rosie Frost , Milos Cernak

Current research in Mixed Reality (MR) presents a wide range of novel use cases for blending virtual elements with the real world. This yet-to-be-ubiquitous technology challenges how users currently work and interact with digital content.…

Teaching that uses projected presentation media such as slide-shows lacks support for dynamic content whose form and behaviors require live changes during a lecture. Recent software alternatives such as the Chalktalk software platform allow…

人机交互 · 计算机科学 2020-08-20 Zhenyi He , Ken Perlin

In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR)…

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work…

Recent progress in mixed reality (MR) and robotics is enabling increasingly sophisticated forms of human-robot collaboration. Building on these developments, we introduce a novel MR framework that allows multiple quadruped robots to operate…

Audio-visual navigation combines sight and hearing to navigate to a sound-emitting source in an unmapped environment. While recent approaches have demonstrated the benefits of audio input to detect and find the goal, they focus on clean and…

声音 · 计算机科学 2023-01-04 Abdelrahman Younes , Daniel Honerkamp , Tim Welschehold , Abhinav Valada

Binaural audio gives the listener an immersive experience and can enhance augmented and virtual reality. However, recording binaural audio requires specialized setup with a dummy human head having microphones in left and right ears. Such a…

计算机视觉与模式识别 · 计算机科学 2021-11-17 Kranti Kumar Parida , Siddharth Srivastava , Gaurav Sharma

This paper reports on the design and outcomes of the ICASSP SP Clarity Challenge: Speech Enhancement for Hearing Aids. The scenario was a listener attending to a target speaker in a noisy, domestic environment. There were multiple…

In Extended Reality (XR), rendering sound that accurately simulates real-world acoustics is pivotal in creating lifelike and believable virtual experiences. However, existing XR spatial audio rendering methods often struggle with real-time…

Acoustical mismatch among training and testing phases degrades outstandingly speech recognition results. This problem has limited the development of real-world nonspecific applications, as testing conditions are highly variant or even…

声音 · 计算机科学 2013-05-13 Rashmi Makhijani , Urmila Shrawankar , V M Thakare

Speech separation approaches for single-channel, dry speech mixtures have significantly improved. However, real-world spatial and reverberant acoustic environments remain challenging, limiting the effectiveness of these approaches for…

Interactive volume visualization using a mixed reality (MR) system helps provide users with an intuitive spatial perception of volumetric data. Due to sophisticated requirements of user interaction and vision when using MR head-mounted…

图形学 · 计算机科学 2023-09-06 Haojie Cheng , Chunxiao Xu , Xujing Chen , Zhenxin Chen , Jiajun Wang , Lingxiao Zhao

Our experience of the world is multisensory, spanning a synthesis of language, sight, sound, touch, taste, and smell. Yet, artificial intelligence has primarily advanced in digital modalities like text, vision, and audio. This paper…

机器学习 · 计算机科学 2026-01-14 Paul Pu Liang

Imagine being able to listen to the birds chirping in a park without hearing the chatter from other hikers, or being able to block out traffic noise on a busy street while still being able to hear emergency sirens and car honks. We…

声音 · 计算机科学 2023-11-02 Bandhav Veluri , Malek Itani , Justin Chan , Takuya Yoshioka , Shyamnath Gollakota

In the last years several solutions were proposed to support people with visual impairments or blindness during road crossing. These solutions focus on computer vision techniques for recognizing pedestrian crosswalks and computing their…

人机交互 · 计算机科学 2015-06-25 Sergio Mascetti , Lorenzo Picinali , Andrea Gerino , Dragan Ahmetovic , Cristian Bernareggi

The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, existing approaches…

多媒体 · 计算机科学 2023-08-21 Sung Jin Um , Dongjin Kim , Jung Uk Kim

Conventional voice conversion modifies voice characteristics from a source speaker to a target speaker, relying on audio input from both sides. However, this process becomes infeasible when clean audio is unavailable, such as in silent…

声音 · 计算机科学 2025-08-05 Yifan Liu , Yu Fang , Zhouhan Lin

In the Clarity project, we will run a series of machine learning challenges to revolutionise speech processing for hearing devices. Over five years, there will be three paired challenges. Each pair will consist of a competition focussed on…

音频与语音处理 · 电气工程与系统科学 2020-08-18 Simone Graetzer , Michael Akeroyd , Jon P. Barker , Trevor J. Cox , John F. Culling , Graham Naylor , Eszter Porter , Rhoddy Viveros Muñoz