中文
相关论文

相关论文: Effect of acoustic scene complexity and visual sce…

200 篇论文

Prosody plays a vital role in verbal communication. Acoustic cues of prosody have been examined extensively. However, prosodic characteristics are not only perceived auditorily, but also visually based on head and facial movements. The…

计算与语言 · 计算机科学 2022-09-14 Hartmut Meister , Isa Samira Winter , Moritz Waeachtler , Pascale Sandmann , Khaled Abdellatif

Metaverse learning environments allow for a seamless and intuitive transition between activities compared to Virtual Reality (VR) learning environments, due to their interconnected design. The design of VR scenes is important for creating…

人机交互 · 计算机科学 2023-11-23 Rahatara Ferdousi , Mohammed Faisal , Fedwa Laamarti , Chunsheng Yang , Abdulmotaleb El Saddik

Urban noise maps and noise visualizations traditionally provide macroscopic representations of noise levels across cities. However, those representations fail at accurately gauging the sound perception associated with these sound…

计算机与社会 · 计算机科学 2024-07-25 Modan Tailleur , Pierre Aumond , Vincent Tourre , Mathieu Lagrange

We present an end-to-end binaural audio rendering approach (Listen2Scene) for virtual reality (VR) and augmented reality (AR) applications. We propose a novel neural-network-based binaural sound propagation method to generate acoustic…

音频与语音处理 · 电气工程与系统科学 2024-02-09 Anton Ratnarajah , Dinesh Manocha

This paper reports on the design and outcomes of the ICASSP SP Clarity Challenge: Speech Enhancement for Hearing Aids. The scenario was a listener attending to a target speaker in a noisy, domestic environment. There were multiple…

This paper describes an acoustic scene classification method which achieved the 4th ranking result in the IEEE AASP challenge of Detection and Classification of Acoustic Scenes and Events 2016. In order to accomplish the ensuing task,…

声音 · 计算机科学 2018-07-16 Sangwook Park , Seongkyu Mun , Younglo Lee , David K. Han , Hanseok Ko

Traditionally, in Audio Recognition pipeline, noise is suppressed by the "frontend", relying on preprocessing techniques such as speech enhancement. However, it is not guaranteed that noise will not cascade into downstream pipelines. To…

声音 · 计算机科学 2022-08-01 Juncheng B Li , Zheng Wang , Shuhui Qu , Florian Metze

Audio spoofing detection has become increasingly important due to the rise in real-world cases. Current spoofing detectors, referred to as spoofing countermeasures (CM), are mainly trained and focused on audio waveforms with a single…

声音 · 计算机科学 2024-08-27 Xuechen Liu , Xin Wang , Junichi Yamagishi

The creation of virtual humans increasingly leverages automated synthesis of speech and gestures, enabling expressive, adaptable agents that effectively engage users. However, the independent development of voice and gesture generation…

图形学 · 计算机科学 2025-07-02 Haoyang Du , Kiran Chhatre , Christopher Peters , Brian Keegan , Rachel McDonnell , Cathy Ennis

We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from…

For augmented (AR) and virtual reality (VR) applications, accurate estimates of the acoustic characteristics of a scene are critical for creating a sense of immersion. However, directly estimating Room-impulse Responses (RIRs) from scene…

音频与语音处理 · 电气工程与系统科学 2025-11-20 Ricardo Falcon-Perez , Ruohan Gao , Gregor Mueckl , Sebastia V. Amengual Gari , Ishwarya Ananthabhotla

Devices capable of detecting and categorizing acoustic scenes have numerous applications such as providing context-aware user experiences. In this paper, we address the task of characterizing acoustic scenes in a workplace setting from…

音频与语音处理 · 电气工程与系统科学 2019-11-12 Arindam Jati , Amrutha Nadarajan , Karel Mundnich , Shrikanth Narayanan

When human listeners try to guess the spatial position of a speech source, they are influenced by the speaker's production level, regardless of the intensity level reaching their ears. Because the perception of distance is a very difficult…

机器人学 · 计算机科学 2023-05-23 Ambre Davat , Véronique Aubergé , Gang Feng

We developed a novel assessment platform with untethered virtual reality, 3-dimensional sounds, and pressure sensing floor mat to help assess the walking balance and negotiation of obstacles given diverse sensory load and/or cognitive load.…

人机交互 · 计算机科学 2020-06-02 Zhu Wang , Anat Lubetzky , Charles Hendee , Marta Gospodarek , Ken Perlin

This paper addresses the issue of active speaker detection (ASD) in noisy environments and formulates a robust active speaker detection (rASD) problem. Existing ASD approaches leverage both audio and visual modalities, but non-speech sounds…

多媒体 · 计算机科学 2024-04-02 Siva Sai Nagender Vasireddy , Chenxu Zhang , Xiaohu Guo , Yapeng Tian

Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic conditions are relatively clean. To explore open-world speaker…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Liangbin Huang , Xiaohua Liao , Chaoqun Cui , Shijing Wang , Zhaolong Huang , Yanlong Du , Wenji Mao

Objective: Three perceptually orthogonal auditory dimensions for multidimensional and multivariate data sonification are identified and experimentally validated. Background: Psychoacoustic investigations have shown that orthogonal…

声音 · 计算机科学 2020-01-22 Tim Ziemer , Holger Schultheis

Remote telepresence via next-generation mixed reality platforms can provide higher levels of immersion for computer-mediated communications, allowing participants to engage in a wide spectrum of activities, previously not possible in 2D…

人机交互 · 计算机科学 2022-04-04 Mohammad Keshavarzi , Michael Zollhoefer , Allen Y. Yang , Patrick Peluse , Luisa Caldas

This paper investigates the application of environmental feature representations for room verification tasks and acoustic meta-data estimation. Audio recordings contain both speaker and non-speaker information. We refer to the…

声音 · 计算机科学 2022-03-10 Desmond Caulley

External Human-Machine Interfaces (eHMIs) have been proposed to facilitate communication between Automated Vehicles (AVs) and pedestrians. However, no attention was given to Deaf and Hard-of-Hearing (DHH) people. We conducted a formative…

人机交互 · 计算机科学 2026-01-21 Wenge Xu , Foroogh Hajiseyedjavadi , Kurtis Weir , Chukwuemeka Eze , Mark Colley