English
Related papers

Related papers: Acoustivision Pro: An Open-Source Interactive Plat…

200 papers

Reconfigurable intelligent surface (RIS) has been widely discussed as new technology to improve wireless communication performance. Based on the unique design of RIS, its elements can reflect, refract, absorb, or focus the incoming waves…

Signal Processing · Electrical Eng. & Systems 2020-09-03 Salah Eddine Zegrar , Liza Afeef , Huseyin Arslan

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-13 Otavio Braga , Olivier Siohan

Fusion of scores is a cornerstone of multimodal biometric systems composed of independent unimodal parts. In this work, we focus on quality-dependent fusion for speaker-face verification. To this end, we propose a universal model which can…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-17 Grigory Antipov , Nicolas Gengembre , Olivier Le Blouch , Gaël Le Lan

Intelligent reflecting surface (IRS) has emerged as a key enabling technology to realize smart and reconfigurable radio environment for wireless communications, by digitally controlling the signal reflection via a large number of passive…

Information Theory · Computer Science 2022-02-02 Beixiong Zheng , Changsheng You , Weidong Mei , Rui Zhang

We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from…

One of the challenges in computational acoustics is the identification of models that can simulate and predict the physical behavior of a system generating an acoustic signal. Whenever such models are used for commercial applications an…

Room reidentification (ReID) is a challenging yet essential task with numerous applications in fields such as augmented reality (AR) and homecare robotics. Existing visual place recognition (VPR) methods, which typically rely on global…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Runmao Yao , Yi Du , Zhuoqun Chen , Haoze Zheng , Chen Wang

In this paper, we propose a novel computer vision-based approach to aid Reconfigurable Intelligent Surface (RIS) for dynamic beam tracking and then implement the corresponding prototype verification system. A camera is attached at the RIS…

Signal Processing · Electrical Eng. & Systems 2022-07-12 Ming Ouyang , Yucong Wang , Feifei Gao , Shun Zhang , Puchu Li , Jian Ren

This paper presents our modeling and architecture approaches for building a highly accurate low-latency language identification system to support multilingual spoken queries for voice assistants. A common approach to solve multilingual…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-02 Chander Chandak , Zeynab Raeesy , Ariya Rastrow , Yuzong Liu , Xiangyang Huang , Siyu Wang , Dong Kwon Joo , Roland Maas

With the widespread application of automatic speech recognition (ASR) systems, their vulnerability to adversarial attacks has been extensively studied. However, most existing adversarial examples are generated on specific individual models,…

Sound · Computer Science 2025-03-26 Weifei Jin , Junjie Su , Hejia Wang , Yulin Ye , Jie Hao

Voice assistants (VAs) are typically evaluated through task performance metrics and self-report questionnaires, but people's voices themselves carry rich paralinguistic cues that reveal affect, effort, and interaction breakdowns. We present…

Human-Computer Interaction · Computer Science 2026-03-23 Yong Ma , Xuesong Zhang , Xuedong Zhang , Natalia Bartłomiejczyk , Seungwoo Je , Adrian Holzer , Morten Fjeld , Andreas Butz

This experience report reflects on researching misophonia as someone who lives with it. Misophonia is an aversive response to everyday sounds (chewing, sniffling, pen clicking) and, for many of us, to associated visual cues (misokinesia).…

Human-Computer Interaction · Computer Science 2026-05-13 Tawfiq Ammari

This paper introduces ClearerVoice-Studio, an open-source, AI-powered speech processing toolkit designed to bridge cutting-edge research and practical application. Unlike broad platforms like SpeechBrain and ESPnet, ClearerVoice-Studio…

Sound · Computer Science 2025-06-25 Shengkui Zhao , Zexu Pan , Bin Ma

Predicting spatially varying Room Impulse Response (RIR) from sparse observations is a critical but highly challenging inverse problem for immersive spatial audio rendering. In this work, we present EIGENET, a geometry-informed multi-modal…

Sound · Computer Science 2026-05-28 Chong Jing , Zitong Lan , Junan Zhang , Zhizheng Wu

The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, existing approaches…

Multimedia · Computer Science 2023-08-21 Sung Jin Um , Dongjin Kim , Jung Uk Kim

Different methods can be employed to render virtual reverberation, often requiring substantial information about the room's geometry and the acoustic characteristics of the surfaces. However, fully comprehensive approaches that account for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-18 Vincent Martin , Isaac Engel , Lorenzo Picinali

In this paper we present a novel algorithm for improved block-online supervised acoustic system identification in adverse noise scenarios by exploiting prior knowledge about the space of Room Impulse Responses (RIRs). The method is based on…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-06 Thomas Haubner , Andreas Brendel , Walter Kellermann

Recent studies on learning-based sound source localization have mainly focused on the localization performance perspective. However, prior work and existing benchmarks overlook a crucial aspect: cross-modal interaction, which is essential…

Multimedia · Computer Science 2024-07-19 Arda Senocak , Hyeonggon Ryu , Junsik Kim , Tae-Hyun Oh , Hanspeter Pfister , Joon Son Chung

This paper considers methods for audio display in a CAVE-type virtual reality theater, a 3 m cube with displays covering all six rigid faces. Headphones are possible since the user's headgear continuously measures ear positions, but…

Sound · Computer Science 2011-06-08 Bowon Lee , Camille Goudeseune , Mark A. Hasegawa-Johnson

For d/Deaf and hard of hearing (DHH) people, captioning is an essential accessibility tool. Significant developments in artificial intelligence (AI) mean that Automatic Speech Recognition (ASR) is now a part of many popular applications.…

Computation and Language · Computer Science 2024-08-30 Korbinian Kuhn , Verena Kersken , Benedikt Reuter , Niklas Egger , Gottfried Zimmermann
‹ Prev 1 8 9 10 Next ›