中文
相关论文

相关论文: SoundShift: Exploring Sound Manipulations for Acce…

200 篇论文

Generative spoken language models produce speech in a wide range of voices, prosody, and recording conditions, seemingly approaching the diversity of natural speech. However, the extent to which generated speech is acoustically diverse…

音频与语音处理 · 电气工程与系统科学 2025-03-12 Matthieu Futeral , Andrea Agostinelli , Marco Tagliasacchi , Neil Zeghidour , Eugene Kharitonov

The growing adoption of electric vehicles, known for their quieter operation compared to internal combustion engine vehicles, raises concerns about their detectability, particularly for vulnerable road users. To address this, regulations…

人机交互 · 计算机科学 2026-02-03 Pavlo Bazilinskyy , Md Shadab Alam , Roberto Merino-Martınez

Virtual Reality (VR) users often experience postural instability, i.e., balance problems, which could be a major barrier to universal usability and accessibility for all, especially for persons with balance impairments. Prior research has…

人机交互 · 计算机科学 2022-02-11 M. Rasel Mahmud , Michael Stewart , Alberto Cordova , John Quarles

Acoustic sensing has proved effective as a foundation for numerous applications in health and human behavior analysis. In this work, we focus on the problem of detecting in-person social interactions in naturalistic settings from audio…

声音 · 计算机科学 2022-03-23 Dawei Liang , Zifan Xu , Yinuo Chen , Rebecca Adaimi , David Harwath , Edison Thomaz

Sound plays a significant role in human memory, yet it is often overlooked by mainstream life-recording methods. Most current UGC (User-Generated Content) creation tools emphasize visual content while lacking user-friendly sound design…

人机交互 · 计算机科学 2024-10-11 Chongjun Zhong , Jiaxing Yu , Yingping Cao , Songruoyao Wu , Wenqi Wu , Kejun Zhang

Dysarthria is a neurological disorder that significantly impairs speech intelligibility, often rendering affected individuals unable to communicate effectively. This necessitates the development of robust dysarthric-to-regular speech…

声音 · 计算机科学 2025-06-23 Shoutrik Das , Nishant Singh , Arjun Gangwar , S Umesh

This demo paper presents sign.mt, an open-source application pioneering real-time multilingual bi-directional translation between spoken and signed languages. Harnessing state-of-the-art open-source models, this tool aims to address the…

计算与语言 · 计算机科学 2024-08-06 Amit Moryossef

We live in a world where 60% of the population can speak two or more languages fluently. Members of these communities constantly switch between languages when having a conversation. As automatic speech recognition (ASR) systems are being…

计算与语言 · 计算机科学 2021-02-16 Siddharth Dalmia , Yuzong Liu , Srikanth Ronanki , Katrin Kirchhoff

In crowded places such as conferences, background noise, overlapping voices, and lively interactions make it difficult to have clear conversations. This situation often worsens the phenomenon known as "cocktail party deafness." We present…

声音 · 计算机科学 2025-12-04 Lixing He , Yunqi Guo , Zhenyu Yan , Guoliang Xing

We introduce and explore a new multimodal input representation for vision-language models: acoustic field video. Unlike conventional video (RGB with stereo/mono audio), our video stream provides a spatially grounded visualization of sound…

人机交互 · 计算机科学 2026-01-27 Daehwa Kim , Chris Harrison

This paper proposes a neural network that performs audio transformations to user-specified sources (e.g., vocals) of a given audio track according to a given description while preserving other sources not mentioned in the description. Audio…

音频与语音处理 · 电气工程与系统科学 2021-04-29 Woosung Choi , Minseok Kim , Marco A. Martínez Ramírez , Jaehwa Chung , Soonyoung Jung

Automatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, particularly in non-English languages. To address this, we fine-tune a voice conversion model on English dysarthric speech (UASpeech) to…

Automatic Speech Recognition (ASR) has advanced with Speech Foundation Models (SFMs), yet performance degrades on dysarthric speech due to variability and limited data. This study as part of the submission to the Speech Accessibility…

音频与语音处理 · 电气工程与系统科学 2025-05-28 Alexandre Ducorroy , Rachid Riad

Target speech separation refers to extracting the target speaker's speech from mixed signals. Despite the recent advances in deep learning based close-talk speech separation, the applications to real-world are still an open issue. Two main…

声音 · 计算机科学 2020-01-03 Rongzhi Gu , Yuexian Zou

This paper is concerned with the task of speaker verification on audio with multiple overlapping speakers. Most speaker verification systems are designed with the assumption of a single speaker being present in a given audio segment.…

音频与语音处理 · 电气工程与系统科学 2023-04-10 Jenthe Thienpondt , Nilesh Madhu , Kris Demuynck

Many hearing-impaired listeners struggle to localize sounds due to poor availability of binaural cues. Listeners with a cochlear implant and a contralateral hearing aid -- so-called bimodal listeners -- are amongst the worst performers, as…

音频与语音处理 · 电气工程与系统科学 2018-03-15 Benjamin Dieudonné , Tom Francart

This project brings music to sight. Music can be a visual masterpiece. Some people naturally experience a visualization of audio - a condition called synesthesia. The type of synesthesia explored is when sounds create colors in the 'mind's…

多媒体 · 计算机科学 2020-12-16 Matthew Joseph Adiletta , Oliver Thomas

The auditory sense of humans is important when it comes to navigation. The importance is especially high in cases when an object of interest is visually partly or fully covered. Interactions with users of technology are mainly focused on…

人机交互 · 计算机科学 2024-04-26 Jan-Niklas Voigt-Antons , Zhirou Sun , Maurizio Vergari , Navid Ashrafi , Francesco Vona , Tanja Kojic

This research presents a proof-of-concept prototype of an all-in-one mixed reality application platform, developed to investigate the needs and expectations of users from mixed reality systems. The study involved an extensive user study…

人机交互 · 计算机科学 2024-04-17 Amir Reza Asadi , Reza Hemadi

In most of practical scenarios, the announcement system must deliver speech messages in a noisy environment, in which the background noise cannot be cancelled out. The local noise reduces speech intelligibility and increases listening…

声音 · 计算机科学 2022-06-28 Tuan Vu Ho , Maori Kobayashi , Masato Akagi
‹ 上一页 1 8 9 10 下一页 ›