中文
相关论文

相关论文: DnR-nonverbal: Cinematic Audio Source Separation D…

200 篇论文

Noisy situations cause huge problems for suffers of hearing loss as hearing aids often make the signal more audible but do not always restore the intelligibility. In noisy settings, humans routinely exploit the audio-visual (AV) nature of…

声音 · 计算机科学 2019-09-24 Mandar Gogate , Kia Dashtipour , Ahsan Adeel , Amir Hussain

Dysarthria is a disability that causes a disturbance in the human speech system and reduces the quality and intelligibility of a person's speech. Because of this effect, the normal speech processing systems can not work properly on impaired…

音频与语音处理 · 电气工程与系统科学 2024-03-22 Aref Farhadipour , Hadi Veisi

Machine learning approaches often require training and evaluation datasets with a clear separation between positive and negative examples. This risks simplifying and even obscuring the inherent subjectivity present in many tasks. Preserving…

Automatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, particularly in non-English languages. To address this, we fine-tune a voice conversion model on English dysarthric speech (UASpeech) to…

This document provides a brief description of the National Institute of Standards and Technology (NIST) speaker recognition evaluation (SRE) conversational telephone speech (CTS) Superset. The CTS Superset has been created in an attempt to…

声音 · 计算机科学 2021-08-17 Seyed Omid Sadjadi

Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to leverage visual information to improve speech recognition…

声音 · 计算机科学 2026-01-27 Junli Chen , Changli Tang , Yixuan Li , Guangzhi Sun , Chao Zhang

During the Covid, online meetings have become an indispensable part of our lives. This trend is likely to continue due to their convenience and broad reach. However, background noise from other family members, roommates, office-mates not…

声音 · 计算机科学 2022-07-22 Wei Sun , Mei Wang , Lili Qiu

We present DRES: a 1.5-hour Dutch realistic elicited (semi-spontaneous) speech dataset from 80 speakers recorded in noisy, public indoor environments. DRES was designed as a test set for the evaluation of state-of-the-art (SOTA) automatic…

音频与语音处理 · 电气工程与系统科学 2026-03-11 Dimme de Groot , Yuanyuan Zhang , Jorge Martinez , Odette Scharenborg

We release the USC Long Single-Speaker (LSS) dataset containing real-time MRI video of the vocal tract dynamics and simultaneous audio obtained during speech production. This unique dataset contains roughly one hour of video and audio data…

We introduce Dataset Concealment (DSC), a rigorous new procedure for evaluating and interpreting objective speech quality estimation models. DSC quantifies and decomposes the performance gap between research results and real-world…

音频与语音处理 · 电气工程与系统科学 2026-01-30 Jaden Pieper , Stephen D. Voran

We present an approach to synthesize whisper by applying a handcrafted signal processing recipe and Voice Conversion (VC) techniques to convert normally phonated speech to whispered speech. We investigate using Gaussian Mixture Models (GMM)…

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here speech enhancement methods have traditionally allowed improved…

音频与语音处理 · 电气工程与系统科学 2020-08-28 Yanpei Shi , Qiang Huang , Thomas Hain

Accurate alignment of dysfluent speech with intended text is crucial for automating the diagnosis of neurodegenerative speech disorders. Traditional methods often fail to model phoneme similarities effectively, limiting their performance.…

Dialogue state tracking plays a crucial role in extracting information in task-oriented dialogue systems. However, preceding research are limited to textual modalities, primarily due to the shortage of authentic human audio datasets. We…

声音 · 计算机科学 2023-12-05 Jihyun Lee , Yejin Jeon , Wonjun Lee , Yunsu Kim , Gary Geunbae Lee

Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size. The goal of this paper is to generate a large scale text-independent…

声音 · 计算机科学 2020-11-05 Arsha Nagrani , Joon Son Chung , Andrew Zisserman

Extracting the speech of a target speaker from mixed audios, based on a reference speech from the target speaker, is a challenging yet powerful technology in speech processing. Recent studies of speaker-independent speech separation, such…

音频与语音处理 · 电气工程与系统科学 2020-10-27 Zining Zhang , Bingsheng He , Zhenjie Zhang

In this paper, SER_AMPEL, a multi-source dataset for speech emotion recognition (SER) is presented. The peculiarity of the dataset is that it is collected with the aim of providing a reference for speech emotion recognition in case of…

音频与语音处理 · 电气工程与系统科学 2023-12-15 Alessandra Grossi , Francesca Gasparini

We address the challenge of making spatial audio datasets by proposing a shared mechanized recording space that can run custom acoustic experiments: a Mechatronic Acoustic Research System (MARS). To accommodate a wide variety of…

音频与语音处理 · 电气工程与系统科学 2023-10-03 Austin Lu , Ethaniel Moore , Arya Nallanthighall , Kanad Sarkar , Manan Mittal , Ryan M. Corey , Paris Smaragdis , Andrew Singer

Natural Language Processing has recently made understanding human interaction easier, leading to improved sentimental analysis and behaviour prediction. However, the choice of words and vocal cues in conversations presents an underexplored…

计算机与社会 · 计算机科学 2022-06-24 Amna Anwar , Eiman Kanjo , Dario Ortega Anderez

This paper proposes a neural network that performs audio transformations to user-specified sources (e.g., vocals) of a given audio track according to a given description while preserving other sources not mentioned in the description. Audio…

音频与语音处理 · 电气工程与系统科学 2021-04-29 Woosung Choi , Minseok Kim , Marco A. Martínez Ramírez , Jaehwa Chung , Soonyoung Jung