中文
相关论文

相关论文: CIPHER: Conformer-based Inference of Phonemes from…

200 篇论文

The predominant metric for evaluating speech recognizers, the Word Error Rate (WER) has been extended in different ways to handle transcripts produced by long-form multi-talker speech recognizers. These systems process long transcripts…

音频与语音处理 · 电气工程与系统科学 2025-08-05 Thilo von Neumann , Christoph Boeddeker , Marc Delcroix , Reinhold Haeb-Umbach

We propose a mixed deep neural network strategy, incorporating parallel combination of Convolutional (CNN) and Recurrent Neural Networks (RNN), cascaded with deep autoencoders and fully connected layers towards automatic identification of…

机器学习 · 计算机科学 2019-04-10 Pramit Saha , Sidney Fels

A crucial part of an accurate and reliable spoken language assessment system is the underlying ASR model. Recently, large-scale pre-trained ASR foundation models such as Whisper have been made available. As the output of these models is…

计算与语言 · 计算机科学 2023-10-11 Rao Ma , Mengjie Qian , Mark J. F. Gales , Kate M. Knill

Brain decoding has emerged as a rapidly advancing and extensively utilized technique within neuroscience. This paper centers on the application of raw electroencephalogram (EEG) signals for decoding human brain activity, offering a more…

机器学习 · 计算机科学 2025-02-04 Zenon Lamprou , Yashar Moshfeghi

When recorded in an enclosed room, a sound signal will most certainly get affected by reverberation. This not only undermines audio quality, but also poses a problem for many human-machine interaction technologies that use speech as their…

声音 · 计算机科学 2018-09-21 Francisco Ibarrola , Leandro Di Persia , Ruben Spies

Cochlear implants(CIs) are arguably the most successful neural implant, having restored hearing to over one million people worldwide. While CI research has focused on modeling the cochlear activations in response to low-level acoustic…

神经与进化计算 · 计算机科学 2024-07-31 Cynthia R. Steinhardt , Menoua Keshishian , Nima Mesgarani , Kim Stachenfeld

Lipreading has a lot of potential applications such as in the domain of surveillance and video conferencing. Despite this, most of the work in building lipreading systems has been limited to classifying silent videos into classes…

音频与语音处理 · 电气工程与系统科学 2019-07-03 Yaman Kumar , Rohit Jain , Khwaja Mohd. Salik , Rajiv Ratn Shah , Yifang yin , Roger Zimmermann

Recent autoregressive transformer-based speech enhancement (SE) methods have shown promising results by leveraging advanced semantic understanding and contextual modeling of speech. However, these approaches often rely on complex…

声音 · 计算机科学 2025-10-03 Luca A. Lanzendörfer , Frédéric Berdoz , Antonis Asonitis , Roger Wattenhofer

We propose a method for joint multichannel speech dereverberation with two spatial-aware tasks: direction-of-arrival (DOA) estimation and speech separation. The proposed method addresses involved tasks as a sequence to sequence mapping…

音频与语音处理 · 电气工程与系统科学 2020-10-23 Yang Jiao

Clinical EEG interpretation requires reasoning over full EEG sessions and integrating signal patterns with clinical context. Existing EEG foundation models are largely designed for short-window decoding and do not incorporate clinical…

人工智能 · 计算机科学 2026-05-12 Peng Cao , Ali Mirzazadeh , Jong Woo Lee , Aleksandar Videnovic , Dina Katabi

In traditional speaker diarization systems, a well-trained speaker model is a key component to extract representations from consecutive and partially overlapping segments in a long speech session. To be more consistent with the back-end…

声音 · 计算机科学 2022-04-01 Yu-Huai Peng , Hung-Shin Lee , Pin-Tuan Huang , Hsin-Min Wang

Classroom environments are particularly challenging for children with hearing impairments, where background noise, multiple talkers, and reverberation degrade speech perception. These difficulties are greater for children than adults, yet…

UniSpeech has achieved superior performance in cross-lingual automatic speech recognition (ASR) by explicitly aligning latent representations to phoneme units using multi-task self-supervised learning. While the learned representations…

音频与语音处理 · 电气工程与系统科学 2023-10-10 Hongfei Xue , Qijie Shao , Peikun Chen , Pengcheng Guo , Lei Xie , Jie Liu

Recent speech-to-speech (S2S) models generate intelligible speech but still lack natural expressiveness, largely due to the absence of a reliable evaluation metric. Existing approaches, such as subjective MOS ratings, low-level acoustic…

声音 · 计算机科学 2025-10-24 Zhiyu Lin , Jingwen Yang , Jiale Zhao , Meng Liu , Sunzhu Li , Benyou Wang

In this paper, we present an improved model for voicing silent speech, where audio is synthesized from facial electromyography (EMG) signals. To give our model greater flexibility to learn its own input features, we directly use EMG signals…

音频与语音处理 · 电气工程与系统科学 2021-06-22 David Gaddy , Dan Klein

The P300 speller is a brain-computer interface that enables people with neuromuscular disorders to communicate based on eliciting event-related potentials (ERP) in electroencephalography (EEG) measurements. One challenge to reliable…

信息论 · 计算机科学 2017-01-13 Vaishakhi Mayya , Boyla Mainsah , Galen Reeves

Electroencephalography (EEG) foundation models hold significant promise for universal Brain-Computer Interfaces (BCIs). However, existing approaches often rely on end-to-end fine-tuning and exhibit limited efficacy under frozen-probing…

机器学习 · 计算机科学 2026-03-20 Jiquan Wang , Sha Zhao , Yangxuan Zhou , Yiming Kang , Shijian Li , Gang Pan

Understanding of neuro-dynamics of a complex higher cognitive process, Working Memory (WM) is challenging. In WM, information processing occurs through four subsystems: phonological loop, visual sketch pad, memory buffer and central…

信号处理 · 电气工程与系统科学 2020-03-13 Pankaj , Jamuna Rajeswaran , Divya Sadana

On-device end-to-end speech recognition poses a high requirement on model efficiency. Most prior works improve the efficiency by reducing model sizes. We propose to reduce the complexity of model architectures in addition to model sizes.…

计算与语言 · 计算机科学 2020-11-12 Peidong Wang , DeLiang Wang

Following the rationale of end-to-end modeling, CTC, RNN-T or encoder-decoder-attention models for automatic speech recognition (ASR) use graphemes or grapheme-based subword units based on e.g. byte-pair encoding (BPE). The mapping from…

音频与语音处理 · 电气工程与系统科学 2021-04-16 Mohammad Zeineldeen , Albert Zeyer , Wei Zhou , Thomas Ng , Ralf Schlüter , Hermann Ney