中文
相关论文

相关论文: CIPHER: Conformer-based Inference of Phonemes from…

200 篇论文

Reverberant speech, denoting the speech signal degraded by reverberation, contains crucial knowledge of both anechoic source speech and room impulse response (RIR). This work proposes a variational Bayesian inference (VBI) framework with…

音频与语音处理 · 电气工程与系统科学 2025-09-10 Pengyu Wang , Ying Fang , Xiaofei Li

Human Activity Recognition (HAR) with wearable sensors is essential for applications in healthcare, fitness, and human-computer interaction. Bio-impedance sensing offers unique advantages for fine-grained motion capture but remains…

计算机视觉与模式识别 · 计算机科学 2025-07-10 Lala Shakti Swarup Ray , Mengxi Liu , Deepika Gurung , Bo Zhou , Sungho Suh , Paul Lukowicz

Speaker verification is a task of confirming an individual's identity through the analysis of their voice. Whispered speech differs from phonated speech in acoustic characteristics, which degrades the performance of speaker verification…

声音 · 计算机科学 2026-05-08 Magdalena Gołębiowska , Piotr Syga

This paper describes our NPU-ASLP system for the Audio-Visual Diarization and Recognition (AVDR) task in the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. Specifically, the weighted prediction error (WPE) and guided…

音频与语音处理 · 电气工程与系统科学 2023-03-14 Pengcheng Guo , He Wang , Bingshen Mu , Ao Zhang , Peikun Chen

Non-invasive brain-computer interfaces that decode spoken commands from electroencephalogram must be both accurate and trustworthy. We present a confidence-aware decoding framework that couples deep ensembles of compact, speech-oriented…

人工智能 · 计算机科学 2025-11-12 Soowon Kim , Byung-Kwan Ko , Seo-Hyun Lee

Deep learning models have become an increasingly preferred option for biometric recognition systems, such as speaker recognition. SincNet, a deep neural network architecture, gained popularity in speaker recognition tasks due to its…

音频与语音处理 · 电气工程与系统科学 2023-06-13 Labib Chowdhury , Mustafa Kamal , Najia Hasan , Nabeel Mohammed

We present a decoder-only Conformer for automatic speech recognition (ASR) that processes speech and text in a single stack without external speech encoders or pretrained large language models (LLM). The model uses a modality-aware sparse…

音频与语音处理 · 电气工程与系统科学 2026-02-16 Jaeyoung Lee , Masato Mimura

Whispered speech lacks vocal fold vibration and fundamental frequency, resulting in degraded acoustic cues and making whisper-to-normal (W2N) conversion challenging, especially with limited parallel data. We propose WhispEar, a…

声音 · 计算机科学 2026-03-10 Zihao Fang , Yingda Shen , Zifan Guan , Tongtong Song , Zhenyi Liu , Zhizheng Wu

Voiced Electromyography (EMG)-to-Speech (V-ETS) models reconstruct speech from muscle activity signals, facilitating applications such as neurolaryngologic diagnostics. Despite its potential, the advancement of V-ETS is hindered by a…

声音 · 计算机科学 2026-01-13 Xiaodan Chen , Xiaoxue Gao , Mathias Quoy , Alexandre Pitti , Nancy F. Chen

Neural network-based speaker recognition has achieved significant improvement in recent years. A robust speaker representation learns meaningful knowledge from both hard and easy samples in the training set to achieve good performance.…

音频与语音处理 · 电气工程与系统科学 2022-10-31 Ruijie Tao , Kong Aik Lee , Zhan Shi , Haizhou Li

Recently Transformer and Convolution neural network (CNN) based models have shown promising results in Automatic Speech Recognition (ASR), outperforming Recurrent neural networks (RNNs). Transformer models are good at capturing…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Anmol Gulati , James Qin , Chung-Cheng Chiu , Niki Parmar , Yu Zhang , Jiahui Yu , Wei Han , Shibo Wang , Zhengdong Zhang , Yonghui Wu , Ruoming Pang

For speech classification tasks, deep learning models often achieve high accuracy but exhibit shortcomings in calibration, manifesting as classifiers exhibiting overconfidence. The significance of calibration lies in its critical role in…

音频与语音处理 · 电气工程与系统科学 2024-06-27 Yaqian Hao , Chenguang Hu , Yingying Gao , Shilei Zhang , Junlan Feng

In recent years, much speech separation research has focused primarily on improving model performance. However, for low-latency speech processing systems, high efficiency is equally important. Therefore, we propose a speech separation model…

声音 · 计算机科学 2026-03-02 Mohan Xu , Kai Li , Guo Chen , Xiaolin Hu

This study addresses unsupervised subword modeling, i.e., learning acoustic feature representations that can distinguish between subword units of a language. We propose a two-stage learning framework that combines self-supervised learning…

音频与语音处理 · 电气工程与系统科学 2021-06-08 Siyuan Feng , Odette Scharenborg

Neuro-steered speaker extraction aims to extract the listener's brain-attended speech signal from a multi-talker speech signal, in which the attention is derived from the cortical activity. This activity is usually recorded using…

音频与语音处理 · 电气工程与系统科学 2023-12-13 Zexu Pan , Gordon Wichern , Francois G. Germain , Sameer Khurana , Jonathan Le Roux

Recent advances in deep learning have had a methodological and practical impact on brain-computer interface research. Among the various deep network architectures, convolutional neural networks have been well suited for…

信号处理 · 电气工程与系统科学 2020-03-06 Wonjun Ko , Eunjin Jeon , Seungwoo Jeong , Heung-Il Suk

Background: Topological data analysis (TDA) has exploded as a tool for analyzing and making sense of high dimensional datasets across a variety of fields. Mapper is a tool from TDA that captures low-dimensional structure from…

Depression is a global health concern with a critical need for increased patient screening. Speech technology offers advantages for remote screening but must perform robustly across patients. We have described two deep learning models…

音频与语音处理 · 电气工程与系统科学 2024-12-30 Y. Lu , A. Harati , T. Rutowski , R. Oliveira , P. Chlebek , E. Shriberg

In this study, we introduce CIKMar, an efficient approach to educational dialogue systems powered by the Gemma Language model. By leveraging a Dual-Encoder ranking system that incorporates both BERT and SBERT model, we have designed CIKMar…

计算与语言 · 计算机科学 2024-08-19 Joanito Agili Lopo , Marina Indah Prasasti , Alma Permatasari

Most state-of-the-art Deep Learning (DL) approaches for speaker recognition work on a short utterance level. Given the speech signal, these algorithms extract a sequence of speaker embeddings from short segments and those are averaged to…

声音 · 计算机科学 2019-07-03 Miquel India , Pooyan Safari , Javier Hernando