中文
相关论文

相关论文: Complex Frequency Domain Linear Prediction: A Tool…

200 篇论文

Studies have shown that in noisy acoustic environments, providing binaural signals to the user of an assistive listening device may improve speech intelligibility and spatial awareness. This paper presents a binaural speech enhancement…

音频与语音处理 · 电气工程与系统科学 2024-03-11 Vikas Tokala , Eric Grinstein , Mike Brookes , Simon Doclo , Jesper Jensen , Patrick A. Naylor

Deep complex U-Net structure and convolutional recurrent network (CRN) structure achieve state-of-the-art performance for monaural speech enhancement. Both deep complex U-Net and CRN are encoder and decoder structures with skip connections,…

声音 · 计算机科学 2024-12-02 Shengkui Zhao , Trung Hieu Nguyen , Bin Ma

Large pre-trained language models (PLMs) have shown remarkable performance across various natural language understanding (NLU) tasks, particularly in low-resource settings. Nevertheless, their potential in Automatic Speech Recognition (ASR)…

计算与语言 · 计算机科学 2023-06-13 Aravind Krishnan , Jesujoba Alabi , Dietrich Klakow

With the increasing demand for audio communication and online conference, ensuring the robustness of Acoustic Echo Cancellation (AEC) under the complicated acoustic scenario including noise, reverberation and nonlinear distortion has become…

声音 · 计算机科学 2022-02-16 Shimin Zhang , Yuxiang Kong , Shubo Lv , Yanxin Hu , Lei Xie

The SpeakerBeam-FE (SBF) method is proposed for speaker extraction. It attempts to overcome the problem of unknown number of speakers in an audio recording during source separation. The mask approximation loss of SBF is sub-optimal, which…

音频与语音处理 · 电气工程与系统科学 2019-03-26 Chenglin Xu , Wei Rao , Eng Siong Chng , Haizhou Li

Concept bottleneck models (CBMs) improve neural network interpretability by introducing an intermediate layer that maps human-understandable concepts to predictions. Recent work has explored the use of vision-language models (VLMs) to…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Xingbo Du , Qiantong Dou , Lei Fan , Rui Zhang

Previous research in speech enhancement has mostly focused on modeling time or time-frequency domain information alone, with little consideration given to the potential benefits of simultaneously modeling both domains. Since these domains…

声音 · 计算机科学 2023-05-16 Feng Dang , Qi Hu , Pengyuan Zhang , Yonghong Yan

Prompt learning has become an efficient paradigm for adapting CLIP to downstream tasks. Compared with traditional fine-tuning, prompt learning optimizes a few parameters yet yields highly competitive results, especially appealing in…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Jianhan Wu , Xiaoyang Qu , Zhangcheng Huang , Jianzong Wang

We introduce Multi-level feature Fusion-based Periodicity Analysis Model (MF-PAM), a novel deep learning-based pitch estimation model that accurately estimates pitch trajectory in noisy and reverberant acoustic environments. Our model…

音频与语音处理 · 电气工程与系统科学 2025-09-11 Woo-Jin Chung , Doyeon Kim , Soo-Whan Chung , Hong-Goo Kang

We propose a learnable mel-frequency cepstral coefficient (MFCC) frontend architecture for deep neural network (DNN) based automatic speaker verification. Our architecture retains the simplicity and interpretability of MFCC-based features…

声音 · 计算机科学 2021-02-23 Xuechen Liu , Md Sahidullah , Tomi Kinnunen

Reverberation conveys critical acoustic cues about the environment, supporting spatial awareness and immersion. For auditory augmented reality (AAR) systems, generating perceptually plausible reverberation in real time remains a key…

音频与语音处理 · 电气工程与系统科学 2025-10-28 Philipp Götz , Gloria Dal Santo , Sebastian J. Schlecht , Vesa Välimäki , Emanuël A. P. Habets

Fake speech detection systems have become a necessity to combat against speech deepfakes. Current systems exhibit poor generalizability on out-of-domain speech samples due to lack to diverse training data. In this paper, we attempt to…

音频与语音处理 · 电气工程与系统科学 2025-08-27 Rishith Sadashiv T N , Abhishek Bedge , Saisha Suresh Bore , Jagabandhu Mishra , Mrinmoy Bhattacharjee , S R Mahadeva Prasanna

Digital predistortion (DPD) is a method commonly used to compensate for the nonlinear effects of power amplifiers (PAs). However, the computational complexity of most DPD algorithms becomes an issue in the downlink of massive multi-user…

信号处理 · 电气工程与系统科学 2022-05-12 Yibo Wu , Ulf Gustavsson , Mikko Valkama , Alexandre Graell i Amat , Henk Wymeersch

Pre-trained language models (PLM) have revolutionized the NLP landscape, achieving stellar performances across diverse tasks. These models, while benefiting from vast training data, often require fine-tuning on specific data to cater to…

计算与语言 · 计算机科学 2023-10-04 Jingwei Sun , Ziyue Xu , Hongxu Yin , Dong Yang , Daguang Xu , Yiran Chen , Holger R. Roth

A novel text-independent speaker identification (SI) method is proposed. This method uses the Mel-frequency Cepstral coefficients (MFCCs) and the dynamic information among adjacent frames as feature sets to capture speaker's…

声音 · 计算机科学 2020-02-04 Zhanyu Ma , Hong Yu

The growing scarcity of spectrum resources and rapid proliferation of wireless devices make efficient radio network management critical. While deep learning-enhanced Cognitive Radio Technology (CRT) provides promising solutions for tasks…

信号处理 · 电气工程与系统科学 2025-05-14 Shuai Chen , Yong Zu , Zhixi Feng , Shuyuan Yang , Mengchang Li

Cognitive Language Processing (CLP), situated at the intersection of Natural Language Processing (NLP) and cognitive science, plays a progressively pivotal role in the domains of artificial intelligence, cognitive intelligence, and brain…

机器学习 · 计算机科学 2024-06-06 Weiguo Chen , Changjian Wang , Kele Xu , Yuan Yuan , Yanru Bai , Dongsong Zhang

In this paper, we propose an effective and robust method for acoustic scene analysis based on spatial information extracted from partially synchronized and/or closely located distributed microphones. In the proposed method, to extract…

声音 · 计算机科学 2018-07-10 Keisuke Imoto

Low-Earth-orbit (LEO) satellites and vehicle-to-everything (V2X) networks are driving integrated communication and navigation (ICAN) toward next-generation intelligent transportation. Affine frequency division multiplexing (AFDM) is a…

信号处理 · 电气工程与系统科学 2026-05-21 Zhenyu Chen , Ke Xiao , Xiaomei Tang , Jing Lei , Muzi Yuan , Guangfu Sun

Syntactic structures used to play a vital role in natural language processing (NLP), but since the deep learning revolution, NLP has been gradually dominated by neural models that do not consider syntactic structures in their design. One…

计算与语言 · 计算机科学 2023-11-28 Haoyi Wu , Kewei Tu