中文
相关论文

相关论文: Acoustic echo suppression using a learning-based m…

200 篇论文

An ideal audio retrieval system efficiently and robustly recognizes a short query snippet from an extensive database. However, the performance of well-known audio fingerprinting systems falls short at high signal distortion levels. This…

音频与语音处理 · 电气工程与系统科学 2024-11-22 Anup Singh , Kris Demuynck , Vipul Arora

Recent work has begun exploring neural acoustic word embeddings---fixed-dimensional vector representations of arbitrary-length speech segments corresponding to words. Such embeddings are applicable to speech retrieval and recognition tasks,…

计算与语言 · 计算机科学 2017-03-14 Wanjia He , Weiran Wang , Karen Livescu

Full-Duplex (FD) Amplify-and-Forward (AF) Multiple-Input Multiple-Output (MIMO) relaying has been the focus of several recent studies, due to the potential for achieving a higher spectral efficiency and lower latency, together with inherent…

信息论 · 计算机科学 2020-08-17 Omid Taghizadeh , Slawomir Stanczak , Hiroki Iimori , Giuseppe Abreu

Recently, self-supervised learning (SSL) techniques have been introduced to solve the monaural speech enhancement problem. Due to the lack of using clean phase information, the enhancement performance is limited in most SSL methods.…

声音 · 计算机科学 2021-12-22 Yi Li , Yang Sun , Syed Mohsen Naqvi

In zero-resource settings where transcribed speech audio is unavailable, unsupervised feature learning is essential for downstream speech processing tasks. Here we compare two recent methods for frame-level acoustic feature learning. For…

计算与语言 · 计算机科学 2020-03-31 Petri-Johan Last , Herman A. Engelbrecht , Herman Kamper

Fine-tuning pretrained language models (LMs) is a popular approach to automatic speech recognition (ASR) error detection during post-processing. While error detection systems often take advantage of statistical language archetypes captured…

计算与语言 · 计算机科学 2021-08-05 Seongmin Park , Dongchan Shin , Sangyoun Paik , Subong Choi , Alena Kazakova , Jihwa Lee

Traditional Active Noise Control (ANC) systems are mostly based on FxLMS algorithms, but such algorithms rely on linear assumptions and are often limited in handling broadband non-stationary noise or nonlinear acoustic paths. Not only that,…

信号处理 · 电气工程与系统科学 2026-04-14 Shuning Dai

Acoustic echo and background noise pose challenges on speech enhancement in hands-free systems and speakerphones. Discriminatively trained end-to-end methods represent a powerful solution for joint acoustic echo control (AEC) and denoising.…

音频与语音处理 · 电气工程与系统科学 2026-05-12 Haljan Lugo Girao , Ernst Seidel , Pejman Mowlaee , Ziyue Zhao , Tim Fingscheidt

The prevalent approach in speech emotion recognition (SER) involves integrating both audio and textual information to comprehensively identify the speaker's emotion, with the text generally obtained through automatic speech recognition…

计算与语言 · 计算机科学 2024-05-29 Jiajun He , Xiaohan Shi , Xingfeng Li , Tomoki Toda

In this paper, we propose a deep-learning framework for environmental sound deepfake detection (ESDD) -- the task of identifying whether the sound scene and sound event in an input audio recording is fake or not. To this end, we conducted…

声音 · 计算机科学 2026-05-04 Lam Pham , Khoi Vu , Dat Tran , Phat Lam , Vu Nguyen , David Fischinger , Son Le

Speech emotion recognition (SER) is an important technology in human-computer interaction. However, achieving high performance is challenging due to emotional complexity and scarce annotated data. To tackle these challenges, we propose a…

声音 · 计算机科学 2026-03-06 Cong Wang , Yizhong Geng , Yuhua Wen , Qifei Li , Yingming Gao , Ruimin Wang , Chunfeng Wang , Hao Li , Ya Li , Wei Chen

In recent years, deep networks have led to dramatic improvements in speech enhancement by framing it as a data-driven pattern recognition problem. In many modern enhancement systems, large amounts of data are used to train a deep network to…

Compression is at the heart of effective representation learning. However, lossy compression is typically achieved through simple parametric models like Gaussian noise to preserve analytic tractability, and the limitations this imposes on…

机器学习 · 计算机科学 2019-11-15 Rob Brekelmans , Daniel Moyer , Aram Galstyan , Greg Ver Steeg

Usually, hearing impaired people use hearing aids which are implemented with speech enhancement algorithms. Estimation of speech and estimation of nose are the components in single channel speech enhancement system. The main objective of…

声音 · 计算机科学 2014-11-10 M. Ravichandra Kumar , B. Ravi Teja

In reverberant conditions with multiple concurrent speakers, each microphone acquires a mixture signal of multiple speakers at a different location. In over-determined conditions where the microphones out-number speakers, we can narrow down…

声音 · 计算机科学 2023-10-31 Zhong-Qiu Wang , Shinji Watanabe

We consider a bidirectional in-band full-duplex (FD) multiple-input multiple-output (MIMO) system subject to imperfect channel state information (CSI), hardware distortion, and limited analog cancellation capability as well as the…

信号处理 · 电气工程与系统科学 2020-06-09 Hiroki Iimori , Giuseppe Thadeu Freitas de Abreu , Koji Ishibashi

Delay-and-Sum (DAS) is the most common algorithm used in photoacoustic (PA) image formation. However, this algorithm results in a reconstructed image with a wide mainlobe and high level of sidelobes. Minimum variance (MV), as an adaptive…

信号处理 · 电气工程与系统科学 2018-05-11 Roya Paridar , Moein Mozaffarzadeh , Mohammad Mehrmohammadi , Mahdi Orooji

We consider the problem of downlink training and channel estimation in frequency division duplex (FDD) massive MIMO systems, where the base station (BS) equipped with a large number of antennas serves a number of single-antenna users…

信息论 · 计算机科学 2016-08-01 Jun Fang , Xingjian Li , Hongbin Li , Feifei Gao

Dysarthria is malfunctioning of motor speech caused by faintness in the human nervous system. It is characterized by the slurred speech along with physical impairment which restricts their communication and creates the lack of confidence…

声音 · 计算机科学 2015-06-09 Megha Rughani , D. Shivakrishna

Speech representation and modelling in high-dimensional spaces of acoustic waveforms, or a linear transformation thereof, is investigated with the aim of improving the robustness of automatic speech recognition to additive noise. The…

计算与语言 · 计算机科学 2015-03-31 Matthew Ager , Zoran Cvetkovic , Peter Sollich
‹ 上一页 1 8 9 10 下一页 ›