中文
相关论文

相关论文: An Invertible Discrete Auditory Transform

200 篇论文

Speech recognition systems have improved dramatically over the last few years, however, their performance is significantly degraded for the cases of accented or impaired speech. This work explores domain adversarial neural networks (DANN)…

声音 · 计算机科学 2020-10-09 Dominika Woszczyk , Stavros Petridis , David Millard

Discrete audio tokens have recently gained considerable attention for their potential to bridge audio and language processing, enabling multimodal language models that can both generate and understand audio. However, preserving key…

The effect of hearing impairment on speech perception was described by Plomp (1978) as a sum of a loss of class A, due to signal attenuation, and a loss of class D, due to signal distortion. While a loss of class A can be compensated by…

音频与语音处理 · 电气工程与系统科学 2021-02-25 Marc René Schädler

Binaural Audio Telepresence (BAT) aims to encode the acoustic scene at the far end into binaural signals for the user at the near end. BAT encompasses an immense range of applications that can vary between two extreme modes of Immersive BAT…

音频与语音处理 · 电气工程与系统科学 2024-05-15 Yicheng Hsu , Mingsian R. Bai

A judicious combination of dictionary learning methods, block sparsity and source recovery algorithm are used in a hierarchical manner to identify the noises and the speakers from a noisy conversation between two people. Conversations are…

声音 · 计算机科学 2016-10-31 K V Vijay Girish , A G Ramakrishnan , T V Ananthapadmanabha

This paper addresses the problem of binaural localization of a single speech source in noisy and reverberant environments. For a given binaural microphone setup, the binaural response corresponding to the direct-path propagation of a single…

声音 · 计算机科学 2016-09-08 Xiaofei Li , Laurent Girin , Radu Horaud , Sharon Gannot

We discuss the problem of adaptive discrete-time signal denoising in the situation where the signal to be recovered admits a "linear oracle" -- an unknown linear estimate that takes the form of convolution of observations with a…

统计理论 · 数学 2021-02-15 Zaid Harchaoui , Anatoli Juditsky , Arkadi Nemirovski , Dmitrii Ostrovskii

Differentiable Wavetable Synthesis (DWTS) is a technique for neural audio synthesis which learns a dictionary of one-period waveforms i.e. wavetables, through end-to-end training. We achieve high-fidelity audio synthesis with as little as…

声音 · 计算机科学 2022-02-15 Siyuan Shan , Lamtharn Hantrakul , Jitong Chen , Matt Avent , David Trevelyan

Acoustic-to-articulatory inversion (AAI) is to obtain the movement of articulators from speech signals. Until now, achieving a speaker-independent AAI remains a challenge given the limited data. Besides, most current works only use audio…

声音 · 计算机科学 2022-04-05 Jianrong Wang , Jinyu Liu , Longxuan Zhao , Shanyu Wang , Ruiguo Yu , Li Liu

Continuous speech can be converted into a discrete sequence by deriving discrete units from the hidden features of self-supervised learned (SSL) speech models. Although SSL models are becoming larger and trained on more data, they are often…

音频与语音处理 · 电气工程与系统科学 2025-02-06 Jakob Poncelet , Yujun Wang , Hugo Van hamme

Distributed Acoustic Sensing (DAS) converts existing fibre-optic cables into dense seismic arrays at near-zero deployment cost, but measures strain rate rather than particle velocity -- the quantity required by virtually all seismological…

地球物理 · 物理学 2026-05-19 Isao Kurosawa

In this paper, a novel and robust algorithm is proposed for adaptive beamforming based on the idea of reconstructing the autocorrelation sequence (ACS) of a random process from a set of measured data. This is obtained from the first column…

信息论 · 计算机科学 2021-06-25 Saeed Mohammadzadeh , Vitor H. Nascimento , Rodrigo C. de Lamare , Osman Kukrer

A Discrete Fourier Transform Method (DFTM) for discrimination between the signal of neutrons and gamma rays in organic scintillation detectors is presented. The method is based on the transformation of signals into the frequency domain…

仪器与探测器 · 物理学 2016-11-03 M. J. Safari , F. Abbasi Davani , H. Afarideh , S. Jamili , E. Bayat

This paper proposes a deep sound-field denoiser, a deep neural network (DNN) based denoising of optically measured sound-field images. Sound-field imaging using optical methods has gained considerable attention due to its ability to achieve…

信号处理 · 电气工程与系统科学 2023-09-22 Kenji Ishikawa , Daiki Takeuchi , Noboru Harada , Takehiro Moriya

Low-dose computed tomography (LDCT) reduces radiation exposure but suffers from image artifacts and loss of detail due to quantum and electronic noise, potentially impacting diagnostic accuracy. Transformer combined with diffusion models…

图像与视频处理 · 电气工程与系统科学 2025-07-01 Qiqing Liu , Guoquan Wei , Zekun Zhou , Yiyang Wen , Liu Shi , Qiegen Liu

Universal sound separation (USS) is a task to separate arbitrary sounds from an audio mixture. Existing USS systems are capable of separating arbitrary sources, given a few examples of the target sources as queries. However, separating…

音频与语音处理 · 电气工程与系统科学 2023-12-01 Yuzhuo Liu , Xubo Liu , Yan Zhao , Yuanyuan Wang , Rui Xia , Pingchuan Tain , Yuxuan Wang

The Fast Fourier Transform (FFT) is the most efficiently known way to compute the Discrete Fourier Transform (DFT) of an arbitrary n-length signal, and has a computational complexity of O(n log n). If the DFT X of the signal x has only k…

信息论 · 计算机科学 2015-01-05 Sameer Pawar , Kannan Ramchandran

We present a novel algorithm, named the 2D-FFAST, to compute a sparse 2D-Discrete Fourier Transform (2D-DFT) featuring both low sample complexity and low computational complexity. The proposed algorithm is based on mixed concepts from…

信息论 · 计算机科学 2015-09-22 Frank Ong , Sameer Pawar , Kannan Ramchandran

Building on the well-established connection between the Hilbert transform and derivative operators, and motivated by recent developments in complex-step differentiation, we introduce the Complex-Step Integral Transform (CSIT): a generalized…

数值分析 · 数学 2025-12-11 Rafael Abreu , Stephanie Durand , Jochen Kamm , Christine Thomas , Monika Pandey

Discrete audio tokenizers are fundamental to empowering large language models with native audio processing and generation capabilities. Despite recent progress, existing approaches often rely on pretrained encoders, semantic distillation,…