中文
相关论文

相关论文: Perceptually Relevant Preservation of Interaural T…

200 篇论文

Audio-visual target speech extraction, which aims to extract a certain speaker's speech from the noisy mixture by looking at lip movements, has made significant progress combining time-domain speech separation models and visual feature…

多媒体 · 计算机科学 2023-03-07 Zhongweiyang Xu , Xulin Fan , Mark Hasegawa-Johnson

In sound field control applications, it is commonly assumed that one has access to an accurate representation of the sound field in the region of interest. This is a problematic assumption since the reconstruction of a sound field from…

音频与语音处理 · 电气工程与系统科学 2026-05-21 David Sundström , Filip Tronarp , Johan Lindström , Andreas Jakobsson

Infrared and visible image fusion aims to integrate complementary information from co-registered source images to produce a single, informative result. Most learning-based approaches train with a combination of structural similarity loss,…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Kaixuan Yang , Wei Xiang , Zhenshuai Chen , Tong Jin , Yunpeng Liu

Recently, the information content (IC) of predictions from a Generative Infinite-Vocabulary Transformer (GIVT) has been used to model musical expectancy and surprisal in audio. We investigate the effectiveness of such modelling using IC…

声音 · 计算机科学 2025-08-08 Mathias Rose Bjare , Stefan Lattner , Gerhard Widmer

Generative Adversarial Network (GAN) based vocoders are superior in both inference speed and synthesis quality when reconstructing an audible waveform from an acoustic representation. This study focuses on improving the discriminator for…

声音 · 计算机科学 2024-04-29 Yicheng Gu , Xueyao Zhang , Liumeng Xue , Haizhou Li , Zhizheng Wu

In this work, we explore the possibility of decoding Imagined Speech brain waves using machine learning techniques. We propose a covariance matrix of Electroencephalogram channels as input features, projection to tangent space of covariance…

信号处理 · 电气工程与系统科学 2021-05-03 Abhiram Singh , Ashwin Gumaste

Recently, the advance in deep learning has brought a considerable improvement in the end-to-end speech recognition field, simplifying the traditional pipeline while producing promising results. Among the end-to-end models, the connectionist…

音频与语音处理 · 电气工程与系统科学 2022-11-29 Ji Won Yoon , Beom Jun Woo , Sunghwan Ahn , Hyeonseung Lee , Nam Soo Kim

Perceptual audio quality measurement systems algorithmically analyze the output of audio processing systems to estimate possible perceived quality degradation using perceptual models of human audition. In this manner, they save the time and…

音频与语音处理 · 电气工程与系统科学 2023-07-14 Pablo M. Delgado , Jürgen Herre

It is widely known in the machine learning community that class noise can be (and often is) detrimental to inducing a model of the data. Many current approaches use a single, often biased, measurement to determine if an instance is noisy. A…

机器学习 · 统计学 2014-03-11 Michael R. Smith , Tony Martinez

Millimeter-wave (mmWave) radar has emerged as a compact and powerful sensing modality for advanced perception tasks that leverage machine learning. It is particularly effective in scenarios where vision-based sensors fail to capture…

信号处理 · 电气工程与系统科学 2026-02-17 Stefan Hägele , Adam Misik , Eckehard Steinbach

We introduce a time-domain framework for efficient multichannel speech enhancement, emphasizing low latency and computational efficiency. This framework incorporates two compact deep neural networks (DNNs) surrounding a multichannel neural…

声音 · 计算机科学 2024-01-17 Tsun-An Hsieh , Jacob Donley , Daniel Wong , Buye Xu , Ashutosh Pandey

We present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual speech separation is a…

We propose the Fr\'echet Audio Distance (FAD), a novel, reference-free evaluation metric for music enhancement algorithms. We demonstrate how typical evaluation metrics for speech enhancement and blind source separation can fail to…

音频与语音处理 · 电气工程与系统科学 2019-01-18 Kevin Kilgour , Mauricio Zuluaga , Dominik Roblek , Matthew Sharifi

The integration of artificial intelligence into hearing assistance marks a paradigm shift from traditional amplification-based systems to intelligent, context-aware audio processing. This systematic literature review evaluates advances in…

声音 · 计算机科学 2025-08-05 Haris Khan , Shumaila Asif , Hassan Nasir , Kamran Aziz Bhatti , Shahzad Amin Sheikh

In this paper we consider a binaural hearing aid setup, where in addition to the head-mounted microphones an external microphone is available. For this setup, we investigate the performance of several relative transfer function (RTF) vector…

音频与语音处理 · 电气工程与系统科学 2022-05-19 Daniel Fejgin , Simon Doclo

Current multispectral object detection methods often retain extraneous background or noise during feature fusion, limiting perceptual performance. To address this, we propose an innovative feature fusion framework based on cross-modal…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Jifeng Shen , Haibo Zhan , Xin Zuo , Heng Fan , Xiaohui Yuan , Jun Li , Wankou Yang

Knowledge Distillation (KD) is a popular area of research for reducing the size of large models while still maintaining good performance. The outputs of larger teacher models are used to guide the training of smaller student models. Given…

音频与语音处理 · 电气工程与系统科学 2020-09-04 Chun-Chieh Chang , Chieh-Chi Kao , Ming Sun , Chao Wang

Active noise control (ANC) systems are commonly designed to achieve maximal sound reduction regardless of the incident direction of the sound. When desired sound is present, the state-of-the-art methods add a separate system to reconstruct…

音频与语音处理 · 电气工程与系统科学 2023-05-15 Tong Xiao , Buye Xu , Chuming Zhao

Integrated Information Theory (IIT) is an audacious attempt to pin down the abstract, phenomenological experiences of consciousness into a rigorous, mathematical framework. We show that IIT's stance in regards to neuronal noise is…

神经元与认知 · 定量生物学 2021-12-10 Refath Bari

Due to the limited isolation of duplexer's stopband transceivers operating in frequency division duplex (FDD) encounter a leakage of the transmitted signal onto the receiving path. Leakage signal with the combination of the second-order…

最优化与控制 · 数学 2024-06-17 A. A. Degtyarev , N. V. Bakholdin , A. Y. Maslovskiy , S. A. Bakhurin