中文
相关论文

相关论文: Gen-A: Generalizing Ambisonics Neural Encoding to …

200 篇论文

Deep neural networks (DNNs) have achieved remarkable success across diverse domains, but their performance can be severely degraded by noisy or corrupted training data. Conventional noise mitigation methods often rely on explicit…

机器学习 · 计算机科学 2025-06-16 Deliang Jin , Gang Chen , Shuo Feng , Yufeng Ling , Haoran Zhu

State-of-the-art stereo matching networks have difficulties in generalizing to new unseen environments due to significant domain differences, such as color, illumination, contrast, and texture. In this paper, we aim at designing a…

计算机视觉与模式识别 · 计算机科学 2019-12-02 Feihu Zhang , Xiaojuan Qi , Ruigang Yang , Victor Prisacariu , Benjamin Wah , Philip Torr

Binaural rendering of ambisonic signals is of broad interest to virtual reality and immersive media. Conventional methods often require manually measured Head-Related Transfer Functions (HRTFs). To address this issue, we collect a paired…

声音 · 计算机科学 2022-11-07 Yin Zhu , Qiuqiang Kong , Junjie Shi , Shilei Liu , Xuzhou Ye , Ju-chiang Wang , Junping Zhang

In this paper, we propose an end-to-end deep learning-based joint transceiver design algorithm for millimeter wave (mmWave) massive multiple-input multiple-output (MIMO) systems, which consists of deep neural network (DNN)-aided pilot…

信息论 · 计算机科学 2021-10-27 Qiyu Hu , Yunlong Cai , Kai Kang , Guanding Yu , Jakob Hoydis , Yonina C. Eldar

Deep neural networks have been used widely to learn the latent structure of datasets, across modalities such as images, shapes, and audio signals. However, existing models are generally modality-dependent, requiring custom architectures and…

机器学习 · 计算机科学 2021-11-12 Yilun Du , Katherine M. Collins , Joshua B. Tenenbaum , Vincent Sitzmann

This paper describes multichannel speech enhancement for improving automatic speech recognition (ASR) in noisy environments. Recently, the minimum variance distortionless response (MVDR) beamforming has widely been used because it works…

We propose a spatial diffuseness feature for deep neural network (DNN)-based automatic speech recognition to improve recognition accuracy in reverberant and noisy environments. The feature is computed in real-time from multiple microphone…

计算与语言 · 计算机科学 2015-09-02 Andreas Schwarz , Christian Huemmer , Roland Maas , Walter Kellermann

We propose a deep beamforming framework for enhancing target speaker(s) in multi-speaker environments. A deep neural network (DNN) is trained to estimate beamforming weights directly from noisy multichannel inputs while satisfying linear…

音频与语音处理 · 电气工程与系统科学 2026-05-21 Ilai Zaidel , Ori Engel , Bar Engel , Sharon Gannot

Acoustic beamformers have been widely used to enhance audio signals. Currently, the best methods are the deep neural network (DNN)-powered variants of the generalized eigenvalue and minimum-variance distortionless response beamformers and…

音频与语音处理 · 电气工程与系统科学 2020-08-12 Yuichiro Koyama , Bhiksha Raj

As is expressed in the adage "a picture is worth a thousand words", when using spoken language to communicate visual information, brevity can be a challenge. This work describes a novel technique for leveraging machine-learned feature…

音频与语音处理 · 电气工程与系统科学 2021-08-27 Andrew Port , Chelhwon Kim , Mitesh Patel

Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart…

音频与语音处理 · 电气工程与系统科学 2024-01-22 Ju Lin , Niko Moritz , Yiteng Huang , Ruiming Xie , Ming Sun , Christian Fuegen , Frank Seide

Deep neural networks (DNNs) trained for image denoising are able to generate high-quality samples with score-based reverse diffusion algorithms. These impressive capabilities seem to imply an escape from the curse of dimensionality, but…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Zahra Kadkhodaie , Florentin Guth , Eero P. Simoncelli , Stéphane Mallat

In recent years, deep learning methods applying unsupervised learning to train deep layers of neural networks have achieved remarkable results in numerous fields. In the past, many genetic algorithms based methods have been successfully…

神经与进化计算 · 计算机科学 2017-11-22 Eli David , Iddo Greental

Recent vision-language pre-training models have exhibited remarkable generalization ability in zero-shot recognition tasks. Previous open-vocabulary 3D scene understanding methods mostly focus on training 3D models using either image or…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Ruihuang Li , Zhengqiang Zhang , Chenhang He , Zhiyuan Ma , Vishal M. Patel , Lei Zhang

Environmental audio tagging aims to predict only the presence or absence of certain acoustic events in the interested acoustic scene. In this paper we make contributions to audio tagging in two parts, respectively, acoustic modeling and…

Conventional vocoders are commonly used as analysis tools to provide interpretable features for downstream tasks such as speech synthesis and voice conversion. They are built under certain assumptions about the signals following signal…

音频与语音处理 · 电气工程与系统科学 2021-10-14 Sergey Nikonorov , Berrak Sisman , Mingyang Zhang , Haizhou Li

In many scientific applications, measured time series are corrupted by noise or distortions. Traditional denoising techniques often fail to recover the signal of interest, particularly when the signal-to-noise ratio is low or when certain…

机器学习 · 计算机科学 2022-11-02 Natalie Klein , Amber J. Day , Harris Mason , Michael W. Malone , Sinead A. Williamson

In this article, we use deep neural networks (DNNs) to develop a wireless end-to-end communication system, in which DNNs are employed for all signal-related functionalities, such as encoding, decoding, modulation, and equalization. However,…

信息论 · 计算机科学 2018-07-03 Hao Ye , Geoffrey Ye Li , Biing-Hwang Fred Juang , Kathiravetpillai Sivanesan

Multichannel processing is widely used for speech enhancement but several limitations appear when trying to deploy these solutions to the real-world. Distributed sensor arrays that consider several devices with a few microphones is a viable…

声音 · 计算机科学 2020-03-17 Nicolas Furnon , Romain Serizel , Irina Illina , Slim Essid

Recently, deep learning (DL) has been emerging as a promising approach for channel estimation and signal detection in wireless communications. The majority of the existing studies investigating the use of DL techniques in this domain focus…

网络与互联网体系结构 · 计算机科学 2024-04-04 Khalid Albagami , Nguyen Van Huynh , Geoffrey Ye Li