中文
相关论文

相关论文: Scalable Neural Vocoder from Range-Null Space Deco…

200 篇论文

This work proposes the use of clean speech vocoder parameters as the target for a neural network performing speech enhancement. These parameters have been designed for text-to-speech synthesis so that they both produce high-quality…

音频与语音处理 · 电气工程与系统科学 2019-04-03 Soumi Maiti , Michael I Mandel

Although deep neural networks have achieved remarkable results for the task of semantic segmentation, they usually fail to generalize towards new domains, especially when performing synthetic-to-real adaptation. Such domain shift is…

计算机视觉与模式识别 · 计算机科学 2021-10-07 Adriano Cardace , Pierluigi Zama Ramirez , Samuele Salti , Luigi Di Stefano

In this work, we consider the time-harmonic Maxwell's equations and their numerical solution with a domain decomposition method. As an innovative feature, we propose a feedforward neural network-enhanced approximation of the interface…

数值分析 · 数学 2023-03-07 T. Knoke , S. Kinnewig , S. Beuchler , A. Demircan , U. Morgner , T. Wick

In this work, we propose a deep beamforming framework for speech enhancement in dynamic acoustic environments. The framework learns time-varying beamformer weights from noisy multichannel signals via a deep neural network, guided by a…

音频与语音处理 · 电气工程与系统科学 2026-02-18 Ilai Zaidel , Sharon Gannot

Neural vocoders model the raw audio waveform and synthesize high-quality audio, but even the highly efficient ones, like MB-MelGAN and LPCNet, fail to run real-time on a low-end device like a smartglass. A pure digital signal processing…

声音 · 计算机科学 2024-01-22 Prabhav Agrawal , Thilo Koehler , Zhiping Xiu , Prashant Serai , Qing He

Precise parcellation of functional networks (FNs) of early developing human brain is the fundamental basis for identifying biomarker of developmental disorders and understanding functional development. Resting-state fMRI (rs-fMRI) enables…

神经元与认知 · 定量生物学 2025-03-05 Sovesh Mohapatra , Minhui Ouyang , Shufang Tan , Jianlin Guo , Lianglong Sun , Yong He , Hao Huang

Subspace clustering aims to cluster unlabeled data that lies in a union of low-dimensional linear subspaces. Deep subspace clustering approaches based on auto-encoders have become very popular to solve subspace clustering problems. However,…

机器学习 · 计算机科学 2019-10-15 Shuai Yang , Wenqi Zhu , Yuesheng Zhu

Neural audio codecs have recently enabled high-fidelity reconstruction at high compression rates, especially for speech. However, speech and non-speech audio exhibit fundamentally different spectral characteristics: speech energy…

音频与语音处理 · 电气工程与系统科学 2025-11-11 Haoran Wang , Jiatong Shi , Jinchuan Tian , Bohan Li , Kai Yu , Shinji Watanabe

Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid motion, and limited spatial resolution. However, while rtMRI acquisitions may provide…

Neural channel decoder, as a data-driven channel decoding strategy, has shown very promising improvement on error-correcting capability over the classical methods. However, the success of those deep learning-based decoder comes at the cost…

信息论 · 计算机科学 2026-05-20 Chengwei Zhang , Yifan Du , Siyu Liao

Fourier-encoded implicit neural representations (INRs) have shown strong capability in modeling continuous signals from discrete samples. However, conventional Fourier feature mappings use a fixed set of frequencies over the entire spatial…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Ligen Shi , Jun Qiu , Yuhang Zheng , Zengyu Pang , Chang Liu

To have a superior generalization, a deep learning neural network often involves a large size of training sample. With increase of hidden layers in order to increase learning ability, neural network has potential degradation in accuracy.…

机器学习 · 计算机科学 2019-01-01 Lianfa Li , Ying Fang , Jun Wu , Jinfeng Wang

We introduce CrossNet, a complex spectral mapping approach to speaker separation and enhancement in reverberant and noisy conditions. The proposed architecture comprises an encoder layer, a global multi-head self-attention module, a…

声音 · 计算机科学 2024-03-07 Vahid Ahmadi Kalkhorani , DeLiang Wang

Models for audio source separation usually operate on the magnitude spectrum, which ignores phase information and makes separation performance dependant on hyper-parameters for the spectral front-end. Therefore, we investigate end-to-end…

声音 · 计算机科学 2018-06-11 Daniel Stoller , Sebastian Ewert , Simon Dixon

Neural audio codecs and autoencoders have emerged as versatile models for audio compression, transmission, feature-extraction, and latent-space generation. However, a key limitation is that most are trained to maximize reconstruction…

声音 · 计算机科学 2025-09-10 Dimitrios Bralios , Jonah Casebeer , Paris Smaragdis

Domain adaptation methods aim to bridge the gap between datasets by enabling knowledge transfer across domains, reducing the need for additional expert annotations. However, many approaches struggle with reliability in the target domain, an…

图像与视频处理 · 电气工程与系统科学 2026-05-14 Arnaud Judge , Nicolas Duchateau , Thierry Judge , Roman A. Sandler , Joseph Z. Sokol , Christian Desrosiers , Olivier Bernard , Pierre-Marc Jodoin

Recently deep neural networks have been successfully applied in channel coding to improve the decoding performance. However, the state-of-the-art neural channel decoders cannot achieve high decoding performance and low complexity…

机器学习 · 计算机科学 2021-02-16 Siyu Liao , Chunhua Deng , Miao Yin , Bo Yuan

Multi-organ segmentation in medical imaging remains challenging due to large anatomical variability, complex inter-organ dependencies, and diverse organ scales and shapes. Conventional encoder-decoder architectures often struggle to capture…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Zhuoyi Fang

This paper presents a parametric variational autoencoder-based human target detection and localization framework working directly with the raw analog-to-digital converter data from the frequency modulated continous wave radar. We propose a…

计算机视觉与模式识别 · 计算机科学 2022-07-14 Michael Stephan , Thomas Stadelmayer , Avik Santra , Georg Fischer , Robert Weigel , Fabian Lurz

We present a neural speech codec that challenges the need for complex residual vector quantization (RVQ) stacks by introducing a simpler, single-stage quantization approach. Our method operates directly on the mel-spectrogram, treating it…

声音 · 计算机科学 2025-09-03 Luis Felipe Chary , Miguel Arjona Ramirez
‹ 上一页 1 8 9 10 下一页 ›