中文
相关论文

相关论文: mdctGAN: Taming transformer-based GAN for speech s…

200 篇论文

Effective speech representations for spoken language models must balance semantic relevance with acoustic fidelity for high-quality reconstruction. However, existing approaches struggle to achieve both simultaneously. To address this, we…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Amir Hussein , Sameer Khurana , Gordon Wichern , Francois G. Germain , Jonathan Le Roux

Compressed sensing (CS) provides an elegant framework for recovering sparse signals from compressed measurements. For example, CS can exploit the structure of natural images and recover an image from only a few random measurements. CS is…

机器学习 · 计算机科学 2019-05-21 Yan Wu , Mihaela Rosca , Timothy Lillicrap

Computed Tomography (CT) enables detailed cross-sectional imaging but continues to face challenges in balancing reconstruction quality and computational efficiency. While deep learning-based methods have significantly improved image quality…

图像与视频处理 · 电气工程与系统科学 2025-10-23 Shaokai Wu , Yuxiang Lu , Yapan Guo , Wei Ji , Suizhi Huang , Fengyu Yang , Shalayiding Sirejiding , Qichen He , Jing Tong , Yanbiao Ji , Yue Ding , Hongtao Lu

Generative adversarial network (GAN) for image super-resolution (SR) has attracted enormous interests in recent years. However, the GAN-based SR methods only use image discriminator to distinguish SR images and high-resolution (HR) images.…

计算机视觉与模式识别 · 计算机科学 2020-03-18 Xuan Zhu , Yue Cheng , Jinye Peng , Rongzhi Wang , Mingnan Le , Xin Liu

Deep learning is at the core of recent spoken language understanding (SLU) related tasks. More precisely, deep neural networks (DNNs) drastically increased the performances of SLU systems, and numerous architectures have been proposed. In…

计算与语言 · 计算机科学 2019-05-07 Titouan Parcollet , Mohamed Morchid , Xavier Bost , Georges Linarès

Despite the breakthroughs in accuracy and speed of single image super-resolution using faster and deeper convolutional neural networks, one central problem remains largely unsolved: how do we recover the finer texture details when we…

We propose Universal MelGAN, a vocoder that synthesizes high-fidelity speech in multiple domains. To preserve sound quality when the MelGAN-based structure is trained with a dataset of hundreds of speakers, we added multi-resolution…

音频与语音处理 · 电气工程与系统科学 2021-03-05 Won Jang , Dan Lim , Jaesam Yoon

Medical image synthesis is a challenging task due to the scarcity of paired data. Several methods have applied CycleGAN to leverage unpaired data, but they often generate inaccurate mappings that shift the anatomy. This problem is further…

图像与视频处理 · 电气工程与系统科学 2023-08-02 Minh Hieu Phan , Zhibin Liao , Johan W. Verjans , Minh-Son To

This paper proposes an efficient reconfigurable hardware design for speech enhancement based on multi band spectral subtraction algorithm and involving both magnitude and phase components. Our proposed design is novel as it estimates…

声音 · 计算机科学 2015-08-26 Tanmay Biswas , Sudhindu Bikash Mandal , Debasree Saha , Amlan Chakrabarti

Multi-talker speech recognition (MTASR) faces unique challenges in disentangling and transcribing overlapping speech. To address these challenges, this paper investigates the role of Connectionist Temporal Classification (CTC) in speaker…

音频与语音处理 · 电气工程与系统科学 2025-01-06 Jiawen Kang , Lingwei Meng , Mingyu Cui , Yuejiao Wang , Xixin Wu , Xunying Liu , Helen Meng

Non-autoregressive (NAR) text-to-speech synthesis relies on length alignment between text sequences and audio representations, constraining naturalness and expressiveness. Existing methods depend on duration modeling or pseudo-alignment…

Large-scale text-to-speech (TTS) systems are limited by the scarcity of clean, multilingual recordings. We introduce Sidon, a fast, open-source speech restoration model that converts noisy in-the-wild speech into studio-quality speech and…

声音 · 计算机科学 2026-01-27 Wataru Nakata , Yuki Saito , Yota Ueda , Hiroshi Saruwatari

Dysarthric speech reconstruction (DSR), which aims to improve the quality of dysarthric speech, remains a challenge, not only because we need to restore the speech to be normal, but also must preserve the speaker's identity. The speaker…

音频与语音处理 · 电气工程与系统科学 2022-02-21 Disong Wang , Songxiang Liu , Xixin Wu , Hui Lu , Lifa Sun , Xunying Liu , Helen Meng

Recent convolutional object detectors exploit multi-scale feature representations added with top-down pathway in order to detect objects at different scales and learn stronger semantic feature responses. In general, during the top-down…

计算机视觉与模式识别 · 计算机科学 2020-11-18 Seong-Ho Lee , Seung-Hwan Bae

This paper presents a super-resolution (SR) method with unpaired training dataset of clinical CT and micro CT volumes. For obtaining very detailed information such as cancer invasion from pre-operative clinical CT volumes of lung cancer…

计算机视觉与模式识别 · 计算机科学 2020-04-08 Tong ZHENG , Hirohisa ODA , Takayasu MORIYA , Takaaki SUGINO , Shota NAKAMURA , Masahiro ODA , Masaki MORI , Hirotsugu TAKABATAKE , Hiroshi NATORI , Kensaku MORI

In this paper, we consider the problem of super-resolution recons-truction. This is a hot topic because super-resolution reconstruction has a wide range of applications in the medical field, remote sensing monitoring, and criminal…

图像与视频处理 · 电气工程与系统科学 2019-07-25 Qi Zhang , Huafeng Wang , Sichen Yang

Neural audio codecs optimized for mel-spectrogram reconstruction often fail to preserve intelligibility. While semantic encoder distillation improves encoded representations, it does not guarantee content preservation in reconstructed…

音频与语音处理 · 电气工程与系统科学 2026-03-09 Junhyeok Lee , Xiluo He , Jihwan Lee , Helin Wang , Shrikanth Narayanan , Thomas Thebaud , Laureano Moro-Velazquez , Jesús Villalba , Najim Dehak

This paper introduces a new synthesis-based defense algorithm for counteracting with a varieties of adversarial attacks developed for challenging the performance of the cutting-edge speech-to-text transcription systems. Our algorithm…

声音 · 计算机科学 2022-10-26 Mohammad Esmaeilpour , Nourhene Chaalia , Patrick Cardinal

We propose Cotatron, a transcription-guided speech encoder for speaker-independent linguistic representation. Cotatron is based on the multispeaker TTS architecture and can be trained with conventional TTS datasets. We train a voice…

音频与语音处理 · 电气工程与系统科学 2020-08-17 Seung-won Park , Doo-young Kim , Myun-chul Joe

Self-supervised learning (SSL) has advanced speech processing. However, existing speech SSL methods typically assume a single sampling rate and struggle with mixed-rate data due to temporal resolution mismatch. To address this limitation,…

声音 · 计算机科学 2026-03-25 Zikang Huang , Meng Ge , Tianrui Wang , Xuanchen Li , Xiaobao Wang , Longbiao Wang , Jianwu Dang
‹ 上一页 1 8 9 10 下一页 ›