English
Related papers

Related papers: Multi-Scale Sub-Band Constant-Q Transform Discrimi…

200 papers

We present STFTCodec, a novel spectral-based neural audio codec that efficiently compresses audio using Short-Time Fourier Transform (STFT). Unlike waveform-based approaches that require large model capacity and substantial memory…

Sound · Computer Science 2025-03-24 Tao Feng , Zhiyuan Zhao , Yifan Xie , Yuqi Ye , Xiangyang Luo , Xun Guan , Yu Li

Speech classification tasks often require powerful language understanding models to grasp useful features, which becomes problematic when limited training data is available. To attain superior classification performance, we propose to…

Computation and Language · Computer Science 2024-07-26 Nicolae-Catalin Ristea , Andrei Anghel , Radu Tudor Ionescu

This paper presents FastSVC, a light-weight cross-domain singing voice conversion (SVC) system, which can achieve high conversion performance, with inference speed 4x faster than real-time on CPUs. FastSVC uses Conformer-based phoneme…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-25 Songxiang Liu , Yuewen Cao , Na Hu , Dan Su , Helen Meng

Despite recent progress in generative adversarial network (GAN)-based vocoders, where the model generates raw waveform conditioned on acoustic features, it is challenging to synthesize high-fidelity audio for numerous speakers across…

Sound · Computer Science 2023-02-17 Sang-gil Lee , Wei Ping , Boris Ginsburg , Bryan Catanzaro , Sungroh Yoon

Recent years have seen increasing interest in applying deep learning methods to the modeling of guitar amplifiers or effect pedals. Existing methods are mainly based on the supervised approach, requiring temporally-aligned data pairs of…

Speaker verification is to judge the similarity between two unknown voices in an open set, where the ideal speaker embedding should be able to condense discriminant information into a compact utterance-level representation that has small…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-10 Hongyu Wang , Hui Li , Bo Li

This paper presents a computationally efficient approach to blind source separation (BSS) of audio signals, applicable even when there are more sources than microphones (i.e., the underdetermined case). When there are as many sources as…

Sound · Computer Science 2021-01-22 Nobutaka Ito , Rintaro Ikeshita , Hiroshi Sawada , Tomohiro Nakatani

The advent of Large Models marks a new era in machine learning, significantly outperforming smaller models by leveraging vast datasets to capture and synthesize complex patterns. Despite these advancements, the exploration into scaling,…

Sound · Computer Science 2024-02-05 Shijia Liao , Shiyi Lan , Arun George Zachariah

In this paper, we introduce MFCCGAN as a novel speech synthesizer based on adversarial learning that adopts MFCCs as input and generates raw speech waveforms. Benefiting the GAN model capabilities, it produces speech with higher…

Sound · Computer Science 2023-10-26 Mohammad Reza Hasanabadi Majid Behdad Davood Gharavian

The performance of speech processing models trained on clean speech drops significantly in noisy conditions. Training with noisy datasets alleviates the problem, but procuring such datasets is not always feasible. Noisy speech simulation…

Sound · Computer Science 2023-05-23 Leander Melroy Maben , Zixun Guo , Chen Chen , Utkarsh Chudiwal , Chng Eng Siong

Most neural vocoders are limited to one type: either GAN or diffusion-based. While state-of-the-art models like Vocos and WaveNeXt use powerful ConvNeXt-based generators, they have only been used in GAN frameworks and have limited…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-26 Wangzixi Zhou , Takuma Okamoto , Yamato Ohtani , Sakriani Sakti , Hisashi Kawai

Speech quality assessment has been a critical component in many voice communication related applications such as telephony and online conferencing. Traditional intrusive speech quality assessment requires the clean reference of the degraded…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-07 Yuchen Liu , Li-Chia Yang , Alex Pawlicki , Marko Stamenovic

This paper presents FastFit, a novel neural vocoder architecture that replaces the U-Net encoder with multiple short-time Fourier transforms (STFTs) to achieve faster generation rates without sacrificing sample quality. We replaced each…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-19 Won Jang , Dan Lim , Heayoung Park

Quantum computing may offer new approaches for advancing machine learning, including in complex tasks such as anomaly detection in network traffic. In this paper, we introduce a quantum generative adversarial network (QGAN) architecture for…

Machine Learning · Computer Science 2025-05-20 Wajdi Hammami , Soumaya Cherkaoui , Shengrui Wang

Denoising diffusion probabilistic models (DDPMs) and generative adversarial networks (GANs) are popular generative models for neural vocoders. The DDPMs and GANs can be characterized by the iterative denoising framework and adversarial…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-04 Yuma Koizumi , Kohei Yatabe , Heiga Zen , Michiel Bacchiani

In this paper, we address the problem of multichannel speech enhancement in the short-time Fourier transform (STFT) domain. A long short-time memory (LSTM) network takes as input a sequence of STFT coefficients associated with a frequency…

Sound · Computer Science 2020-09-24 Xiaofei LI , Radu Horaud

In this paper, we propose an enhanced triplet method that improves the encoding process of embeddings by jointly utilizing generative adversarial mechanism and multitasking optimization. We extend our triplet encoder with Generative…

Sound · Computer Science 2018-03-28 Wenhao Ding , Liang He

Generative Adversarial Networks (GANs) currently achieve the state-of-the-art sound synthesis quality for pitched musical instruments using a 2-channel spectrogram representation consisting of log magnitude and instantaneous frequency (the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-24 Chitralekha Gupta , Purnima Kamath , Lonce Wyse

Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, which suffer from significant information loss. Even with…

The method used to measure relationships between face embeddings plays a crucial role in determining the performance of face clustering. Existing methods employ the Jaccard similarity coefficient instead of the cosine distance to enhance…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Dafeng Zhang , Yongqi Song , Shizhuo Liu
‹ Prev 1 8 9 10 Next ›