English
Related papers

Related papers: FreGrad: Lightweight and Fast Frequency-aware Diff…

200 papers

In this work, we propose a new mathematical vocoder algorithm(modified spectral inversion) that generates a waveform from acoustic features without phase estimation. The main benefit of using our proposed method is that it excludes the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Hyun Gon Ryu , Jeong-Hoon Kim , Simon See

Recent researches indicate that utilizing the frequency information of input data can enhance the performance of networks. However, the existing popular convolutional structure is not designed specifically for utilizing the frequency…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Zhaowen Li , Xu Zhao , Peigeng Ding , Zongxin Gao , Yuting Yang , Ming Tang , Jinqiao Wang

Attention mechanisms have emerged as important tools that boost the performance of deep models by allowing them to focus on key parts of learned embeddings. However, current attention mechanisms used in speaker recognition tasks fail to…

Sound · Computer Science 2022-07-21 Amirhossein Hajavi , Ali Etemad

In this paper, we show that a simple self-supervised pre-trained audio model can achieve comparable inference efficiency to more complicated pre-trained models with speech transformer encoders. These speech transformers rely on mixing…

Sound · Computer Science 2024-02-09 Sungho Jeon , Ching-Feng Yeh , Hakan Inan , Wei-Ning Hsu , Rashi Rungta , Yashar Mehdad , Daniel Bikel

Ultrashort laser pulses enable attosecond-scale measurements and drive breakthroughs across science and technology, but their routine use hinges on reliable pulse characterization. Frequency-Resolved Optical Gating (FROG) is a leading…

Optics · Physics 2026-04-29 Abhimanyu Borthakur , Jack Eden Hirschman , Sergio Carbajo

Diffusion models have recently advanced photorealistic human synthesis, although practical talking-head generation (THG) remains constrained by high inference latency, temporal instability such as flicker and identity drift, and imperfect…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Soumya Mazumdar , Vineet Kumar Rakesh

The embodied intelligence bridges the physical world and information space. As its typical physical embodiment, humanoid robots have shown great promise through robot learning algorithms in recent years. In this study, a hardware platform,…

Robotics · Computer Science 2025-10-17 Jiaxin Huang , Hanyu Liu , Yunsheng Ma , Jian Shen , Yilin Zheng , Jiayi Wen , Baishu Wan , Pan Li , Zhigong Song

High-quality speech corpora are essential foundations for most speech applications. However, such speech data are expensive and limited since they are collected in professional recording environments. In this work, we propose an…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-11 Haoyu Li , Yang Ai , Junichi Yamagishi

Learning visuomotor policies via behavior cloning typically involves mimicking expert demonstrations collected by human operators. However, natural human demonstrations inherently contain high-frequency noise, such as intermittent jerks,…

Robotics · Computer Science 2026-05-28 Junlin Wang

This paper introduces a lightweight deep learning model for real-time speech enhancement, designed to operate efficiently on resource-constrained devices. The proposed model leverages a compact architecture that facilitates rapid inference…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-23 Shuubham Ojha , Felix Gervits , Carol Espy-Wilson

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-30 Jaemin Jung , Junseok Ahn , Chaeyoung Jung , Tan Dat Nguyen , Youngjoon Jang , Joon Son Chung

Video diffusion models have made substantial progress in various video generation applications. However, training models for long video generation tasks require significant computational and data resources, posing a challenge to developing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Yu Lu , Yuanzhi Liang , Linchao Zhu , Yi Yang

We describe a new convolutional framework for waveform evaluation, WEnets, and build a Narrowband Audio Waveform Evaluation Network, or NAWEnet, using this framework. NAWEnet is single-ended (or no-reference) and was trained three separate…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-20 Andrew A. Catellier , Stephen D. Voran

Speech enhancement (SE) aims to extract the clean waveform from noise-contaminated measurements to improve the speech quality and intelligibility. Although learning-based methods can perform much better than traditional counterparts, the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-23 Haoyin Yan , Jie Zhang , Cunhang Fan , Yeping Zhou , Peiqi Liu

Building a voice conversion system for noisy target speakers, such as users providing noisy samples or Internet found data, is a challenging task since the use of contaminated speech in model training will apparently degrade the conversion…

Sound · Computer Science 2022-07-05 Liumeng Xue , Shan Yang , Na Hu , Dan Su , Lei Xie

We propose FrePolad: frequency-rectified point latent diffusion, a point cloud generation pipeline integrating a variational autoencoder (VAE) with a denoising diffusion probabilistic model (DDPM) for the latent distribution. FrePolad…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Chenliang Zhou , Fangcheng Zhong , Param Hanji , Zhilin Guo , Kyle Fogarty , Alejandro Sztrajman , Hongyun Gao , Cengiz Oztireli

Existing methods for deepfake audio detection have demonstrated some effectiveness. However, they still face challenges in generalizing to new forgery techniques and evolving attack patterns. This limitation mainly arises because the models…

To circumvent the inherent fidelity bottlenecks and optimization misalignment of VAE-based latent diffusion, pixel-space diffusion models have emerged as a compelling end-to-end paradigm. However, existing pixel diffusion models often…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Lichen Ma , Zipeng Guo , Yu He , Xiaolong Fu , Luohang Liu , Jingling Fu , Junshi Huang , Yan Li

It is challenging to accelerate the training process while ensuring both high-quality generated voices and acceptable inference speed. In this paper, we propose a novel neural vocoder called InstructSing, which can converge much faster…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-11 Chang Zeng , Chunhui Wang , Xiaoxiao Miao , Jian Zhao , Zhonglin Jiang , Yong Chen

Despite the latest remarkable advances in generative modeling, efficient generation of high-quality 3D assets from textual prompts remains a difficult task. A key challenge lies in data scarcity: the most extensive 3D datasets encompass…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Antoine Mercier , Ramin Nakhli , Mahesh Reddy , Rajeev Yasarla , Hong Cai , Fatih Porikli , Guillaume Berger