English
Related papers

Related papers: SpecMaskGIT: Masked Generative Modeling of Audio S…

200 papers

Recent advances in motion diffusion models have enabled spatially controllable text-to-motion generation. However, these models struggle to achieve high-precision control while maintaining high-quality motion generation. To address these…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Ekkasit Pinyoanuntapong , Muhammad Usama Saleem , Korrawe Karunratanakul , Pu Wang , Hongfei Xue , Chen Chen , Chuan Guo , Junli Cao , Jian Ren , Sergey Tulyakov

Accurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications. However, the impact of adjusting the speech rate on speech quality, such as intelligibility…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-07 Sefik Emre Eskimez , Xiaofei Wang , Manthan Thakker , Chung-Hsien Tsai , Canrun Li , Zhen Xiao , Hemin Yang , Zirun Zhu , Min Tang , Jinyu Li , Sheng Zhao , Naoyuki Kanda

Deploying ASR models at an industrial scale poses significant challenges in hardware resource management, especially for long-form transcription tasks where audio may last for hours. Large Conformer models, despite their capabilities, are…

Sound · Computer Science 2025-02-21 Khanh Le , Tuan Vu Ho , Dung Tran , Duc Thanh Chau

Analyzing medical data to find abnormalities is a time-consuming and costly task, particularly for rare abnormalities, requiring tremendous efforts from medical experts. Artificial intelligence has become a popular tool for the automatic…

Real-world audio recordings often contain multiple speakers and various degradations, which limit both the quantity and quality of speech data available for building state-of-the-art speech processing models. Although end-to-end approaches…

Sound · Computer Science 2026-01-27 Kohei Asai , Wataru Nakata , Yuki Saito , Hiroshi Saruwatari

We present self-speculative masked diffusions, a new class of masked diffusion generative models for discrete data that require significantly fewer function evaluations to generate samples. Standard masked diffusion models predict…

Machine Learning · Statistics 2026-03-09 Andrew Campbell , Valentin De Bortoli , Jiaxin Shi , Arnaud Doucet

Generative models have thrived in computer vision, enabling unprecedented image processes. Yet the results in audio remain less advanced. Our project targets real-time sound synthesis from a reduced set of high-level parameters, including…

Sound · Computer Science 2019-06-25 Adrien Bitton , Philippe Esling , Antoine Caillon , Martin Fouilleul

Personalizing a speech synthesis system is a highly desired application, where the system can generate speech with the user's voice with rare enrolled recordings. There are two main approaches to build such a system in recent works: speaker…

Sound · Computer Science 2022-08-01 Sung-Feng Huang , Chyi-Jiunn Lin , Da-Rong Liu , Yi-Chen Chen , Hung-yi Lee

This research presents Muskits-ESPnet, a versatile toolkit that introduces new paradigms to Singing Voice Synthesis (SVS) through the application of pretrained audio models in both continuous and discrete approaches. Specifically, we…

We propose an end-to-end speech synthesizer, Fast DCTTS, that synthesizes speech in real time on a single CPU thread. The proposed model is composed of a carefully-tuned lightweight network designed by applying multiple network reduction…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-02 Minsu Kang , Jihyun Lee , Simin Kim , Injung Kim

Representation and generative learning, as reconstruction-based methods, have demonstrated their potential for mutual reinforcement across various domains. In the field of point cloud processing, although existing studies have adopted…

Computer Vision and Pattern Recognition · Computer Science 2024-08-16 Hongliang Zeng , Ping Zhang , Fang Li , Jiahua Wang , Tingyu Ye , Pengteng Guo

The rapid development of large-scale text-to-speech (TTS) models has led to significant advancements in modeling diverse speaker prosody and voices. However, these models often face issues such as slow inference speeds, reliance on complex…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Yinghao Aaron Li , Xilin Jiang , Cong Han , Nima Mesgarani

Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, which suffer from significant information loss. Even with…

Significant progress has been made in training large generative models for natural language and images. Yet, the advancement of 3D generative models is hindered by their substantial resource demands for training, along with inefficient,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Ka-Hei Hui , Aditya Sanghi , Arianna Rampini , Kamal Rahimi Malekshan , Zhengzhe Liu , Hooman Shayani , Chi-Wing Fu

In speech synthesis and speech enhancement systems, melspectrograms need to be precise in acoustic representations. However, the generated spectrograms are over-smooth, that could not produce high quality synthesized speech. Inspired by…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-04 Leyuan Sheng , Dong-Yan Huang , Evgeniy N. Pavlovskiy

A key challenge in synthesizing audios from silent videos is the inherent trade-off between synthesis quality and inference efficiency in existing methods. For instance, flow matching based models rely on modeling instantaneous velocity,…

Sound · Computer Science 2025-09-09 Xiaoran Yang , Jianxuan Yang , Xinyue Guo , Haoyu Wang , Ningning Pan , Gongping Huang

Recently, phase processing is attracting increasinginterest in speech enhancement community. Some researchersintegrate phase estimations module into speech enhancementmodels by using complex-valued short-time Fourier transform(STFT)…

Sound · Computer Science 2019-01-03 Xingjian Du , Mengyao Zhu , Xuan Shi , Xinpeng Zhang , Wen Zhang , Jingdong Chen

Recent conditional image generation methods produce images of remarkable diversity, fidelity and realism. However, the majority of these methods allow conditioning only on labels or text prompts, which limits their level of control over the…

Computer Vision and Pattern Recognition · Computer Science 2023-02-14 Dina Bashkirova , Jose Lezama , Kihyuk Sohn , Kate Saenko , Irfan Essa

Audio token modeling has become a powerful framework for speech synthesis, with two-stage approaches employing semantic tokens remaining prevalent. In this paper, we aim to simplify this process by introducing a semantic knowledge…

Sound · Computer Science 2024-09-18 Gerard I. Gállego , Roy Fejgin , Chunghsin Yeh , Xiaoyu Liu , Gautam Bhattacharya

Recent advances in text-to-speech (TTS) synthesis, such as Tacotron and WaveRNN, have made it possible to construct a fully neural network based TTS system, by coupling the two components together. Such a system is conceptually simple as it…

‹ Prev 1 3 4 5 6 7 10 Next ›