English
Related papers

Related papers: On Generating *-Sound Nets with Substitution

200 papers

Graded Type Theory provides a mechanism to track and reason about resource usage in type systems. In this paper, we develop GraD, a novel version of such a graded dependent type system that includes functions, tensor products, additive…

Programming Languages · Computer Science 2021-01-07 Pritam Choudhury , Harley Eades , Richard A. Eisenberg , Stephanie C Weirich

Recent advances in text-to-speech, particularly those based on Graph Neural Networks (GNNs), have significantly improved the expressiveness of short-form synthetic speech. However, generating human-parity long-form speech with high dynamic…

Sound · Computer Science 2023-10-10 Dake Guo , Xinfa Zhu , Liumeng Xue , Tao Li , Yuanjun Lv , Yuepeng Jiang , Lei Xie

Spontaneous conversations in real-world settings such as those found in child-centered recordings have been shown to be amongst the most challenging audio files to process. Nevertheless, building speech processing models handling such a…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-12 Marvin Lavechin , Ruben Bousbib , Hervé Bredin , Emmanuel Dupoux , Alejandrina Cristia

The prosodic aspects of speech signals produced by current text-to-speech systems are typically averaged over training material, and as such lack the variety and liveliness found in natural speech. To avoid monotony and averaged prosody…

Computation and Language · Computer Science 2019-06-05 Vincent Wan , Chun-an Chan , Tom Kenter , Jakub Vit , Rob Clark

Video-conditioned audio generation, including Video-to-Sound (V2S) and Visual Text-to-Speech (VisualTTS), has traditionally been treated as distinct tasks, leaving the potential for a unified generative framework largely underexplored. In…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-23 Xin Cheng , Yuyue Wang , Xihua Wang , Yihan Wu , Kaisi Guan , Yijing Chen , Peng Zhang , Xiaojiang Liu , Meng Cao , Ruihua Song

This paper proposes a neural network that performs audio transformations to user-specified sources (e.g., vocals) of a given audio track according to a given description while preserving other sources not mentioned in the description. Audio…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-29 Woosung Choi , Minseok Kim , Marco A. Martínez Ramírez , Jaehwa Chung , Soonyoung Jung

Topology inference is a powerful tool to better understand the behaviours of network systems (NSs). Different from most of prior works, this paper is dedicated to inferring the directed topology of NSs from noisy observations, where the…

Systems and Control · Electrical Eng. & Systems 2025-09-03 Qing Jiao , Yushan Li , Jianping He

A sound field synthesis method enhancing perceptual quality is proposed. Sound field synthesis using multiple loudspeakers enables spatial audio reproduction with a broad listening area; however, synthesis errors at high frequencies called…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-27 Keisuke Kimura , Shoichi Koyama , Hiroshi Saruwatari

Diffusion models create data from noise by inverting the forward paths of data towards noise and have emerged as a powerful generative modeling technique for high-dimensional, perceptual data such as images and videos. Rectified flow is a…

The popularity of applying machine learning techniques in musical domains has created an inherent availability of freely accessible pre-trained neural network (NN) models ready for use in creative applications. This work outlines the…

Human-Computer Interaction · Computer Science 2020-12-07 Rohan Proctor , Charles Patrick Martin

Alternators have recently been introduced as a framework for modeling time-dependent data. They often outperform other popular frameworks, such as state-space models and diffusion models, on challenging time-series tasks. This paper…

Machine Learning · Computer Science 2025-05-20 Mohammad R. Rezaei , Adji Bousso Dieng

Being able to control the acoustic events (AEs) to which we want to listen would allow the development of more controllable hearable devices. This paper addresses the AE sound selection (or removal) problems, that we define as the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-11 Tsubasa Ochiai , Marc Delcroix , Yuma Koizumi , Hiroaki Ito , Keisuke Kinoshita , Shoko Araki

Polyphonic music generation is still a challenge direction due to its correct between generating melody and harmony. Most of the previous studies used RNN-based models. However, the RNN-based models are hard to establish the relationship…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-08 Jiuyang Zhou , Hong Zhu , Xingping Wang

Numerous models have shown great success in the fields of speech recognition as well as speech synthesis, but models for speech to speech processing have not been heavily explored. We propose Speech to Speech Synthesis Network (STSSN), a…

Sound · Computer Science 2026-02-20 Bjorn Johnson , Jared Levy

We propose WHISPER-GPT: A generative large language model (LLM) for speech and music that allows us to work with continuous audio representations and discrete tokens simultaneously as part of a single architecture. There has been a huge…

Sound · Computer Science 2024-12-20 Prateek Verma

In this paper, we introduce a novel training framework designed to comprehensively address the acoustic howling issue by examining its fundamental formation process. This framework integrates a neural network (NN) module into the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-29 Hao Zhang , Yixuan Zhang , Meng Yu , Dong Yu

High-order wave-making theories are becoming available but are limited to certain ranges of waves and wavemaker types in their applicability. Alternatively, machine learning can be considered to find nonlinear functional relationships.…

Fluid Dynamics · Physics 2022-02-24 Yulin Xie , Xizeng Zhao

Pointer generator networks have been used successfully for abstractive summarization. Along with the capability to generate novel words, it also allows the model to copy from the input text to handle out-of-vocabulary words. In this paper,…

Machine Learning · Computer Science 2019-02-01 Kushal Chawla , Kundan Krishna , Balaji Vasan Srinivasan

In this paper we propose WaveGlow: a flow-based network capable of generating high quality speech from mel-spectrograms. WaveGlow combines insights from Glow and WaveNet in order to provide fast, efficient and high-quality audio synthesis,…

Sound · Computer Science 2018-11-02 Ryan Prenger , Rafael Valle , Bryan Catanzaro

We propose a novel quantum generative model paradigm that fundamentally avoids the issue of extremely small post-selection probabilities present in previous models. Unlike existing methods that require multi-step noise addition and…

Quantum Physics · Physics 2026-05-05 Xin Wang , Rebing Wu