English
Related papers

Related papers: Simi-SFX: A similarity-based conditioning method f…

200 papers

Recent MIDI-to-audio synthesis methods using deep neural networks have successfully generated high-quality, expressive instrumental tracks. However, these methods require MIDI annotations for supervised training, limiting the diversity of…

Sound · Computer Science 2025-06-12 Osamu Take , Taketo Akama

Controllable Singing Voice Synthesis (SVS) aims to generate expressive singing voices reflecting user intent. While recent SVS systems achieve high audio quality, most rely on probabilistic modeling, limiting precise control over attributes…

Sound · Computer Science 2025-09-10 Yerin Ryu , Inseop Shin , Chanwoo Kim

This Ph.D. thesis focuses on developing a system for high-quality speech synthesis and voice conversion. Vocoder-based speech analysis, manipulation, and synthesis plays a crucial role in various kinds of statistical parametric speech…

Sound · Computer Science 2021-01-26 Mohammed Salah Al-Radhi

Even though they differ in the physical domain, digital video and audio share many characteristics. Both are temporal data streams often stored in buffers with 8-bit values. This paper investigates a method for creating harmonic sounds with…

Human-Computer Interaction · Computer Science 2016-03-01 Carl Thomé

Music performance synthesis aims to synthesize a musical score into a natural performance. In this paper, we borrow recent advances in text-to-speech synthesis and present the Deep Performer -- a novel system for score-to-audio music…

Sound · Computer Science 2022-02-22 Hao-Wen Dong , Cong Zhou , Taylor Berg-Kirkpatrick , Julian McAuley

Deep neural networks have shown promise for music audio signal processing applications, often surpassing prior approaches, particularly as end-to-end models in the waveform domain. Yet results to date have tended to be constrained by low…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-11 William Mitchell , Scott H. Hawley

Generating controllable indoor scenes is fundamental to applications in game development, architectural visualization, and embodied AI. However, existing approaches either support a limited input modalities or rely on implicit generation…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Wentang Chen , Shougao Zhang , Yiman Zhang , Tianhao Zhou , Ruihui Li

Texture synthesis techniques based on matching the Gram matrix of feature activations in neural networks have achieved spectacular success in the image domain. In this paper we extend these techniques to the audio domain. We demonstrate…

Sound · Computer Science 2018-06-22 Joseph Antognini , Matt Hoffman , Ron J. Weiss

Parametric sound field synthesis methods, such as the Spatial Decomposition Method (SDM) and Higher-Order Spatial Impulse Response Rendering (HO-SIRR), are widely used for the analysis and auralization of sound fields. This paper studies…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-04 Alan Pawlak , Hyunkook Lee , Aki Mäkivirta , Thomas Lund

Modeling real-world sound is a fundamental problem in the creative use of machine learning and many other fields, including human speech processing and bioacoustics. Transformer-based generative models and some prior work (e.g., DDSP) are…

Sound · Computer Science 2022-10-21 Masato Hagiwara , Maddie Cusimano , Jen-Yu Liu

This comprehensive paper delves into the forefront of personalized voice synthesis within artificial intelligence (AI), spotlighting the Dynamic Individual Voice Synthesis Engine (DIVSE). DIVSE represents a groundbreaking leap in…

Sound · Computer Science 2024-01-01 Fan Shi

In recent years, text-to-audio systems have achieved remarkable success, enabling the generation of complete audio segments directly from text descriptions. While these systems also facilitate music creation, the element of human creativity…

Sound · Computer Science 2025-04-15 Weixuan Yuan , Qadeer Khan , Vladimir Golkov

Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrowed stereo image, especially when created in non-professional settings without specialized equipment…

Sound · Computer Science 2026-05-06 Jan Melechovsky , Ambuj Mehrish , Abhinaba Roy , Dorien Herremans

Speech generated by parametric synthesizers generally suffers from a typical buzziness, similar to what was encountered in old LPC-like vocoders. In order to alleviate this problem, a more suited modeling of the excitation should be…

Sound · Computer Science 2020-01-06 Thomas Drugman , Geoffrey Wilfart , Thierry Dutoit

Articulatory trajectories like electromagnetic articulography (EMA) provide a low-dimensional representation of the vocal tract filter and have been used as natural, grounded features for speech synthesis. Differentiable digital signal…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-05 Yisi Liu , Bohan Yu , Drake Lin , Peter Wu , Cheol Jun Cho , Gopala Krishna Anumanchipalli

Generation of dynamic, scalable multi-species bird soundscapes remains a significant challenge in computer music and algorithmic sound design. Birdsongs involve rapid frequency-modulated chirps, complex amplitude envelopes, distinctive…

Sound · Computer Science 2025-11-25 Ellie L. Zhang , Duoduo Liao , Callie C. Liao

Synthetic datasets are widely used in many applications, such as missing data imputation, examining non-stationary scenarios, in simulations, training data-driven models, and analyzing system robustness. Typically, synthetic data are based…

Methodology · Statistics 2025-02-05 Ofek Aloni , Gal Perelman , Barak Fishbain

Expressive speech synthesis, like audiobook synthesis, is still challenging for style representation learning and prediction. Deriving from reference audio or predicting style tags from text requires a huge amount of labeled data, which is…

Sound · Computer Science 2022-06-28 Yihan Wu , Xi Wang , Shaofei Zhang , Lei He , Ruihua Song , Jian-Yun Nie

Current methods for creating drum loop audio in digital music production, such as using one-shot samples or resampling, often demand non-trivial efforts of creators. While recent generative models achieve high fidelity and adhere to text,…

Traditional audiometry often provides an incomplete characterization of the functional impact of hearing loss on speech understanding, particularly for supra-threshold deficits common in presbycusis. This motivates the development of more…

Sound · Computer Science 2025-06-16 Stefan Bleeck
‹ Prev 1 3 4 5 6 7 10 Next ›