中文
相关论文

相关论文: An Audio Envelope Generator Derived from Industria…

200 篇论文

Signals can be interpreted as composed of a rapidly varying component modulated by a slower varying envelope. Identifying this envelope is an essential operation in signal processing, with applications in areas ranging from seismology to…

声音 · 计算机科学 2021-10-25 Carlos Tarjano , Valdecy Pereira

Video to sound generation aims to generate realistic and natural sound given a video input. However, previous video-to-sound generation methods can only generate a random or average timbre without any controls or specializations of the…

多媒体 · 计算机科学 2022-11-22 Chenye Cui , Yi Ren , Jinglin Liu , Rongjie Huang , Zhou Zhao

End-to-end generation of musical audio using deep learning techniques has seen an explosion of activity recently. However, most models concentrate on generating fully mixed music in response to abstract conditioning information. In this…

Differentiable digital signal processing (DDSP) techniques, including methods for audio synthesis, have gained attention in recent years and lend themselves to interpretability in the parameter space. However, current differentiable…

Musical expressivity and coherence are indispensable in music composition and performance, while often neglected in modern AI generative models. In this work, we introduce a listening-based data-processing technique that captures the…

声音 · 计算机科学 2025-03-18 Jingwei Liu

A method for musical audio synthesis using autoencoding neural networks is proposed. The autoencoder is trained to compress and reconstruct magnitude short-time Fourier transform frames. The autoencoder produces a spectrogram by activating…

音频与语音处理 · 电气工程与系统科学 2020-04-29 Joseph Colonel , Christopher Curro , Sam Keene

We propose a convex distributed optimization algorithm for synthesizing robust controllers for large-scale continuous time systems subject to exogenous disturbances. Given a large scale system, instead of solving the larger centralized…

最优化与控制 · 数学 2018-03-02 Mohamadreza Ahmadi , Murat Cubuktepe , Ufuk Topcu , Takashi Tanaka

The field of computer vision has witnessed phenomenal progress in recent years partially due to the development of deep convolutional neural networks. However, deep learning models are notoriously sensitive to adversarial examples which are…

计算机视觉与模式识别 · 计算机科学 2020-10-28 Haofeng Li , Yirui Zeng , Guanbin Li , Liang Lin , Yizhou Yu

We present a hybrid neural network and rule-based system that generates pop music. Music produced by pure rule-based systems often sounds mechanical. Music produced by machine learning sounds better, but still lacks hierarchical temporal…

声音 · 计算机科学 2017-10-09 Yifei Teng , An Zhao , Camille Goudeseune

The recent rise in capabilities of AI-based music generation tools has created an upheaval in the music industry, necessitating the creation of accurate methods to detect such AI-generated content. This can be done using audio-based…

In speech processing pipelines, improving the quality and intelligibility of real-world recordings is crucial. While supervised regression is the primary method for speech enhancement, audio tokenization is emerging as a promising…

声音 · 计算机科学 2025-07-18 Luca Della Libera , Cem Subakan , Mirco Ravanelli

Providing security to audio songs for maintaining its intellectual property right (IPR) is one of chanllenging fields in commercial world especially in creative industry. In this paper, an effective approach has been incorporated to…

密码学与安全 · 计算机科学 2012-02-24 Uttam Kr. Mondal , J. K. Mandal

Automatic speech recognition (ASR) systems are of vital importance nowadays in commonplace tasks such as speech-to-text processing and language translation. This created the need for an ASR system that can operate in realistic crowded…

音频与语音处理 · 电气工程与系统科学 2020-12-29 Sherif Abdulatif , Karim Armanious , Karim Guirguis , Jayasankar T. Sajeev , Bin Yang

In real-world applications of large language models, outputs are often required to be confined: selecting items from predefined product or document sets, generating phrases that comply with safety standards, or conforming to specialized…

计算与语言 · 计算机科学 2025-04-15 Haotian Ye , Himanshu Jain , Chong You , Ananda Theertha Suresh , Haowei Lin , James Zou , Felix Yu

Recent advancements in speech synthesis technology have enriched our daily lives, with high-quality and human-like audio widely adopted across real-world applications. However, malicious exploitation like voice-cloning fraud poses severe…

声音 · 计算机科学 2025-11-11 Zhisheng Zhang , Derui Wang , Yifan Mi , Zhiyong Wu , Jie Gao , Yuxin Cao , Kai Ye , Minhui Xue , Jie Hao

This paper introduces Open-Amp, a synthetic data framework for generating large-scale and diverse audio effects data. Audio effects are relevant to many musical audio processing and Music Information Retrieval (MIR) tasks, such as modelling…

音频与语音处理 · 电气工程与系统科学 2024-11-25 Alec Wright , Alistair Carson , Lauri Juvela

Recently, deep learning-based generative models have been introduced to generate singing voices. One approach is to predict the parametric vocoder features consisting of explicit speech parameters. This approach has the advantage that the…

音频与语音处理 · 电气工程与系统科学 2024-06-14 Tae-Woo Kim , Min-Su Kang , Gyeong-Hoon Lee

There has been a recent surge in adversarial attacks on deep learning based automatic speech recognition (ASR) systems. These attacks pose new challenges to deep learning security and have raised significant concerns in deploying ASR…

密码学与安全 · 计算机科学 2021-03-08 Shehzeen Hussain , Paarth Neekhara , Shlomo Dubnov , Julian McAuley , Farinaz Koushanfar

This paper describes an experimental system designed for development of real time voice synthesis applications. The system is composed from a DSP coprocessor card, equipped with an TMS320C25 or TMS320C50 chip, voice acquisition module…

声音 · 计算机科学 2008-03-04 Radu Arsinte , Attila Ferencz , Costin Miron

This paper presents a novel approach to neural instrument sound synthesis using a two-stage semi-supervised learning framework capable of generating pitch-accurate, high-quality music samples from an expressive timbre latent space. Existing…

声音 · 计算机科学 2025-10-07 Christian Limberg , Fares Schulz , Zhe Zhang , Stefan Weinzierl