中文
相关论文

相关论文: White-box Audio VST Effect Programming

200 篇论文

Many audio synthesizers can produce the same signal given different parameter configurations, meaning the inversion from sound to parameters is an inherently ill-posed problem. We show that this is largely due to intrinsic symmetries of the…

声音 · 计算机科学 2025-06-10 Ben Hayes , Charalampos Saitis , György Fazekas

We propose \textbf{IC-Effect}, an instruction-guided, DiT-based framework for few-shot video VFX editing that synthesizes complex effects (\eg flames, particles and cartoon characters) while strictly preserving spatial and temporal…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Yuanhang Li , Yiren Song , Junzhe Bai , Xinran Liang , Hu Yang , Libiao Jin , Qi Mao

Gray-box models offer significant benefit over black-box approaches for equipment emulator development for equipment since their integration of physics provides more confidence in the model outside of the training domain. However,…

机器学习 · 计算机科学 2023-10-04 Mohamed Kandil , J. J. McArthur

The COVID-19 pandemic shifted many events in our daily lives into the virtual domain. While virtual conference systems provide an alternative to physical meetings, larger events require a muted audience to avoid an accumulation of…

多媒体 · 计算机科学 2023-10-30 Tamay Aykut , Markus Hofbauer , Christopher Kuhn , Eckehard Steinbach , Bernd Girod

This thesis develops a Transformer model based on Whisper, which extracts melodies and chords from music audio and records them into ABC notation. A comprehensive data processing workflow is customized for ABC notation, including data…

声音 · 计算机科学 2024-10-23 Hongyao Zhang , Bohang Sun

Musicians and audio engineers sculpt and transform their sounds by connecting multiple processors, forming an audio processing graph. However, most deep-learning methods overlook this real-world practice and assume fixed graph settings. To…

声音 · 计算机科学 2023-05-09 Sungho Lee , Jaehyun Park , Seungryeol Paik , Kyogu Lee

In real-world voice conversion applications, environmental noise in source speech and user demands for expressive output pose critical challenges. Traditional ASR-based methods ensure noise robustness but suppress prosody richness, while…

音频与语音处理 · 电气工程与系统科学 2025-08-11 Yuepeng Jiang , Ziqian Ning , Shuai Wang , Chengjia Wang , Mengxiao Bi , Pengcheng Zhu , Zhonghua Fu , Lei Xie

The virtual world is being established in which digital humans are created indistinguishable from real humans. Producing their audio-related capabilities is crucial since voice conveys extensive personal characteristics. We aim to create a…

声音 · 计算机科学 2023-05-10 Wei Xue , Yiwen Wang , Qifeng Liu , Yike Guo

The hear-through functionality on hearing devices, which allows hearing equivalent to the open-ear while providing the possibility to modify the sound pressure at the eardrum in a desired manner, has drawn great attention from researchers…

音频与语音处理 · 电气工程与系统科学 2022-02-21 Wenyu Jin , Tim Schoof , Henning Schepker

Modulations are a critical part of sound design and music production, enabling the creation of complex and evolving audio. Modern synthesizers provide envelopes, low frequency oscillators (LFOs), and more parameter automation tools that…

声音 · 计算机科学 2025-10-08 Christopher Mitcheltree , Hao Hao Tan , Joshua D. Reiss

Large-scale, weakly-supervised speech recognition models, such as Whisper, have demonstrated impressive results on speech recognition across domains and languages. However, their application to long audio transcription via buffered or…

声音 · 计算机科学 2023-07-12 Max Bain , Jaesung Huh , Tengda Han , Andrew Zisserman

Many machine learning algorithms and classifiers are available only via API queries as a ``black-box'' -- that is, the downstream user has no ability to change, re-train, or fine-tune the model on a particular target distribution. Indeed,…

This paper is concerned with synthesizing programs based on black-box oracles: we are interested in the case where there exists an executable implementation of a component or library, but its internal structure is unknown. We are provided…

编程语言 · 计算机科学 2020-10-13 Bruce Collie , Jackson Woodruff , Michael F. P. O'Boyle

Dysfluencies and variations in speech pronunciation can severely degrade speech recognition performance, and for many individuals with moderate-to-severe speech disorders, voice operated systems do not work. Current speech recognition…

Video-to-speech synthesis involves reconstructing the speech signal of a speaker from a silent video. The implicit assumption of this task is that the sound signal is either missing or contains a high amount of noise/corruption such that it…

声音 · 计算机科学 2024-10-28 Triantafyllos Kefalas , Yannis Panagakis , Maja Pantic

Pivotuner is a VST3/AU MIDI effect plugin that automatically tunes note data in an adaptive pure intonation, in real time. Where previously pure intonation was out of reach for most musicians due to difficulty and impracticality, Pivotuner…

多媒体 · 计算机科学 2023-06-07 Dmitri Volkov

The generation of realistic, context-aware audio is important in real-world applications such as video game development. While existing video-to-audio (V2A) methods mainly focus on Foley sound generation, they struggle to produce…

The goal of voice conversion is to transform source speech into a target voice, keeping the content unchanged. In this paper, we focus on self-supervised representation learning for voice conversion. Specifically, we compare discrete and…

音频与语音处理 · 电气工程与系统科学 2022-06-09 Benjamin van Niekerk , Marc-André Carbonneau , Julian Zaïdi , Mathew Baas , Hugo Seuté , Herman Kamper

Spatial audio offers more immersive video consumption experiences to viewers; however, creating and editing spatial audio often expensive and requires specialized equipment and skills, posing a high barrier for amateur video creators. We…

人机交互 · 计算机科学 2024-04-24 Zheng Ning , Zheng Zhang , Jerrick Ban , Kaiwen Jiang , Ruohong Gan , Yapeng Tian , Toby Jia-Jun Li

Procedural audio, often referred to as "digital Foley", generates sound from scratch using computational processes. It represents an innovative approach to sound-effects creation. However, the development and adoption of procedural audio…

声音 · 计算机科学 2025-01-30 Nelly Garcia , Joshua Reiss