中文
相关论文

相关论文: DawDreamer: Bridging the Gap Between Digital Audio…

200 篇论文

Musicians and audio engineers sculpt and transform their sounds by connecting multiple processors, forming an audio processing graph. However, most deep-learning methods overlook this real-world practice and assume fixed graph settings. To…

声音 · 计算机科学 2023-05-09 Sungho Lee , Jaehyun Park , Seungryeol Paik , Kyogu Lee

Large general-purpose transformer models have recently become the mainstay in the realm of speech analysis. In particular, Whisper achieves state-of-the-art results in relevant tasks such as speech recognition, translation, language…

声音 · 计算机科学 2024-05-07 Antonio Bevilacqua , Paolo Saviano , Alessandro Amirante , Simon Pietro Romano

While generative models for music composition are increasingly capable, their adoption by musicians is hindered by text-prompting, an asynchronous workflow disconnected from the embodied, responsive nature of instrumental performance. To…

声音 · 计算机科学 2025-11-04 Louis Bradshaw , Alexander Spangher , Stella Biderman , Simon Colton

Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and onset and offset timestamps. This enables precise control…

声音 · 计算机科学 2026-02-10 Yisu Liu , Chenxing Li , Wanqian Zhang , Wenfu Wang , Meng Yu , Ruibo Fu , Zheng Lin , Weiping Wang , Dong Yu

For the task of speech separation, previous study usually treats multi-channel and single-channel scenarios as two research tracks with specialized solutions developed respectively. Instead, we propose a simple and unified architecture -…

声音 · 计算机科学 2023-03-15 Shuo Wang , Xiangyu Kong , Xiulian Peng , Mahmood Movassagh , Vinod Prakash , Yan Lu

While recent video-to-audio (V2A) models can generate realistic background audio from visual input, they largely overlook speech, an essential part of many video soundtracks. This paper proposes a new task, video-to-soundtrack (V2ST)…

多媒体 · 计算机科学 2025-07-15 Wenjie Tian , Xinfa Zhu , Haohe Liu , Zhixian Zhao , Zihao Chen , Chaofan Ding , Xinhan Di , Junjie Zheng , Lei Xie

Binaural audio gives the listener the feeling of being in the recording place and enhances the immersive experience if coupled with AR/VR. But the problem with binaural audio recording is that it requires a specialized setup which is not…

声音 · 计算机科学 2021-08-12 Kranti Kumar Parida , Siddharth Srivastava , Neeraj Matiyali , Gaurav Sharma

Current audio foundation models typically rely on rigid, task-specific supervision, addressing isolated factors of audio rather than the whole. In contrast, human intelligence processes audio holistically, seamlessly bridging physical…

Large-scale wireless testbeds are being increasingly used in developing and evaluating new solutions for next generation wireless networks. Among others, high-fidelity FPGA-based emulation platforms have unique capabilities for faithfully…

网络与互联网体系结构 · 计算机科学 2023-04-17 Davide Villa , Miead Tehrani-Moayyed , Pedram Johari , Stefano Basagni , Tommaso Melodia

Sound morphing is the process of gradually and smoothly transforming one sound into another to generate novel and perceptually hybrid sounds that simultaneously resemble both. Recently, diffusion-based text-to-audio models have produced…

音频与语音处理 · 电气工程与系统科学 2024-08-15 Purnima Kamath , Chitralekha Gupta , Suranga Nanayakkara

This paper introduces a novel interface for experiencing music through haptic impulses to the palm of the hand. It presents a practical implementation of the system exploring the realm of musical haptics through the translation of MIDI data…

人机交互 · 计算机科学 2024-08-14 Raj Varshith Moora , Gowdham Prabhakar

The rise of "bedroom producers" has democratized music creation, while challenging producers to objectively evaluate their work. To address this, we present AI TrackMate, an LLM-based music chatbot designed to provide constructive feedback…

声音 · 计算机科学 2024-12-10 Yi-Lin Jiang , Chia-Ho Hsiung , Yen-Tung Yeh , Lu-Rong Chen , Bo-Yu Chen

We introduce Dreamento (Dream engineering toolbox), an open-source Python package for dream engineering using sleep electroencephalography (EEG) wearables. Dreamento main functions are (1) real-time recording, monitoring, analysis, and…

人机交互 · 计算机科学 2023-01-04 Mahdad Jafarzadeh Esfahani , Amir Hossein Daraie , Paul Zerr , Frederik D. Weber , Martin Dresler

Recent Diffusion Transformers (DiTs) have shown impressive capabilities in generating high-quality single-modality content, including images, videos, and audio. However, it is still under-explored whether the transformer-based diffuser can…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Kai Wang , Shijian Deng , Jing Shi , Dimitrios Hatzinakos , Yapeng Tian

Cloud platforms allow users to execute tasks directly from their web browser and are a key enabling technology not only for commerce but also for computational science. Research software is often developed by scientists with limited…

Computational analysis of performed music is a key component of music information research, as performance shapes much of the music we hear. Music performance analysis studies the acoustic variations introduced by performers and how these…

声音 · 计算机科学 2026-05-06 Corentin Guichaoua , Daniel Bedoya , Elaine Chew

The increasing adoption of Jupyter notebooks in data science and machine learning workflows has created a gap between exploratory code development and production-ready software systems. While notebooks excel at iterative development and…

软件工程 · 计算机科学 2025-11-11 Hanya Elhashemy , Youssef Lotfy , Yongjian Tang

In contemporary popular music production, drum sound design is commonly performed by cumbersome browsing and processing of pre-recorded samples in sound libraries. One can also use specialized synthesis hardware, typically controlled…

声音 · 计算机科学 2022-06-30 Javier Nistal , Cyran Aouameur , Ithan Velarde , Stefan Lattner

For more than the past twenty years, Csound has been one of the leaders in the world of the computer music research, implementing innovative synthesis methods and making them available beyond the academic environments from which they often…

声音 · 计算机科学 2016-10-18 Gleb G. Rogozinsky , Eugene Cherny , Ivan Osipenko

Direct speech-to-speech translation (S2ST) translates speech from one language into another using a single model. However, due to the presence of linguistic and acoustic diversity, the target speech follows a complex multimodal…

计算与语言 · 计算机科学 2023-10-12 Qingkai Fang , Yan Zhou , Yang Feng