中文
相关论文

相关论文: Composer's Assistant 2: Interactive Multi-Track MI…

200 篇论文

We propose Beat Transformer, a novel Transformer encoder architecture for joint beat and downbeat tracking. Different from previous models that track beats solely based on the spectrogram of an audio mixture, our model deals with demixed…

声音 · 计算机科学 2022-09-16 Jingwei Zhao , Gus Xia , Ye Wang

Music performance is a distinctly human activity, intrinsically linked to the performer's ability to convey, evoke, or express emotion. Machines cannot perform music in the human sense; they can produce, reproduce, execute, or synthesize…

To apply neural sequence models such as the Transformers to music generation tasks, one has to represent a piece of music by a sequence of tokens drawn from a finite set of pre-defined vocabulary. Such a vocabulary usually involves tokens…

声音 · 计算机科学 2021-01-08 Wen-Yi Hsiao , Jen-Yu Liu , Yin-Cheng Yeh , Yi-Hsuan Yang

Language-guided navigation is a cornerstone of embodied AI, enabling agents to interpret language instructions and navigate complex environments. However, expert-provided instructions are limited in quantity, while synthesized annotations…

人工智能 · 计算机科学 2025-08-11 Zongtao He , Liuyi Wang , Lu Chen , Chengju Liu , Qijun Chen

Modern keyboards allow a musician to play multiple instruments at the same time by assigning zones -- fixed pitch ranges of the keyboard -- to different instruments. In this paper, we aim to further extend this idea and examine the…

声音 · 计算机科学 2021-10-22 Hao-Wen Dong , Chris Donahue , Taylor Berg-Kirkpatrick , Julian McAuley

In this work, we implement music production for silent film clips using LLM-driven method. Given the strong professional demands of film music production, we propose the FilmComposer, simulating the actual workflows of professional…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Zhifeng Xie , Qile He , Youjia Zhu , Qiwei He , Mengtian Li

Existing generative models, such as diffusion and auto-regressive networks, are inherently static, relying on a fixed set of pretrained parameters to handle all inputs. In contrast, humans flexibly adapt their internal generative…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Minh-Tuan Tran , Xuan-May Le , Quan Hung Tran , Mehrtash Harandi , Dinh Phung , Trung Le

Recently, multi-instrument music generation has become a hot topic. Different from single-instrument generation, multi-instrument generation needs to consider inter-track harmony besides intra-track coherence. This is usually achieved by…

声音 · 计算机科学 2023-05-29 Xipin Wei , Junhui Chen , Zirui Zheng , Li Guo , Lantian Li , Dong Wang

In order to collaborate and co-create with humans, an AI system must be capable of both reactive and anticipatory behavior. We present a case study of such a system in the domain of musical improvisation. We consider a duo consisting of a…

人工智能 · 计算机科学 2019-06-06 Oscar Thörn , Peter Fögel , Peter Knudsen , Luis de Miranda , Alessandro Saffiotti

With rapid advances in generative artificial intelligence, the text-to-music synthesis task has emerged as a promising direction for music generation. Nevertheless, achieving precise control over multi-track generation remains an open…

声音 · 计算机科学 2024-12-18 Yao Yao , Peike Li , Boyu Chen , Alex Wang

Creating and editing high-quality 3D content remains a central challenge in computer graphics. We address this challenge by introducing CompoSE, a novel method for Compositional Synthesis and Editing of 3D shapes via part-aware control. Our…

图形学 · 计算机科学 2026-05-20 Habib Slim , Shariq Farooq Bhat , Mohamed Elhoseiny , Yifan Wang , Mike Roberts

Building upon Diff-A-Riff, a latent diffusion model for musical instrument accompaniment generation, we present a series of improvements targeting quality, diversity, inference speed, and text-driven control. First, we upgrade the…

声音 · 计算机科学 2024-10-31 Javier Nistal , Marco Pasini , Stefan Lattner

While existing Singing Voice Synthesis systems achieve high-fidelity solo performances, they are constrained by global timbre control, failing to address dynamic multi-singer arrangement and vocal texture within a single song. To address…

声音 · 计算机科学 2026-02-10 Jiatao Chen , Xing Tang , Xiaoyue Duan , Yutang Feng , Jinchao Zhang , Jie Zhou

MIDI velocity is crucial for capturing expressive dynamics in human performances. In practical scenarios, a music score with inaccurate velocities may be available alongside the performance audio (e.g., music education and free online…

音频与语音处理 · 电气工程与系统科学 2026-03-03 Zhanhong He , Roberto Togneri , David Huang

This paper presents the first step in a research project situated within the field of musical agents. The objective is to achieve, through training, the tuning of the desired musical relationship between a live musical input and a real-time…

声音 · 计算机科学 2025-10-01 Balthazar Bujard , Jérôme Nika , Fédéric Bevilacqua , Nicolas Obin

This work presents a dual-agent \ac{llm}-based reasoning framework for automated planar mechanism synthesis that tightly couples linguistic specification with symbolic representation and simulation. From a natural-language task description,…

人工智能 · 计算机科学 2025-10-09 João Pedro Gandarela , Thiago Rios , Stefan Menzel , André Freitas

This Thesis discusses the development of technologies for the automatic resynthesis of music recordings using digital synthesizers. First, the main issue is identified in the understanding of how Music Information Processing (MIP) methods…

声音 · 计算机科学 2022-05-03 Federico Simonetta

This paper presents <Dialogue in Resonance>, an interactive music piece for a human pianist and a computer-controlled piano that integrates real-time automatic music transcription into a score-driven framework. Unlike previous approaches…

声音 · 计算机科学 2025-05-23 Hayeon Bang , Taegyun Kwon , Juhan Nam

This paper presents a fully automated procedure for controller synthesis for multi-agent systems under coupled constraints. Each agent has dynamics consisting of two terms: the first one models the coupled constraints and the other one is…

系统与控制 · 计算机科学 2016-09-20 Alexandros Nikou , Dimitris Boskos , Jana Tumova , Dimos V. Dimarogonas

Computer-generated visualisations can accompany recorded or live music to create novel audiovisual experiences for audiences. We present a system to streamline the creation of audio-driven visualisations based on audio feature extraction…

多媒体 · 计算机科学 2021-06-21 Max Graf , Harold Chijioke Opara , Mathieu Barthet