中文
相关论文

相关论文: SegTune: Structured and Fine-Grained Control for S…

200 篇论文

As smart homes become increasingly prevalent, intelligent models are widely used for tasks such as anomaly detection and behavior prediction. These models are typically trained on static datasets, making them brittle to behavioral drift…

人工智能 · 计算机科学 2025-08-06 Zhiyao Xu , Dan Zhao , Qingsong Zou , Qing Li , Yong Jiang , Yuhang Wang , Jingyu Xiao

Many of the music generation systems based on neural networks are fully autonomous and do not offer control over the generation process. In this research, we present a controllable music generation system in terms of tonal tension. We…

声音 · 计算机科学 2020-10-15 Rui Guo , Ivor Simpson , Thor Magnusson , Chris Kiefer , Dorien Herremans

Dance plays an important role as an artistic form and expression in human culture, yet automatically generating dance sequences is a significant yet challenging endeavor. Existing approaches often neglect the critical aspect of…

计算机视觉与模式识别 · 计算机科学 2026-03-11 Hongsong Wang , Ying Zhu , Xin Geng , Liang Wang

RNN-Transducer (RNN-T) is a widely adopted architecture in speech recognition, integrating acoustic and language modeling in an end-to-end framework. However, the RNN-T predictor tends to over-rely on consecutive word dependencies in…

声音 · 计算机科学 2025-02-21 Khanh Le , Tuan Vu Ho , Dung Tran , Duc Thanh Chau

We propose Expotion (Facial Expression and Motion Control for Multimodal Music Generation), a generative model leveraging multimodal visual controls - specifically, human facial expressions and upper-body motion - as well as text prompts to…

声音 · 计算机科学 2025-07-08 Fathinah Izzati , Xinyue Li , Gus Xia

Generic generation and manipulation of text is challenging and has limited success compared to recent deep generative modeling in visual domain. This paper aims at generating plausible natural language sentences, whose attributes are…

机器学习 · 计算机科学 2018-09-14 Zhiting Hu , Zichao Yang , Xiaodan Liang , Ruslan Salakhutdinov , Eric P. Xing

MusicGen is a music generation language model (LM) that can be conditioned on textual descriptions and melodic features. We introduce MusicGen-Chord, which extends this capability by incorporating chord progression features. This model…

声音 · 计算机科学 2024-12-03 Jongmin Jung , Andreas Jansson , Dasaem Jeong

We present SegINR, a novel approach to neural Text-to-Speech (TTS) that addresses sequence alignment without relying on an auxiliary duration predictor and complex autoregressive (AR) or non-autoregressive (NAR) frame-level sequence…

音频与语音处理 · 电气工程与系统科学 2024-10-22 Minchan Kim , Myeonghun Jeong , Joun Yeop Lee , Nam Soo Kim

Recent advances in text-to-music generation (TTM) have yielded high-quality results, but often at the cost of extensive compute and the use of large proprietary internal data. To improve the affordability and openness of TTM training, an…

声音 · 计算机科学 2026-01-22 Wei-Jaw Lee , Fang-Chih Hsieh , Xuanjun Chen , Fang-Duo Tsai , Yi-Hsuan Yang

We propose structured prompt tuning, a simple and effective method to improve prompt tuning. Instead of prepending a sequence of tunable embeddings to the input, we generate the soft prompt embeddings through a hypernetwork. Our approach…

计算与语言 · 计算机科学 2022-05-26 Chi-Liang Liu , Hung-yi Lee , Wen-tau Yih

We present Text2midi-InferAlign, a novel technique for improving symbolic music generation at inference time. Our method leverages text-to-audio alignment and music structural alignment rewards during inference to encourage the generated…

声音 · 计算机科学 2025-05-20 Abhinaba Roy , Geeta Puri , Dorien Herremans

We examine the problem of learning a probabilistic model for melody directly from musical sequences belonging to the same genre. This is a challenging task as one needs to capture not only the rich temporal structure evident in music, but…

机器学习 · 计算机科学 2012-07-03 Athina Spiliopoulou , Amos Storkey

We consider the problem of sentence specified dynamic video thumbnail generation. Given an input video and a user query sentence, the goal is to generate a video thumbnail that not only provides the preview of the video content, but also…

计算机视觉与模式识别 · 计算机科学 2020-09-01 Mrigank Rochan , Mahesh Kumar Krishna Reddy , Yang Wang

We study cross-modal recommendation of music tracks to be used as soundtracks for videos. This problem is known as the music supervision task. We build on a self-supervised system that learns a content association between music and video.…

多媒体 · 计算机科学 2023-06-13 Laure Prétet , Gaël Richard , Clément Souchier , Geoffroy Peeters

Loops--short audio segments designed for seamless repetition--are central to many music genres, particularly those rooted in dance and electronic styles. However, current generative music models struggle to produce truly loopable audio, as…

Sequence generation applications require satisfying semantic constraints, such as ensuring that programs are correct, using certain keywords, or avoiding undesirable content. Language models, whether fine-tuned or prompted with few-shot…

计算与语言 · 计算机科学 2022-11-02 Sean Welleck , Ximing Lu , Peter West , Faeze Brahman , Tianxiao Shen , Daniel Khashabi , Yejin Choi

Storytelling video generation (SVG) aims to produce coherent and visually rich multi-scene videos that follow a structured narrative. Existing methods primarily employ LLM for high-level planning to decompose a story into scene-level…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Zun Wang , Jialu Li , Han Lin , Jaehong Yoon , Mohit Bansal

We study the problem of symbolic music generation (e.g., generating piano rolls), with a technical focus on non-differentiable rule guidance. Musical rules are often expressed in symbolic form on note characteristics, such as note density…

Many practices have been presented in music generation recently. While stylistic music generation using deep learning techniques has became the main stream, these models still struggle to generate music with high musicality, different…

声音 · 计算机科学 2021-05-12 Shuqi Dai , Xichu Ma , Ye Wang , Roger B. Dannenberg

Recent advances in text-to-audio generation enable models to translate natural-language descriptions into diverse musical output. However, the robustness of these systems under semantically equivalent prompt variations remains largely…

声音 · 计算机科学 2026-05-06 Jiahui Wu
‹ 上一页 1 8 9 10 下一页 ›