中文
相关论文

相关论文: SLEEPING-DISCO 9M: A large-scale pre-training data…

200 篇论文

Sounding Video Generation (SVG) is an audio-video joint generation task challenged by high-dimensional signal spaces, distinct data formats, and different patterns of content information. To address these issues, we introduce a novel…

计算机视觉与模式识别 · 计算机科学 2024-10-03 Mingzhen Sun , Weining Wang , Yanyuan Qiao , Jiahui Sun , Zihan Qin , Longteng Guo , Xinxin Zhu , Jing Liu

Deep generative models have emerged as a transformative tool in medical imaging, offering substantial potential for synthetic data generation. However, recent empirical studies highlight a critical vulnerability: these models can memorize…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Antonio Scardace , Lemuel Puglisi , Francesco Guarnera , Sebastiano Battiato , Daniele Ravì

Music generation introduces challenging complexities to large language models. Symbolic structures of music often include vertical harmonization as well as horizontal counterpoint, urging various adaptations and enhancements for large-scale…

声音 · 计算机科学 2024-07-30 Seungyeon Rhyu , Kichang Yang , Sungjun Cho , Jaehyeon Kim , Kyogu Lee , Moontae Lee

The current landscape of research leveraging large language models (LLMs) is experiencing a surge. Many works harness the powerful reasoning capabilities of these models to comprehend various modalities, such as text, speech, images,…

声音 · 计算机科学 2024-12-10 Shansong Liu , Atin Sakkeer Hussain , Qilong Wu , Chenshuo Sun , Ying Shan

Music-driven choreography is a challenging problem with a wide variety of industrial applications. Recently, many methods have been proposed to synthesize dance motions from music for a single dancer. However, generating dance motion for a…

多媒体 · 计算机科学 2023-03-28 Nhat Le , Thang Pham , Tuong Do , Erman Tjiputra , Quang D. Tran , Anh Nguyen

Music mixing involves combining individual tracks into a cohesive mixture, a task characterized by subjectivity where multiple valid solutions exist for the same input. Existing automatic mixing systems treat this task as a deterministic…

音频与语音处理 · 电气工程与系统科学 2025-11-12 Eloi Moliner , Marco A. Martínez-Ramírez , Junghyun Koo , Wei-Hsiang Liao , Kin Wai Cheuk , Joan Serrà , Vesa Välimäki , Yuki Mitsufuji

At the core of many important machine learning problems faced by online streaming services is a need to model how users interact with the content they are served. Unfortunately, there are no public datasets currently available that enable…

信息检索 · 计算机科学 2020-10-16 Brian Brost , Rishabh Mehrotra , Tristan Jehan

Recent advances in text-to-image (T2I) generation have enabled visually coherent image synthesis from descriptions, but generating images containing multiple given subjects remains challenging. As the number of reference identities…

机器学习 · 计算机科学 2026-04-10 Yucheng Zhou , Dubing Chen , Huan Zheng , Jianbing Shen

Musical performance requires prediction to operate instruments, to perform in groups and to improvise. In this paper, we investigate how a number of digital musical instruments (DMIs), including two of our own, have applied predictive…

声音 · 计算机科学 2018-12-21 Charles P. Martin , Kai Olav Ellefsen , Jim Torresen

Multimodal Large Language Models (MLLMs) have achieved remarkable visual reasoning abilities in natural images, text-rich documents, and graphic designs. However, their ability to interpret music sheets remains underexplored. To bridge this…

计算机视觉与模式识别 · 计算机科学 2025-08-15 Jian Chen , Wenye Ma , Penghang Liu , Wei Wang , Tengwei Song , Ming Li , Chenguang Wang , Jiayu Qin , Ruiyi Zhang , Changyou Chen

In music source separation (MSS), obtaining isolated sources or stems is highly costly, making pre-training on unlabeled data a promising approach. Although source-agnostic unsupervised learning like mixture-invariant training (MixIT) has…

音频与语音处理 · 电气工程与系统科学 2025-05-13 Kohei Saijo , Yoshiaki Bando

Current deep networks are very data-hungry and benefit from training on largescale datasets, which are often time-consuming to collect and annotate. By contrast, synthetic data can be generated infinitely using generative models such as…

计算机视觉与模式识别 · 计算机科学 2023-10-11 Weijia Wu , Yuzhong Zhao , Hao Chen , Yuchao Gu , Rui Zhao , Yefei He , Hong Zhou , Mike Zheng Shou , Chunhua Shen

Music-to-dance generation aims to synthesize human dance motion conditioned on musical input. Despite recent progress, significant challenges remain due to the semantic gap between music and dance motion, as music offers only abstract cues,…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Qing Wang , Xiaohang Yang , Yilan Dong , Naveen Raj Govindaraj , Gregory Slabaugh , Shanxin Yuan

Composing music for video is essential yet challenging, leading to a growing interest in automating music generation for video applications. Existing approaches often struggle to achieve robust music-video correspondence and generative…

声音 · 计算机科学 2025-04-21 Heda Zuo , Weitao You , Junxian Wu , Shihong Ren , Pei Chen , Mingxu Zhou , Yujia Lu , Lingyun Sun

Large language models (LLMs) have become a dominant and important tool for NLP researchers in a wide range of tasks. Today, many researchers use LLMs in synthetic data generation, task evaluation, fine-tuning, distillation, and other…

计算与语言 · 计算机科学 2024-05-29 Ajay Patel , Colin Raffel , Chris Callison-Burch

Downscaling techniques are one of the most prominent applications of Deep Learning (DL) in Earth System Modeling. A robust DL downscaling model can generate high-resolution fields from coarse-scale numerical model simulations, saving the…

机器学习 · 计算机科学 2025-08-28 Elena Tomasi , Gabriele Franch , Marco Cristoforetti

Music information retrieval (MIR) has gone through an explosive development with the advancement of deep learning in recent years. However, music genres like electronic dance music (EDM) has always been relatively less investigated compared…

声音 · 计算机科学 2023-05-13 Xinyu Li

Reviews of songs play an important role in online music service platforms. Prior research shows that users can make quicker and more informed decisions when presented with meaningful song reviews. However, reviews of music songs are…

信息检索 · 计算机科学 2022-05-31 Jingya Zang , Cuiyun Gao , Yupan Chen , Ruifeng Xu , Lanjun Zhou , Xuan Wang

Scientific figure interpretation is a crucial capability for AI-driven scientific assistants built on advanced Large Vision Language Models. However, current datasets and benchmarks primarily focus on simple charts or other relatively…

In this work, we provide a broad comparative analysis of strategies for pre-training audio understanding models for several tasks in the music domain, including labelling of genre, era, origin, mood, instrumentation, key, pitch, vocal…

‹ 上一页 1 8 9 10 下一页 ›