English
Related papers

Related papers: SLEEPING-DISCO 9M: A large-scale pre-training data…

200 papers

Sounding Video Generation (SVG) is an audio-video joint generation task challenged by high-dimensional signal spaces, distinct data formats, and different patterns of content information. To address these issues, we introduce a novel…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Mingzhen Sun , Weining Wang , Yanyuan Qiao , Jiahui Sun , Zihan Qin , Longteng Guo , Xinxin Zhu , Jing Liu

Deep generative models have emerged as a transformative tool in medical imaging, offering substantial potential for synthetic data generation. However, recent empirical studies highlight a critical vulnerability: these models can memorize…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Antonio Scardace , Lemuel Puglisi , Francesco Guarnera , Sebastiano Battiato , Daniele Ravì

Music generation introduces challenging complexities to large language models. Symbolic structures of music often include vertical harmonization as well as horizontal counterpoint, urging various adaptations and enhancements for large-scale…

Sound · Computer Science 2024-07-30 Seungyeon Rhyu , Kichang Yang , Sungjun Cho , Jaehyeon Kim , Kyogu Lee , Moontae Lee

The current landscape of research leveraging large language models (LLMs) is experiencing a surge. Many works harness the powerful reasoning capabilities of these models to comprehend various modalities, such as text, speech, images,…

Sound · Computer Science 2024-12-10 Shansong Liu , Atin Sakkeer Hussain , Qilong Wu , Chenshuo Sun , Ying Shan

Music-driven choreography is a challenging problem with a wide variety of industrial applications. Recently, many methods have been proposed to synthesize dance motions from music for a single dancer. However, generating dance motion for a…

Multimedia · Computer Science 2023-03-28 Nhat Le , Thang Pham , Tuong Do , Erman Tjiputra , Quang D. Tran , Anh Nguyen

Music mixing involves combining individual tracks into a cohesive mixture, a task characterized by subjectivity where multiple valid solutions exist for the same input. Existing automatic mixing systems treat this task as a deterministic…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-12 Eloi Moliner , Marco A. Martínez-Ramírez , Junghyun Koo , Wei-Hsiang Liao , Kin Wai Cheuk , Joan Serrà , Vesa Välimäki , Yuki Mitsufuji

At the core of many important machine learning problems faced by online streaming services is a need to model how users interact with the content they are served. Unfortunately, there are no public datasets currently available that enable…

Information Retrieval · Computer Science 2020-10-16 Brian Brost , Rishabh Mehrotra , Tristan Jehan

Recent advances in text-to-image (T2I) generation have enabled visually coherent image synthesis from descriptions, but generating images containing multiple given subjects remains challenging. As the number of reference identities…

Machine Learning · Computer Science 2026-04-10 Yucheng Zhou , Dubing Chen , Huan Zheng , Jianbing Shen

Musical performance requires prediction to operate instruments, to perform in groups and to improvise. In this paper, we investigate how a number of digital musical instruments (DMIs), including two of our own, have applied predictive…

Sound · Computer Science 2018-12-21 Charles P. Martin , Kai Olav Ellefsen , Jim Torresen

Multimodal Large Language Models (MLLMs) have achieved remarkable visual reasoning abilities in natural images, text-rich documents, and graphic designs. However, their ability to interpret music sheets remains underexplored. To bridge this…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Jian Chen , Wenye Ma , Penghang Liu , Wei Wang , Tengwei Song , Ming Li , Chenguang Wang , Jiayu Qin , Ruiyi Zhang , Changyou Chen

In music source separation (MSS), obtaining isolated sources or stems is highly costly, making pre-training on unlabeled data a promising approach. Although source-agnostic unsupervised learning like mixture-invariant training (MixIT) has…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-13 Kohei Saijo , Yoshiaki Bando

Current deep networks are very data-hungry and benefit from training on largescale datasets, which are often time-consuming to collect and annotate. By contrast, synthetic data can be generated infinitely using generative models such as…

Computer Vision and Pattern Recognition · Computer Science 2023-10-11 Weijia Wu , Yuzhong Zhao , Hao Chen , Yuchao Gu , Rui Zhao , Yefei He , Hong Zhou , Mike Zheng Shou , Chunhua Shen

Music-to-dance generation aims to synthesize human dance motion conditioned on musical input. Despite recent progress, significant challenges remain due to the semantic gap between music and dance motion, as music offers only abstract cues,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Qing Wang , Xiaohang Yang , Yilan Dong , Naveen Raj Govindaraj , Gregory Slabaugh , Shanxin Yuan

Composing music for video is essential yet challenging, leading to a growing interest in automating music generation for video applications. Existing approaches often struggle to achieve robust music-video correspondence and generative…

Sound · Computer Science 2025-04-21 Heda Zuo , Weitao You , Junxian Wu , Shihong Ren , Pei Chen , Mingxu Zhou , Yujia Lu , Lingyun Sun

Large language models (LLMs) have become a dominant and important tool for NLP researchers in a wide range of tasks. Today, many researchers use LLMs in synthetic data generation, task evaluation, fine-tuning, distillation, and other…

Computation and Language · Computer Science 2024-05-29 Ajay Patel , Colin Raffel , Chris Callison-Burch

Downscaling techniques are one of the most prominent applications of Deep Learning (DL) in Earth System Modeling. A robust DL downscaling model can generate high-resolution fields from coarse-scale numerical model simulations, saving the…

Machine Learning · Computer Science 2025-08-28 Elena Tomasi , Gabriele Franch , Marco Cristoforetti

Music information retrieval (MIR) has gone through an explosive development with the advancement of deep learning in recent years. However, music genres like electronic dance music (EDM) has always been relatively less investigated compared…

Sound · Computer Science 2023-05-13 Xinyu Li

Reviews of songs play an important role in online music service platforms. Prior research shows that users can make quicker and more informed decisions when presented with meaningful song reviews. However, reviews of music songs are…

Information Retrieval · Computer Science 2022-05-31 Jingya Zang , Cuiyun Gao , Yupan Chen , Ruifeng Xu , Lanjun Zhou , Xuan Wang

Scientific figure interpretation is a crucial capability for AI-driven scientific assistants built on advanced Large Vision Language Models. However, current datasets and benchmarks primarily focus on simple charts or other relatively…

In this work, we provide a broad comparative analysis of strategies for pre-training audio understanding models for several tasks in the music domain, including labelling of genre, era, origin, mood, instrumentation, key, pitch, vocal…

‹ Prev 1 8 9 10 Next ›