English
Related papers

Related papers: Advancing the Foundation Model for Music Understan…

200 papers

In this study, we aim to determine if generalized sounds and music can share a common emotional space, improving predictions of emotion in terms of arousal and valence. We propose the use of multiple datasets as a multi-domain learning…

Sound · Computer Science 2024-08-15 Federico Simonetta , Francesca Certo , Stavros Ntalampiras

Multimodal sentiment analysis remains a challenging task due to the inherent heterogeneity across modalities. Such heterogeneity often manifests as asynchronous signals, imbalanced information between modalities, and interference from…

Multimedia · Computer Science 2025-11-26 Yadong Liu , Shangfei Wang

Modeling of music audio semantics has been previously tackled through learning of mappings from audio data to high-level tags or latent unsupervised spaces. The resulting semantic spaces are theoretically limited, either because the chosen…

Information Retrieval · Computer Science 2017-12-18 Francisco Raposo , David Martins de Matos , Ricardo Ribeiro , Suhua Tang , Yi Yu

Music exists in various modalities, such as score images, symbolic scores, MIDI, and audio. Translations between each modality are established as core tasks of music information retrieval, such as automatic music transcription…

Sound · Computer Science 2026-04-08 Jongmin Jung , Dongmin Kim , Sihun Lee , Seola Cho , Hyungjoon Soh , Irmak Bukey , Chris Donahue , Dasaem Jeong

Music editing plays a vital role in modern music production, with applications in film, broadcasting, and game development. Recent advances in music generation models have enabled diverse editing tasks such as timbre transfer, instrument…

Sound · Computer Science 2025-12-17 Yash Vishe , Eric Xue , Xunyi Jiang , Zachary Novack , Junda Wu , Julian McAuley , Xin Xu

Optical Music Recognition (OMR) is concerned with transcribing sheet music into a machine-readable format. The transcribed copy should allow musicians to compose, play and edit music by taking a picture of a music sheet. Complete…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Elona Shatri , György Fazekas

Deep learning-based methods have achieved encouraging performances in the field of magnetic resonance (MR) image reconstruction. Nevertheless, to properly learn a powerful and robust model, these methods generally require large quantities…

Image and Video Processing · Electrical Eng. & Systems 2023-04-18 Ruoyou Wu , Cheng Li , Juan Zou , Qiegen Liu , Hairong Zheng , Shanshan Wang

Recent commercial systems such as Suno demonstrate strong capabilities in long-form song generation, while academic research remains largely non-reproducible due to the lack of publicly available training data, hindering fair comparison and…

The high computational complexity of the multiple signal classification (MUSIC) algorithm is mainly caused by the subspace decomposition and spectrum search, especially for frequent real-time applications or massive sensors. In this paper,…

Signal Processing · Electrical Eng. & Systems 2025-06-16 Yiming Fang , Li Chen , Ang Chen , Weidong Wang

Multimodal video-audio-text understanding and generation can benefit from datasets that are narrow but rich. The narrowness allows bite-sized challenges that the research community can make progress on. The richness ensures we are making…

Computer Vision and Pattern Recognition · Computer Science 2022-04-29 Thomas Hayes , Songyang Zhang , Xi Yin , Guan Pang , Sasha Sheng , Harry Yang , Songwei Ge , Qiyuan Hu , Devi Parikh

Multimodal learning is a rapidly growing research field that has revolutionized multitasking and generative modeling in AI. While much of the research has focused on dealing with unstructured data (e.g., language, images, audio, or video),…

Artificial Intelligence · Computer Science 2024-03-11 Marco D Alessandro , Enrique Calabrés , Mikel Elkano

In this paper, we present our solutions for emotion recognition in the sub-challenges of Multimodal Emotion Recognition Challenge (MER2024). To mitigate the modal competition issue between audio and text, we adopt an early fusion strategy…

Multimedia · Computer Science 2024-10-01 Mengying Ge , Mingyang Li , Dongkai Tang , Pengbo Li , Kuo Liu , Shuhao Deng , Songbai Pu , Long Liu , Yang Song , Tao Zhang

Current audio foundation models typically rely on rigid, task-specific supervision, addressing isolated factors of audio rather than the whole. In contrast, human intelligence processes audio holistically, seamlessly bridging physical…

Reasoning has become a defining capability of modern foundation models, yet its development in the audio modality remains limited. Audio poses challenges that are distinct from those of text and vision. It is continuous, temporally dense,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 Zhihan Guo , Wenqian Cui , Guan-Ting Lin , Daxin Tan , Jingyao Li , Qiyong Zheng , Dingdong Wang , Jing Xiong , Han Shi , Jiaya Jia , Irwin King

Computer-assisted pronunciation training (CAPT) manages to facilitate second-language (L2) learners to practice pronunciation skills by offering timely and instructive feedback. To examine pronunciation proficiency from multiple facets,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-08 Bi-Cheng Yan , Ming-Kang Tsai , Berlin Chen

Recent advances in music foundation models have improved audio representation learning, yet their effectiveness across diverse musical traditions remains limited. We introduce CultureMERT-95M, a multi-culturally adapted foundation model…

Sound · Computer Science 2025-06-24 Angelos-Nikolaos Kanatas , Charilaos Papaioannou , Alexandros Potamianos

Does popular music from the 60s sound different than that of the 90s? Prior study has shown that there would exist some variations of patterns and regularities related to instrumentation changes and growing loudness across multi-decadal…

Sound · Computer Science 2024-07-09 Qiqi He , Xuchen Song , Weituo Hao , Ju-Chiang Wang , Wei-Tsung Lu , Wei Li

In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as robust acoustic and visual representations of raw video. Three…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-18 Haotian Wang , Yuxuan Xi , Hang Chen , Jun Du , Yan Song , Qing Wang , Hengshun Zhou , Chenxi Wang , Jiefeng Ma , Pengfei Hu , Ya Jiang , Shi Cheng , Jie Zhang , Yuzhe Weng

Recent Multimodal Large Language Models (MLLMs) exhibit impressive abilities to perceive images and follow open-ended instructions. The capabilities of MLLMs depend on two crucial factors: the model architecture to facilitate the feature…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Tianyu Yu , Jinyi Hu , Yuan Yao , Haoye Zhang , Yue Zhao , Chongyi Wang , Shan Wang , Yinxv Pan , Jiao Xue , Dahai Li , Zhiyuan Liu , Hai-Tao Zheng , Maosong Sun

Sentiment analysis models exhibit complementary strengths, yet existing approaches lack a unified framework for effective integration. We present SentiFuse, a flexible and model-agnostic framework that integrates heterogeneous sentiment…

Computation and Language · Computer Science 2026-02-03 Hieu Minh Duong , Rupa Ghosh , Cong Hoan Nguyen , Eugene Levin , Todd Gary , Long Nguyen
‹ Prev 1 8 9 10 Next ›