中文
相关论文

相关论文: Advancing the Foundation Model for Music Understan…

200 篇论文

In this study, we aim to determine if generalized sounds and music can share a common emotional space, improving predictions of emotion in terms of arousal and valence. We propose the use of multiple datasets as a multi-domain learning…

声音 · 计算机科学 2024-08-15 Federico Simonetta , Francesca Certo , Stavros Ntalampiras

Multimodal sentiment analysis remains a challenging task due to the inherent heterogeneity across modalities. Such heterogeneity often manifests as asynchronous signals, imbalanced information between modalities, and interference from…

多媒体 · 计算机科学 2025-11-26 Yadong Liu , Shangfei Wang

Modeling of music audio semantics has been previously tackled through learning of mappings from audio data to high-level tags or latent unsupervised spaces. The resulting semantic spaces are theoretically limited, either because the chosen…

信息检索 · 计算机科学 2017-12-18 Francisco Raposo , David Martins de Matos , Ricardo Ribeiro , Suhua Tang , Yi Yu

Music exists in various modalities, such as score images, symbolic scores, MIDI, and audio. Translations between each modality are established as core tasks of music information retrieval, such as automatic music transcription…

声音 · 计算机科学 2026-04-08 Jongmin Jung , Dongmin Kim , Sihun Lee , Seola Cho , Hyungjoon Soh , Irmak Bukey , Chris Donahue , Dasaem Jeong

Music editing plays a vital role in modern music production, with applications in film, broadcasting, and game development. Recent advances in music generation models have enabled diverse editing tasks such as timbre transfer, instrument…

声音 · 计算机科学 2025-12-17 Yash Vishe , Eric Xue , Xunyi Jiang , Zachary Novack , Junda Wu , Julian McAuley , Xin Xu

Optical Music Recognition (OMR) is concerned with transcribing sheet music into a machine-readable format. The transcribed copy should allow musicians to compose, play and edit music by taking a picture of a music sheet. Complete…

计算机视觉与模式识别 · 计算机科学 2020-06-23 Elona Shatri , György Fazekas

Deep learning-based methods have achieved encouraging performances in the field of magnetic resonance (MR) image reconstruction. Nevertheless, to properly learn a powerful and robust model, these methods generally require large quantities…

图像与视频处理 · 电气工程与系统科学 2023-04-18 Ruoyou Wu , Cheng Li , Juan Zou , Qiegen Liu , Hairong Zheng , Shanshan Wang

Recent commercial systems such as Suno demonstrate strong capabilities in long-form song generation, while academic research remains largely non-reproducible due to the lack of publicly available training data, hindering fair comparison and…

The high computational complexity of the multiple signal classification (MUSIC) algorithm is mainly caused by the subspace decomposition and spectrum search, especially for frequent real-time applications or massive sensors. In this paper,…

信号处理 · 电气工程与系统科学 2025-06-16 Yiming Fang , Li Chen , Ang Chen , Weidong Wang

Multimodal video-audio-text understanding and generation can benefit from datasets that are narrow but rich. The narrowness allows bite-sized challenges that the research community can make progress on. The richness ensures we are making…

计算机视觉与模式识别 · 计算机科学 2022-04-29 Thomas Hayes , Songyang Zhang , Xi Yin , Guan Pang , Sasha Sheng , Harry Yang , Songwei Ge , Qiyuan Hu , Devi Parikh

Multimodal learning is a rapidly growing research field that has revolutionized multitasking and generative modeling in AI. While much of the research has focused on dealing with unstructured data (e.g., language, images, audio, or video),…

人工智能 · 计算机科学 2024-03-11 Marco D Alessandro , Enrique Calabrés , Mikel Elkano

In this paper, we present our solutions for emotion recognition in the sub-challenges of Multimodal Emotion Recognition Challenge (MER2024). To mitigate the modal competition issue between audio and text, we adopt an early fusion strategy…

多媒体 · 计算机科学 2024-10-01 Mengying Ge , Mingyang Li , Dongkai Tang , Pengbo Li , Kuo Liu , Shuhao Deng , Songbai Pu , Long Liu , Yang Song , Tao Zhang

Current audio foundation models typically rely on rigid, task-specific supervision, addressing isolated factors of audio rather than the whole. In contrast, human intelligence processes audio holistically, seamlessly bridging physical…

Reasoning has become a defining capability of modern foundation models, yet its development in the audio modality remains limited. Audio poses challenges that are distinct from those of text and vision. It is continuous, temporally dense,…

音频与语音处理 · 电气工程与系统科学 2026-05-21 Zhihan Guo , Wenqian Cui , Guan-Ting Lin , Daxin Tan , Jingyao Li , Qiyong Zheng , Dingdong Wang , Jing Xiong , Han Shi , Jiaya Jia , Irwin King

Computer-assisted pronunciation training (CAPT) manages to facilitate second-language (L2) learners to practice pronunciation skills by offering timely and instructive feedback. To examine pronunciation proficiency from multiple facets,…

音频与语音处理 · 电气工程与系统科学 2025-10-08 Bi-Cheng Yan , Ming-Kang Tsai , Berlin Chen

Recent advances in music foundation models have improved audio representation learning, yet their effectiveness across diverse musical traditions remains limited. We introduce CultureMERT-95M, a multi-culturally adapted foundation model…

声音 · 计算机科学 2025-06-24 Angelos-Nikolaos Kanatas , Charilaos Papaioannou , Alexandros Potamianos

Does popular music from the 60s sound different than that of the 90s? Prior study has shown that there would exist some variations of patterns and regularities related to instrumentation changes and growing loudness across multi-decadal…

声音 · 计算机科学 2024-07-09 Qiqi He , Xuchen Song , Weituo Hao , Ju-Chiang Wang , Wei-Tsung Lu , Wei Li

In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as robust acoustic and visual representations of raw video. Three…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Haotian Wang , Yuxuan Xi , Hang Chen , Jun Du , Yan Song , Qing Wang , Hengshun Zhou , Chenxi Wang , Jiefeng Ma , Pengfei Hu , Ya Jiang , Shi Cheng , Jie Zhang , Yuzhe Weng

Recent Multimodal Large Language Models (MLLMs) exhibit impressive abilities to perceive images and follow open-ended instructions. The capabilities of MLLMs depend on two crucial factors: the model architecture to facilitate the feature…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Tianyu Yu , Jinyi Hu , Yuan Yao , Haoye Zhang , Yue Zhao , Chongyi Wang , Shan Wang , Yinxv Pan , Jiao Xue , Dahai Li , Zhiyuan Liu , Hai-Tao Zheng , Maosong Sun

Sentiment analysis models exhibit complementary strengths, yet existing approaches lack a unified framework for effective integration. We present SentiFuse, a flexible and model-agnostic framework that integrates heterogeneous sentiment…

计算与语言 · 计算机科学 2026-02-03 Hieu Minh Duong , Rupa Ghosh , Cong Hoan Nguyen , Eugene Levin , Todd Gary , Long Nguyen
‹ 上一页 1 8 9 10 下一页 ›