中文
相关论文

相关论文: MuQ-Eval: An Open-Source Per-Sample Quality Metric…

200 篇论文

As multimedia data flourishes on the Internet, quality assessment (QA) of multimedia data becomes paramount for digital media applications. Since multimedia data includes multiple modalities including audio, image, video, and audio-visual…

图像与视频处理 · 电气工程与系统科学 2024-07-30 Yuqin Cao , Xiongkuo Min , Yixuan Gao , Wei Sun , Weisi Lin , Guangtao Zhai

Research into the prediction and analysis of perceived audio quality is hampered by the scarcity of openly available datasets of audio signals accompanied by corresponding subjective quality scores. To address this problem, we present the…

音频与语音处理 · 电气工程与系统科学 2024-01-02 Matteo Torcoli , Chih-Wei Wu , Sascha Dick , Phillip A. Williams , Mhd Modar Halimeh , William Wolcott , Emanuel A. P. Habets

Modern audio systems universally employ mel-scale representations derived from 1940s Western psychoacoustic studies, potentially encoding cultural biases that create systematic performance disparities. We present a comprehensive evaluation…

声音 · 计算机科学 2026-04-14 Shivam Chauhan , Ajay Pundhir

Recent advancements in music generation are raising multiple concerns about the implications of AI in creative music processes, current business models and impacts related to intellectual property management. A relevant discussion and…

声音 · 计算机科学 2025-07-07 Roser Batlle-Roca , Wei-Hsiang Liao , Xavier Serra , Yuki Mitsufuji , Emilia Gómez

Recent advancements have brought generated music closer to human-created compositions, yet evaluating these models remains challenging. While human preference is the gold standard for assessing quality, translating these subjective…

机器学习 · 计算机科学 2025-06-25 Florian Grötschla , Ahmet Solak , Luca A. Lanzendörfer , Roger Wattenhofer

Recently, an increasing number of multimodal (text and audio) benchmarks have emerged, primarily focusing on evaluating models' understanding capability. However, exploration into assessing generative capabilities remains limited,…

Aesthetics serve as an implicit and important criterion in song generation tasks that reflect human perception beyond objective metrics. However, evaluating the aesthetics of generated songs remains a fundamental challenge, as the…

音频与语音处理 · 电气工程与系统科学 2025-05-19 Jixun Yao , Guobin Ma , Huixin Xue , Huakang Chen , Chunbo Hao , Yuepeng Jiang , Haohe Liu , Ruibin Yuan , Jin Xu , Wei Xue , Hao Liu , Lei Xie

Conversational recommendation has advanced rapidly with large language models (LLMs), yet music remains a uniquely challenging domain in which effective recommendations require reasoning over audio content beyond what text or metadata can…

声音 · 计算机科学 2026-01-26 Rohan Surana , Amit Namburi , Gagan Mundada , Abhay Lal , Zachary Novack , Julian McAuley , Junda Wu

This study presents a machine learning framework for assessing similarity between audio content and predicting sentiment score. We construct a dataset containing audio samples from music covers on YouTube along with the audio of the…

声音 · 计算机科学 2024-11-04 Aris J. Aristorenas

In an earlier study, we gathered perceptual evaluations of the audio, video, and audiovisual quality for 360 audiovisual content. This paper investigates perceived audiovisual quality prediction based on objective quality metrics and…

多媒体 · 计算机科学 2021-12-24 Randy Frans Fela , Nick Zacharov , Søren Forchhammer

With the rise of AI-generated content (AIGC), generating perceptually natural and feeling-aligned music from multimodal inputs has become a central challenge. Existing approaches often rely on explicit emotion labels that require costly…

声音 · 计算机科学 2025-12-02 Jiaying Hong , Ting Zhu , Thanet Markchom , Huizhi Liang

In recent years, artificial intelligence (AI)-driven video generation has gained significant attention. Consequently, there is a growing need for accurate video quality assessment (VQA) metrics to evaluate the perceptual quality of…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Zhichao Zhang , Wei Sun , Xinyue Li , Jun Jia , Xiongkuo Min , Zicheng Zhang , Chunyi Li , Zijian Chen , Puyi Wang , Fengyu Sun , Shangling Jui , Guangtao Zhai

The Internet is integral to modern life, influencing communication, business, and lifestyles globally. As dependence on Internet services grows, the demand for high-quality service delivery increases. Service providers must maintain high…

网络与互联网体系结构 · 计算机科学 2025-03-11 Parsa Hassani Shariat Panahi , Amir Hossein Jalilvand , Abolfazl Diyanat

Multimodal models that jointly process audio and language hold great promise in audio understanding and are increasingly being adopted in the music domain. By allowing users to query via text and obtain information about a given audio…

声音 · 计算机科学 2024-08-05 Benno Weck , Ilaria Manco , Emmanouil Benetos , Elio Quinton , George Fazekas , Dmitry Bogdanov

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we bridge this critical gap by establishing a comprehensive…

The complex nature of musical emotion introduces inherent bias in both recognition and generation, particularly when relying on a single audio encoder, emotion classifier, or evaluation metric. In this work, we conduct a study on Music…

音频与语音处理 · 电气工程与系统科学 2025-05-01 Yuanchao Li , Azalea Gui , Dimitra Emmanouilidou , Hannes Gamper

Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortions are often asymmetric, where one modality may be severely degraded while the other…

多媒体 · 计算机科学 2026-05-05 Mayesha Maliha R. Mithila , Mylene C. Q. Farias

We propose the Fr\'echet Audio Distance (FAD), a novel, reference-free evaluation metric for music enhancement algorithms. We demonstrate how typical evaluation metrics for speech enhancement and blind source separation can fail to…

音频与语音处理 · 电气工程与系统科学 2019-01-18 Kevin Kilgour , Mauricio Zuluaga , Dominik Roblek , Matthew Sharifi

Despite significant recent advances in generative acoustic text-to-music (TTM) modeling, robust evaluation of these models lags behind, relying in particular on the popular Fr\'echet Audio Distance (FAD). In this work, we rigorously study…

The rapidly evolving multimodal Large Language Models (LLMs) urgently require new benchmarks to uniformly evaluate their performance on understanding and textually describing music. However, due to semantic gaps between Music Information…

声音 · 计算机科学 2024-06-14 Zihao Wang , Shuyu Li , Tao Zhang , Qi Wang , Pengfei Yu , Jinyang Luo , Yan Liu , Ming Xi , Kejun Zhang