中文
相关论文

相关论文: MuQ-Eval: An Open-Source Per-Sample Quality Metric…

200 篇论文

The use of large language models (LLMs) for evaluating outputs is becoming an increasingly effective and scalable approach. However, it remains uncertain whether this capability extends beyond task-specific evaluations to more general…

计算与语言 · 计算机科学 2025-11-13 Rhitabrat Pokharel , Ameeta Agrawal

Evaluating the quality of slide-based multimedia instruction is challenging. Existing methods like manual assessment, reference-based metrics, and large language model evaluators face limitations in scalability, context capture, or bias. In…

计算与语言 · 计算机科学 2025-05-06 Joy Lim Jia Yin , Daniel Zhang-Li , Jifan Yu , Haoxuan Li , Shangqing Tu , Yuanchun Wang , Zhiyuan Liu , Huiqin Liu , Lei Hou , Juanzi Li , Bin Xu

Artificial Intelligence (AI)-generated feedback in educational settings has garnered considerable attention due to its potential to enhance learning outcomes. However, a comprehensive understanding of the linguistic characteristics of…

计算与语言 · 计算机科学 2025-05-01 Antoun Yaacoub , Zainab Assaghir , Lionel Prevost , Jérôme Da-Rugna

We introduce UEval, a benchmark to evaluate unified models, i.e., models capable of generating both images and text. UEval comprises 1,000 expert-curated questions that require both images and text in the model output, sourced from 8…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Bo Li , Yida Yin , Wenhao Chai , Xingyu Fu , Zhuang Liu

Mixed-precision quantization (MPQ) is crucial for deploying deep neural networks on resource-constrained devices, but finding the optimal bit-width for each layer represents a complex combinatorial optimization problem. Current…

机器学习 · 计算机科学 2026-03-24 Mehmet Emre Akbulut , Hazem Hesham Yousef Shalby , Fabrizio Pittorino , Manuel Roveri

Despite prolific work on evaluating generative models, little research has been done on the quality evaluation of an individual generated sample. To address this problem, a lightweight generated sample quality evaluation (LGSQE) method is…

图像与视频处理 · 电气工程与系统科学 2022-11-10 Ganning Zhao , Vasileios Magoulianitis , Suya You , C. -C. Jay Kuo

ODAQ (Open Dataset of Audio Quality) provides a comprehensive framework for exploring both monaural and binaural audio quality degradations across a range of distortion classes and signals, accompanied by subjective quality ratings. A…

音频与语音处理 · 电气工程与系统科学 2025-12-12 Pablo M. Delgado , Sascha Dick , Christoph Thompson , Chih-Wei Wu , Phillip A. Williams

One of the challenging problems in Music Information Retrieval is the acquisition of enough non-copyrighted audio recordings for model training and evaluation. This study compares two Transformer-based neural network models for chord…

声音 · 计算机科学 2025-08-11 Martyna Majchrzak , Jacek Mańdziuk

In this paper, we highlight a problem of evaluation metrics adopted in the open-vocabulary segmentation. That is, the evaluation process still heavily relies on closed-set metrics on zero-shot or cross-dataset pipelines without considering…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Hao Zhou , Tiancheng Shen , Xu Yang , Hai Huang , Xiangtai Li , Lu Qi , Ming-Hsuan Yang

Generative speech technologies are progressing rapidly, but evaluating the perceptual quality of synthetic speech remains a core challenge. Existing methods typically rely on scalar scores or binary decisions, which lack interpretability…

Music generation aims to create music segments that align with human aesthetics based on diverse conditional information. Despite advancements in generating music from specific textual descriptions (e.g., style, genre, instruments), the…

声音 · 计算机科学 2025-04-21 Jiahao Song , Yuzhao Wang

In recent years, AI-generated music has made significant progress, with several models performing well in multimodal and complex musical genres and scenes. While objective metrics can be used to evaluate generative music, they often lack…

声音 · 计算机科学 2023-08-29 Zeyu Xiong , Weitao Wang , Jing Yu , Yue Lin , Ziyan Wang

We introduce TCUQ, a single pass, label free uncertainty monitor for streaming TinyML that converts short horizon temporal consistency captured via lightweight signals on posteriors and features into a calibrated risk score with an O(W )…

机器学习 · 计算机科学 2025-08-19 Ismail Lamaakal , Chaymae Yahyati , Khalid El Makkaoui , Ibrahim Ouahbi , Yassine Maleh

Numerous studies in the field of music generation have demonstrated impressive performance, yet virtually no models are able to directly generate music to match accompanying videos. In this work, we develop a generative music AI framework,…

声音 · 计算机科学 2024-06-03 Jaeyong Kang , Soujanya Poria , Dorien Herremans

The rapid advancement of generative models has led to a growing volume of AI-generated videos, making the automatic quality assessment of such videos increasingly important. Existing AI-generated content video quality assessment (AIGC-VQA)…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Minghao Zou , Gen Liu , Guanghui Yue , Baoquan Zhao , Zhihua Wang , Paul L. Rosin , Hantao Liu , Wei Zhou

We present ArtifactNet, a lightweight framework that detects AI-generated music by reframing the problem as forensic physics -- extracting and analyzing the physical artifacts that neural audio codecs inevitably imprint on generated audio.…

声音 · 计算机科学 2026-04-21 Heewon Oh

We have developed reduced reference parametric models for estimating perceived quality in audiovisual multimedia services. We have created 144 unique configurations for audiovisual content including various application and network…

多媒体 · 计算机科学 2016-04-26 Edip Demirbilek , Jean-Charles Grégoire

Recent multimodal large language models (MLLMs) have demonstrated strong capabilities in image quality assessment (IQA) tasks. However, adapting such large-scale models is computationally expensive and still relies on substantial Mean…

计算机视觉与模式识别 · 计算机科学 2026-01-29 Xinyue Li , Zhichao Zhang , Zhiming Xu , Shubo Xu , Xiongkuo Min , Yitong Chen , Guangtao Zhai

Current app ranking and recommendation systems are mainly based on user-generated information, e.g., number of downloads and ratings. However, new apps often have few (or even no) user feedback, suffering from the classic cold-start…

信息检索 · 计算机科学 2021-09-09 Dan Su , Jiqiang Liu , Sencun Zhu , Xiaoyang Wang , Wei Wang , Xiangliang Zhang

In this paper, we introduce OmniEval, a benchmark for evaluating omni-modality models like MiniCPM-O 2.6, which encompasses visual, auditory, and textual inputs. Compared with existing benchmarks, our OmniEval has several distinctive…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Yiman Zhang , Ziheng Luo , Qiangyu Yan , Wei He , Borui Jiang , Xinghao Chen , Kai Han