中文
相关论文

相关论文: Aligning Text-to-Music Evaluation with Human Prefe…

200 篇论文

Diffusion and flow-matching models have revolutionized automatic text-to-audio generation in recent times. These models are increasingly capable of generating high quality and faithful audio outputs capturing to speech and acoustic events.…

We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production…

音频与语音处理 · 电气工程与系统科学 2025-09-04 Dyah A. M. G. Wisnu , Ryandhimas E. Zezario , Stefano Rini , Hsin-Min Wang , Yu Tsao

Music Recommendation Systems (MRSs) are a cornerstone of modern streaming platforms. Existing recommendation models, spanning both recall and ranking stages, predominantly rely on collaborative filtering, which fails to exploit the…

信息检索 · 计算机科学 2026-04-24 Yizhi Zhou , Jia-Qi Yang , De-Chuan Zhan , Da-Wei Zhou

Tag-based music retrieval is crucial to browse large-scale music libraries efficiently. Hence, automatic music tagging has been actively explored, mostly as a classification task, which has an inherent limitation: a fixed vocabulary. On the…

信息检索 · 计算机科学 2020-11-02 Minz Won , Sergio Oramas , Oriol Nieto , Fabien Gouyon , Xavier Serra

We propose a novel objective evaluation metric for synthesized audio in text-to-audio (TTA), aiming to improve the performance of TTA models. In TTA, subjective evaluation of the synthesized sound is an important, but its implementation…

声音 · 计算机科学 2025-07-02 Minoru Kishi , Ryosuke Sakai , Shinnosuke Takamichi , Yusuke Kanamori , Yuki Okamoto

Many speech processing methods based on deep learning require an automatic and differentiable audio metric for the loss function. The DPAM approach of Manocha et al. learns a full-reference metric trained directly on human judgments, and…

音频与语音处理 · 电气工程与系统科学 2021-02-11 Pranay Manocha , Zeyu Jin , Richard Zhang , Adam Finkelstein

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we bridge this critical gap by establishing a comprehensive…

We introduce MAVE (Mamba with Cross-Attention for Voice Editing and Synthesis), a novel autoregressive architecture for text-conditioned voice editing and high-fidelity text-to-speech (TTS) synthesis, built on a cross-attentive Mamba…

声音 · 计算机科学 2025-10-07 Baher Mohammad , Magauiya Zhussip , Stamatios Lefkimmiatis

Generative artificial intelligence has made significant strides, producing text indistinguishable from human prose and remarkably photorealistic images. Automatically measuring how close the generated data distribution is to the target…

Many audio processing tasks require perceptual assessment. The ``gold standard`` of obtaining human judgments is time-consuming, expensive, and cannot be used as an optimization criterion. On the other hand, automated metrics are efficient…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Pranay Manocha , Adam Finkelstein , Richard Zhang , Nicholas J. Bryan , Gautham J. Mysore , Zeyu Jin

Evaluating audio generation systems, including text-to-music (TTM), text-to-speech (TTS), and text-to-audio (TTA), remains challenging due to the subjective and multi-dimensional nature of human perception. Existing methods treat mean…

声音 · 计算机科学 2025-08-13 Chien-Chun Wang , Kuan-Tang Huang , Cheng-Yeh Yang , Hung-Shin Lee , Hsin-Min Wang , Berlin Chen

Recent progress in text-to-music generation has enabled models to synthesize high-quality musical segments, full compositions, and even respond to fine-grained control signals, e.g. chord progressions. State-of-the-art (SOTA) systems differ…

声音 · 计算机科学 2025-09-05 Or Tal , Felix Kreuk , Yossi Adi

This paper introduces effective design choices for text-to-music retrieval systems. An ideal text-based retrieval system would support various input queries such as pre-defined tags, unseen tags, and sentence-level descriptions. In reality,…

信息检索 · 计算机科学 2022-11-29 SeungHeon Doh , Minz Won , Keunwoo Choi , Juhan Nam

Significant advancements have been made in video generative models recently. Unlike image generation, video generation presents greater challenges, requiring not only generating high-quality frames but also ensuring temporal consistency…

计算机视觉与模式识别 · 计算机科学 2024-07-24 Jiahe Liu , Youran Qu , Qi Yan , Xiaohui Zeng , Lele Wang , Renjie Liao

Diffusion based Text-To-Music (TTM) models generate music corresponding to text descriptions. Typically UNet based diffusion models condition on text embeddings generated from a pre-trained large language model or from a cross-modality…

音频与语音处理 · 电气工程与系统科学 2025-01-28 Jisi Zhang , Pablo Peso Parada , Md Asif Jalal , Karthikeyan Saravanan

Manual sound design with a synthesizer is inherently iterative: an artist compares the synthesized output to a mental target, adjusts parameters, and repeats until satisfied. Iterative sound-matching automates this workflow by continually…

声音 · 计算机科学 2025-10-10 Amir Salimi , Abram Hindle , Osmar R. Zaiane

This paper presents NOMAD (Non-Matching Audio Distance), a differentiable perceptual similarity metric that measures the distance of a degraded signal against non-matching references. The proposed method is based on learning deep feature…

声音 · 计算机科学 2024-01-22 Alessandro Ragano , Jan Skoglund , Andrew Hines

With the similarity between music and speech synthesis from symbolic input and the rapid development of text-to-speech (TTS) techniques, it is worthwhile to explore ways to improve the MIDI-to-audio performance by borrowing from TTS…

声音 · 计算机科学 2023-03-22 Xuan Shi , Erica Cooper , Xin Wang , Junichi Yamagishi , Shrikanth Narayanan

The popular success of text-based large language models (LLM) has streamlined the attention of the multimodal community to combine other modalities like vision and audio along with text to achieve similar multimodal capabilities. In this…

计算与语言 · 计算机科学 2025-05-20 Debarpan Bhattacharya , Apoorva Kulkarni , Sriram Ganapathy

Aligning large generative models with human feedback is a critical challenge. In speech synthesis, this is particularly pronounced due to the lack of a large-scale human preference dataset, which hinders the development of models that truly…