English
Related papers

Related papers: The ICASSP 2026 Automatic Song Aesthetics Evaluati…

200 papers

Generative AI models for music and the arts in general are increasingly complex and hard to understand. The field of eXplainable AI (XAI) seeks to make complex and opaque AI models such as neural networks more understandable to people. One…

Sound · Computer Science 2024-02-06 Nick Bryan-Kinns , Bingyuan Zhang , Songyan Zhao , Berker Banar

This technical report describes two methods that were developed for Task 2 of the DCASE 2020 challenge. The challenge involves an unsupervised learning to detect anomalous sounds, thus only normal machine working condition samples are…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-22 Alexandrine Ribeiro , Luis Miguel Matos , Pedro Jose Pereira , Eduardo C. Nunes , Andre L. Ferreira , Paulo Cortez , Andre Pilastri

Scientific peer review faces mounting strain as submission volumes surge, making it increasingly difficult to sustain review quality, consistency, and timeliness. Recent advances in AI have led the community to consider its use in peer…

Deep learning models are typically evaluated to measure and compare their performance on a given task. The metrics that are commonly used to evaluate these models are standard metrics that are used for different tasks. In the field of music…

Sound · Computer Science 2022-04-05 Carlos Hernandez-Olivan , Jorge Abadias Puyuelo , Jose R. Beltran

Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In this paper, we propose…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-30 Dapeng Wu , Shun Lei , Wei Tan , Guangzheng Li , Yunzhe Wang , Huaicheng Zhang , Lishi Zuo , Zhiyong Wu

This document provides the results of the tests of acoustic parameter estimation algorithms on the Acoustic Characterization of Environments (ACE) Challenge Evaluation dataset which were subsequently submitted and written up into papers for…

Sound · Computer Science 2017-06-28 James Eaton , Nikolay D. Gaubitch , Alastair H. Moore , Patrick A. Naylor

Formalized in ISO 12913, the "soundscape" approach is a paradigmatic shift towards perception-based urban sound management, aiming to alleviate the substantial socioeconomic costs of noise pollution to advance the United Nations Sustainable…

This pictorial aims to critically consider the nature of text-to-audio and text-to-music generative tools in the context of explainable AI. As a group of experimental musicians and researchers, we are enthusiastic about the creative…

Sound · Computer Science 2024-08-15 Jesse Allison , Drew Farrar , Treya Nash , Carlos Román , Morgan Weeks , Fiona Xue Ju

Understanding and identifying musical shape plays an important role in music education and performance assessment. To simplify the otherwise time- and cost-intensive musical shape evaluation, in this paper we explore how artificial…

Sound · Computer Science 2024-01-08 Xiaoquan Li , Stephan Weiss , Yijun Yan , Yinhe Li , Jinchang Ren , John Soraghan , Ming Gong

Computational visual aesthetics has recently become an active research area. Existing state-of-art methods formulate this as a binary classification task where a given image is predicted to be beautiful or not. In many applications such as…

Computer Vision and Pattern Recognition · Computer Science 2017-04-06 Parag S. Chandakkar , Vijetha Gattupalli , Baoxin Li

We present a thorough analysis of the findings of the latest iteration of the Singing Voice Conversion Challenge, a scientific event aiming to compare and understand different voice conversion systems in a controlled environment. Compared…

Although text-to-audio generation has made remarkable progress in realism and diversity, the development of evaluation metrics has not kept pace. Widely-adopted approaches, typically based on embedding similarity like CLAPScore, effectively…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-22 Chun-Yi Kuan , Kai-Wei Chang , Hung-yi Lee

In recent years, artificial intelligence (AI) has made significant progress in the field of music generation, driving innovation in music creation and applications. This paper provides a systematic review of the latest research advancements…

Sound · Computer Science 2024-09-06 Yanxu Chen , Linshu Huang , Tian Gou

Speech emotion recognition (SER) in naturalistic conditions presents a significant challenge for the speech processing community. Challenges include disagreement in labeling among annotators and imbalanced data distributions. This paper…

Machine Learning · Computer Science 2025-06-13 Thanathai Lertpetchpun , Tiantian Feng , Dani Byrd , Shrikanth Narayanan

We present Task 5 of the DCASE 2025 Challenge: an Audio Question Answering (AQA) benchmark spanning multiple domains of sound understanding. This task defines three QA subsets (Bioacoustics, Temporal Soundscapes, and Complex QA) to test…

The development of artificial intelligent composition has resulted in the increasing popularity of machine-generated pieces, with frequent copyright disputes consequently emerging. There is an insufficient amount of research on the…

Artificial Intelligence · Computer Science 2020-10-16 Yang Deng , Ziyao Xu , Li Zhou , Huanping Liu , Anqi Huang

Computational harmony analysis is important for MIR tasks such as automatic segmentation, corpus analysis and automatic chord label estimation. However, recent research into the ambiguous nature of musical harmony, causing limited…

Sound · Computer Science 2023-10-18 Hendrik Vincent Koops , Gianluca Micchi , Ilaria Manco , Elio Quinton

This paper proposes a benchmark of submissions to Detection and Classification Acoustic Scene and Events 2021 Challenge (DCASE) Task 4 representing a sampling of the state-of-the-art in Sound Event Detection task. The submissions are…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-02 Francesca Ronchini , Romain Serizel

Image aesthetics assessment (IAA) is a challenging task due to its highly subjective nature. Most of the current studies rely on large-scale datasets (e.g., AVA and AADB) to learn a general model for all kinds of photography images.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Ran Yi , Haoyuan Tian , Zhihao Gu , Yu-Kun Lai , Paul L. Rosin

Music has been identified as a promising medium to enhance the accessibility and experience of visual art for people who are blind or have low vision (BLV). However, composing music and designing soundscapes for visual art is a…

Human-Computer Interaction · Computer Science 2024-05-24 Stephen James Krol , Maria Teresa Llano , Matthew Butler , Cagatay Goncu