English
Related papers

Related papers: The ICASSP 2026 Automatic Song Aesthetics Evaluati…

200 papers

In this article we present an account of the state-of-the-art in acoustic scene classification (ASC), the task of classifying environments from the sounds they produce. Starting from a historical review of previous research in this area, we…

Sound · Computer Science 2015-04-08 Daniele Barchiesi , Dimitrios Giannoulis , Dan Stowell , Mark D. Plumbley

The aesthetic quality assessment task is crucial for developing a human-aligned quantitative evaluation system for AIGC. However, its inherently complex nature, spanning visual perception, cognition, and emotion, poses fundamental…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Henglin Liu , Nisha Huang , Chang Liu , Jiangpeng Yan , Huijuan Huang , Jixuan Ying , Tong-Yee Lee , Pengfei Wan , Xiangyang Ji

Human beings often assess the aesthetic quality of an image coupled with the identification of the image's semantic content. This paper addresses the correlation issue between automatic aesthetic quality assessment and semantic recognition.…

Computer Vision and Pattern Recognition · Computer Science 2017-04-05 Yueying Kao , Ran He , Kaiqi Huang

Singing Accompaniment Generation (SAG) is the process of generating instrumental music for a given clean vocal input. However, existing SAG techniques use source-separated vocals as input and overfit to separation artifacts. This creates a…

Sound · Computer Science 2025-09-18 Junan Zhang , Yunjia Zhang , Xueyao Zhang , Zhizheng Wu

This paper reviews the challenge on Sparse Neural Rendering that was part of the Advances in Image Manipulation (AIM) workshop, held in conjunction with ECCV 2024. This manuscript focuses on the competition set-up, the proposed methods and…

In the fields of Experimental and Computational Aesthetics, numerous image datasets have been created over the last two decades. In the present work, we provide a comparative overview of twelve image datasets that include aesthetic ratings…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Ralf Bartho , Katja Thoemmes , Christoph Redies

Neural network models for audio tasks, such as automatic speech recognition (ASR) and acoustic scene classification (ASC), are susceptible to noise contamination for real-life applications. To improve audio quality, an enhancement module,…

Audio tagging aims to infer descriptive labels from audio clips. Audio tagging is challenging due to the limited size of data and noisy labels. In this paper, we describe our solution for the DCASE 2018 Task 2 general audio tagging…

Computer Vision and Pattern Recognition · Computer Science 2019-07-24 Kele Xu , Boqing Zhu , Qiuqiang Kong , Haibo Mi , Bo Ding , Dezhi Wang , Huaimin Wang

This paper presents our work for the ICASSP 2026 Environmental Sound Deepfake Detection (ESDD) Challenge. The challenge is based on the large-scale EnvSDD dataset that consists of various synthetic environmental sounds. We focus on…

Sound · Computer Science 2025-12-09 Candy Olivia Mawalim , Haotian Zhang , Shogo Okada

Recent advancements in audio-visual generative modeling have been propelled by progress in deep learning and the availability of data-rich benchmarks. However, the growth is not attributed solely to models and benchmarks. Universally…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Lucas Goncalves , Prashant Mathur , Chandrashekhar Lavania , Metehan Cekic , Marcello Federico , Kyu J. Han

In industry, machine anomalous sound detection (ASD) is in great demand. However, collecting enough abnormal samples is difficult due to the high cost, which boosts the rapid development of unsupervised ASD algorithms. Autoencoder (AE)…

Sound · Computer Science 2023-11-16 Yifan Zhou , Dongxing Xu , Haoran Wei , Yanhua Long

Audio-to-score alignment (A2SA) is a multimodal task consisting in the alignment of audio signals to music scores. Recent literature confirms the benefits of Automatic Music Transcription (AMT) for A2SA at the frame-level. In this work, we…

Sound · Computer Science 2022-01-03 Federico Simonetta , Stavros Ntalampiras , Federico Avanzini

Many speech enhancement methods try to learn the relationship between noisy and clean speech, obtained using an acoustic room simulator. We point out several limitations of enhancement methods relying on clean speech targets; the goal of…

Computation and Language · Computer Science 2018-12-26 Geonmin Kim , Hwaran Lee , Bo-Kyeong Kim , Sang-Hoon Oh , Soo-Young Lee

AI systems for high quality music generation typically rely on extremely large musical datasets to train the AI models. This creates barriers to generating music beyond the genres represented in dominant datasets such as Western Classical…

Sound · Computer Science 2024-07-19 Nick Bryan-Kinns , Zijin Li

SemEval-2024 Task 8 provides a challenge to detect human-written and machine-generated text. There are 3 subtasks for different detection scenarios. This paper proposes a system that mainly deals with Subtask B. It aims to detect if given…

Computation and Language · Computer Science 2024-04-02 Renhua Gu , Xiangfeng Meng

The ACM Recommender Systems Challenge 2018 focused on the task of automatic music playlist continuation, which is a form of the more general task of sequential recommendation. Given a playlist of arbitrary length with some additional…

Information Retrieval · Computer Science 2019-09-04 Hamed Zamani , Markus Schedl , Paul Lamere , Ching-Wei Chen

The article explores the use of agentic software engineering (ASE) in the development of innovative audio software. It begins with a review of background work that lays out the challenges of longevity, interoperability and barriers to entry…

Software Engineering · Computer Science 2026-05-15 Matthew John Yee-King

This paper develops automatic song translation (AST) for tonal languages and addresses the unique challenge of aligning words' tones with melody of a song in addition to conveying the original meaning. We propose three criteria for…

Computation and Language · Computer Science 2022-03-28 Fenfei Guo , Chen Zhang , Zhirui Zhang , Qixin He , Kejun Zhang , Jun Xie , Jordan Boyd-Graber

It is well established that listening to music is an issue for those with hearing loss, and hearing aids are not a universal solution. How can machine learning be used to address this? This paper details the first application of the open…

Recent advances in instruction-guided image editing underscore the need for effective automated evaluation. While Vision-Language Models (VLMs) have been explored as judges, open-source models struggle with alignment, and proprietary models…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Sherry X. Chen , Yi Wei , Luowei Zhou , Suren Kumar
‹ Prev 1 3 4 5 6 7 10 Next ›