中文
相关论文

相关论文: Audio Foundation Models Outperform Symbolic Repres…

200 篇论文

In this work we present a new approach for the task of predicting fingerings for piano music. While prior neural approaches have often treated this as a sequence tagging problem with independent predictions, we put forward a checklist…

机器学习 · 计算机科学 2022-09-14 Nikita Srivatsan , Taylor Berg-Kirkpatrick

Real-time music tracking systems follow a musical performance and at any time report the current position in a corresponding score. Most existing methods approach this problem exclusively in the audio domain, typically using online time…

声音 · 计算机科学 2025-05-09 Silvan Peter , Patricia Hu , Gerhard Widmer

Automatically generating symbolic music-music scores tailored to specific human needs-can be highly beneficial for musicians and enthusiasts. Recent studies have shown promising results using extensive datasets and advanced transformer…

声音 · 计算机科学 2024-07-08 Yangyang Shu , Haiming Xu , Ziqin Zhou , Anton van den Hengel , Lingqiao Liu

Deep neural networks have shown promise for music audio signal processing applications, often surpassing prior approaches, particularly as end-to-end models in the waveform domain. Yet results to date have tended to be constrained by low…

音频与语音处理 · 电气工程与系统科学 2020-06-11 William Mitchell , Scott H. Hawley

Non-intrusive assessment of speech quality and intelligibility is essential when clean reference signals are unavailable. In this work, we propose a multimodal framework that integrates audio features and visual cues to predict PESQ and…

音频与语音处理 · 电气工程与系统科学 2025-06-12 Shafique Ahmed , Ryandhimas E. Zezario , Nasir Saleem , Amir Hussain , Hsin-Min Wang , Yu Tsao

Self-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech. Although SSL has been proven effective in speech and audio, its…

Speech-based depression detection tools could aid early screening. Here, we propose an interpretable speech foundation model approach to enhance the clinical applicability of such tools. We introduce a speech-level Audio Spectrogram…

声音 · 计算机科学 2026-03-26 Qingkun Deng , Saturnino Luz , Sofia de la Fuente Garcia

Large deep-learning models for music, including those focused on learning general-purpose music audio representations, are often assumed to require substantial training data to achieve high performance. If true, this would pose challenges…

声音 · 计算机科学 2025-05-12 Christos Plachouras , Emmanouil Benetos , Johan Pauwels

This project performs multimodal sentiment analysis using the CMU-MOSEI dataset, using transformer-based models with early fusion to integrate text, audio, and visual modalities. We employ BERT-based encoders for each modality, extracting…

计算与语言 · 计算机科学 2025-07-16 Jugal Gajjar , Kaustik Ranaware

ITU-R BS.1387 states a method for objective assessment of perceived audio quality. This Recommendation, known also as PEAQ (Perceptual Evaluation of Audio Quality) is based on a psychoacoustic model of the human ear and was standardized by…

音频与语音处理 · 电气工程与系统科学 2019-10-31 Luis F. Abanto-Leon , Guillermo Kemper Vasquez , Joel Telles

Multimodal music emotion analysis leverages both audio and MIDI modalities to enhance performance. While mainstream approaches focus on complex feature extraction networks, we propose that shortening the length of audio sequence features to…

声音 · 计算机科学 2025-09-24 Dinghao Zou , Yicheng Gong , Xiaokang Li , Xin Cao , Sunbowen Lee

With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the dissonance between the…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Trevine Oorloff , Surya Koppisetti , Nicolò Bonettini , Divyaraj Solanki , Ben Colman , Yaser Yacoob , Ali Shahriyari , Gaurav Bharaj

Over the past decades, numerous Image Quality Assessment (IQA) models have emerged, aiming to predict the perceptual quality of images. However, individual models are often biased toward certain types of image content or distortions,…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Zhongling Wang , Raymond Zhou , Shahrukh Athar , Wenbo Yang , Zhou Wang

Quality assessment and aesthetics assessment aim to evaluate the perceived quality and aesthetics of visual content. Current learning-based methods suffer greatly from the scarcity of labeled data and usually perform sub-optimally in terms…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Qizhi Xie , Kun Yuan , Yunpeng Qu , Mingda Wu , Ming Sun , Chao Zhou , Jihong Zhu

Psychoacoustical so-called "timbre spaces" map perceptual similarity ratings of instrument sounds onto low-dimensional embeddings via multidimensional scaling, but suffer from scalability issues and are incapable of generalization. Recent…

声音 · 计算机科学 2025-07-11 Haokun Tian , Stefan Lattner , Charalampos Saitis

Auscultation is a vital diagnostic tool, yet its utility is often limited by subjective interpretation. While general-purpose Audio-Language Models (ALMs) excel in general domains, they struggle with the nuances of physiological signals. We…

We propose a model to estimate the fundamental frequency in monophonic audio, often referred to as pitch estimation. We acknowledge the fact that obtaining ground truth annotations at the required temporal and frequency resolution is a…

音频与语音处理 · 电气工程与系统科学 2020-09-07 Beat Gfeller , Christian Frank , Dominik Roblek , Matt Sharifi , Marco Tagliasacchi , Mihajlo Velimirović

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we bridge this critical gap by establishing a comprehensive…

Sound field reproduction methods based on numerical optimization, which aim to minimize the error between synthesized and desired sound fields, are useful in many practical scenarios because of their flexibility in the array geometry of…

音频与语音处理 · 电气工程与系统科学 2021-11-23 Shoichi Koyama , Keisuke Kimura , Natsuki Ueno

In deep learning research, many melody extraction models rely on redesigning neural network architectures to improve performance. In this paper, we propose an input feature modification and a training objective modification based on two…

声音 · 计算机科学 2023-08-08 Keren Shao , Ke Chen , Taylor Berg-Kirkpatrick , Shlomo Dubnov