English
Related papers

Related papers: Audio Foundation Models Outperform Symbolic Repres…

200 papers

The relationship between perceptual loudness and physical attributes of sound is an important subject in both computer music and psychoacoustics. Early studies of "equal-loudness contour" can trace back to the 1920s and the measured…

Sound · Computer Science 2022-11-01 Yang Qu , Yutian Qin , Lecheng Chao , Hangkai Qian , Ziyu Wang , Gus Xia

In this work, we investigate multimodal foundation models (MFMs) for EmoFake detection (EFD) and hypothesize that they will outperform audio foundation models (AFMs). MFMs due to their cross-modal pre-training, learns emotional patterns…

Automatically estimating the performance difficulty of a music piece represents a key process in music education to create tailored curricula according to the individual needs of the students. Given its relevance, the Music Information…

Sound · Computer Science 2025-05-30 Pedro Ramoneda , Minhee Lee , Dasaem Jeong , J. J. Valero-Mas , Xavier Serra

Timbre and pitch are the two main perceptual properties of musical sounds. Depending on the target applications, we sometimes prefer to focus on one of them, while reducing the effect of the other. Researchers have managed to hand-craft…

Sound · Computer Science 2018-11-09 Yun-Ning Hung , Yi-An Chen , Yi-Hsuan Yang

Symbolic music generation has seen rapid progress with artificial neural networks, yet remains underexplored in the biologically plausible domain of spiking neural networks (SNNs), where both standardized benchmarks and comprehensive…

Sound · Computer Science 2025-08-28 Qian Liang , Menghaoran Tang , Yi Zeng

Recently, self-supervised pre-training has shown significant improvements in many areas of machine learning, including speech and NLP. We propose using large self-supervised pre-trained models for both audio and text modality with…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-24 Krishna D N

In this study, we investigate multimodal foundation models (MFMs) for emotion recognition from non-verbal sounds. We hypothesize that MFMs, with their joint pre-training across multiple modalities, will be more effective in non-verbal…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 Orchid Chetia Phukan , Mohd Mujtaba Akhtar , Girish , Swarup Ranjan Behera , Sishir Kalita , Arun Balaji Buduru , Rajesh Sharma , S. R Mahadeva Prasanna

Over the past few decades, computational methods have been developed to estimate perceptual audio quality. These methods, also referred to as objective quality measures, are usually developed and intended for a specific application domain.…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-25 Matteo Torcoli , Thorsten Kastner , Jürgen Herre

Recent advances in music foundation models have improved audio representation learning, yet their effectiveness across diverse musical traditions remains limited. We introduce CultureMERT-95M, a multi-culturally adapted foundation model…

Sound · Computer Science 2025-06-24 Angelos-Nikolaos Kanatas , Charilaos Papaioannou , Alexandros Potamianos

In a noisy environment, a lossy speech signal can be automatically restored by a listener if he/she knows the language well. That is, with the built-in knowledge of a "language model", a listener may effectively suppress noise interference…

Machine Learning · Computer Science 2019-07-03 Chien-Feng Liao , Yu Tsao , Xugang Lu , Hisashi Kawai

Speech synthesis and music audio generation from symbolic input differ in many aspects but share some similarities. In this study, we investigate how text-to-speech synthesis techniques can be used for piano MIDI-to-audio synthesis tasks.…

Sound · Computer Science 2022-02-25 Erica Cooper , Xin Wang , Junichi Yamagishi

Existing audio-to-MIDI tools extract notes but discard the timbral characteristics that define an instrument's identity. We present Instrumental, a system that recovers continuous synthesizer parameters from audio by coupling a…

Sound · Computer Science 2026-03-18 Philipp Bogdan

With the growing amount of musical data available, automatic instrument recognition, one of the essential problems in Music Information Retrieval (MIR), is drawing more and more attention. While automatic recognition of single instruments…

Sound · Computer Science 2023-06-16 Lifan Zhong , Erica Cooper , Junichi Yamagishi , Nobuaki Minematsu

We introduce Multi-level feature Fusion-based Periodicity Analysis Model (MF-PAM), a novel deep learning-based pitch estimation model that accurately estimates pitch trajectory in noisy and reverberant acoustic environments. Our model…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-11 Woo-Jin Chung , Doyeon Kim , Soo-Whan Chung , Hong-Goo Kang

We introduce a structure-aware approach for symbolic piano accompaniment that decouples high-level planning from note-level realization. A lightweight transformer predicts an interpretable, per-measure style plan conditioned on…

Sound · Computer Science 2026-02-18 Wanyu Zang , Yang Yu , Meng Yu

We consider audio decoding as an inverse problem and solve it through diffusion posterior sampling. Explicit conditioning functions are developed for input signal measurements provided by an example of a transform domain perceptual audio…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-13 Pedro J. Villasana T. , Lars Villemoes , Janusz Klejsa , Per Hedelin

We present our preliminary work to determine if patient's vocal acoustic, linguistic, and facial patterns could predict clinical ratings of depression severity, namely Patient Health Questionnaire depression scale (PHQ-8). We proposed a…

Computer Vision and Pattern Recognition · Computer Science 2017-12-01 Aven Samareh , Yan Jin , Zhangyang Wang , Xiangyu Chang , Shuai Huang

Music Information Retrieval (MIR) research is increasingly leveraging representation learning to obtain more compact, powerful music audio representations for various downstream MIR tasks. However, current representation evaluation methods…

Sound · Computer Science 2023-12-13 Christos Plachouras , Pablo Alonso-Jiménez , Dmitry Bogdanov

Symbolic music representation is a fundamental challenge in computational musicology. While grid-based representations effectively preserve pitch-time spatial correspondence, their inherent data sparsity leads to low encoding efficiency.…

Sound · Computer Science 2026-01-29 Lekai Qian , Haoyu Gu , Dehan Li , Boyu Cao , Qi Liu

We explore the use of neural synthesis for acoustic guitar from string-wise MIDI input. We propose four different systems and compare them with both objective metrics and subjective evaluation against natural audio and a sample-based…

Sound · Computer Science 2023-09-15 Nicolas Jonason , Xin Wang , Erica Cooper , Lauri Juvela , Bob L. T. Sturm , Junichi Yamagishi
‹ Prev 1 4 5 6 7 8 10 Next ›