English
Related papers

Related papers: Mesostructures: Beyond Spectrogram Loss in Differe…

200 papers

We present a hybrid neural network and rule-based system that generates pop music. Music produced by pure rule-based systems often sounds mechanical. Music produced by machine learning sounds better, but still lacks hierarchical temporal…

Sound · Computer Science 2017-10-09 Yifei Teng , An Zhao , Camille Goudeseune

We study indeterminacies in realization of ornaments and how they can be incorporated in a stochastic performance model applicable for music information processing such as score-performance matching. We point out the importance of temporal…

Artificial Intelligence · Computer Science 2016-08-04 Eita Nakamura , Nobutaka Ono , Shigeki Sagayama , Kenji Watanabe

Spatially localized oscillations in periodically forced systems are intriguing phenomena. They may occur in spatially homogeneous media (oscillons), but quite often emerge in heterogeneous media, such as the auditory system, where localized…

Pattern Formation and Solitons · Physics 2020-04-21 Yuval Edri , Ehud Meron , Arik Yochelis

Automatic lyrics to polyphonic audio alignment is a challenging task not only because the vocals are corrupted by background music, but also there is a lack of annotated polyphonic corpus for effective acoustic modeling. In this work, we…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-26 Chitralekha Gupta , Emre Yılmaz , Haizhou Li

Self-supervised pre-training models have been used successfully in several machine learning domains. However, only a tiny amount of work is related to music. In our work, we treat a spectrogram of music as a series of patches and design a…

Sound · Computer Science 2022-10-31 Leyi Zhao , Yi Li

Music segmentation refers to the dual problem of identifying boundaries between, and labeling, distinct music segments, e.g., the chorus, verse, bridge etc. in popular music. The performance of a range of music segmentation algorithms has…

Sound · Computer Science 2021-08-31 Matthew C. McCallum

Generating musical audio directly with neural networks is notoriously difficult because it requires coherently modeling structure at many different timescales. Fortunately, most music is also highly structured and can be represented as…

Medical image segmentation remains challenging in low-data regimes, where scarce annotations often yield poor generalization and ambiguous boundaries with missing fine structures. Recent self-supervised pretraining has improved…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Zhiquan Chen , Haitao Wang , Guowei Zou , Hejun Wu

The audio spectrogram is a time-frequency representation that has been widely used for audio classification. One of the key attributes of the audio spectrogram is the temporal resolution, which depends on the hop size used in the Short-Time…

Sound · Computer Science 2024-01-15 Haohe Liu , Xubo Liu , Qiuqiang Kong , Wenwu Wang , Mark D. Plumbley

In modelling time series data coming from different sources, frequencies can easily vary since some variable can be measured at higher frequencies, others, at lower frequencies. Given data measured over spatial units and at varying…

Methodology · Statistics 2025-03-05 Vladimir A. Malabanan , Joseph Ryan G. Lansangan , Erniel B. Barrios

Microstructured materials, such as architected metamaterials and phononic crystals, exhibit complex wave propagation phenomena due to their internal structure. While full-scale numerical simulations can capture these effects, they are…

Computational Physics · Physics 2025-05-21 Gianluca Rizzi , Angela Madeo

Connecting large libraries of digitized audio recordings to their corresponding sheet music images has long been a motivation for researchers to develop new cross-modal retrieval systems. In recent years, retrieval systems based on…

Information Retrieval · Computer Science 2019-06-27 Stefan Balke , Matthias Dorfer , Luis Carvalho , Andreas Arzt , Gerhard Widmer

We propose multi-microphone complex spectral mapping, a simple way of applying deep learning for time-varying non-linear beamforming, for speaker separation in reverberant conditions. We aim at both speaker separation and dereverberation.…

Sound · Computer Science 2021-05-25 Zhong-Qiu Wang , Peidong Wang , DeLiang Wang

Hand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. However, these manually designed spatial features are hard to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-16 Rongzhi Gu , Shi-Xiong Zhang , Lianwu Chen , Yong Xu , Meng Yu , Dan Su , Yuexian Zou , Dong Yu

Languages have long been described according to their perceived rhythmic attributes. The associated typologies are of interest in psycholinguistics as they partly predict newborns' abilities to discriminate between languages and provide…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-29 François Deloche , Laurent Bonnasse-Gahot , Judit Gervain

Generative models in vision have seen rapid progress due to algorithmic improvements and the availability of high-quality image datasets. In this paper, we offer contributions in both these areas to enable similar progress in audio…

Machine Learning · Computer Science 2017-04-06 Jesse Engel , Cinjon Resnick , Adam Roberts , Sander Dieleman , Douglas Eck , Karen Simonyan , Mohammad Norouzi

The representation of basic elements of music in terms of discrete audio signals is often used in software for musical creation and design. Nevertheless, there is no unified approach that relates these elements to the discrete samples of…

Supervised deep learning approaches to underdetermined audio source separation achieve state-of-the-art performance but require a dataset of mixtures along with their corresponding isolated source signals. Such datasets can be extremely…

Classical auditory-periphery models, exemplified by Bruce et al., 2018, provide high-fidelity simulations but are stochastic and computationally demanding, limiting large-scale experimentation and low-latency use. Prior neural encoders…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-23 Eylon Zohar , Israel Nelken , Boaz Rafaely

Music, being a multifaceted stimulus evolving at multiple timescales, modulates brain function in a manifold way that encompasses not only the distinct stages of auditory perception but also higher cognitive processes like memory and…

Neurons and Cognition · Quantitative Biology 2018-02-06 Dimitrios A. Adamos , Nikolaos Laskaris , Sifis Micheloyannis
‹ Prev 1 3 4 5 6 7 10 Next ›