English
Related papers

Related papers: Mesostructures: Beyond Spectrogram Loss in Differe…

200 papers

Subwavelength resonance is a vital acoustic phenomenon in contrasting media. The narrow bandgap width of single-layer resonator has prompted the exploration of multi-layer metamaterials as an effective alternative, which consist of…

Mathematical Physics · Physics 2024-11-15 Youjun Deng , Lingzheng Kong , Hongjie Li , Hongyu Liu , Liyan Zhu

Long-context modeling is essential for symbolic music generation, since motif repetition and developmental variation can span thousands of musical events, yet practical workflows frequently rely on resource-limited hardware. We introduce…

Sound · Computer Science 2026-03-03 Yungang Yi , Weihua Li , Matthew Kuo , Catherine Shi , Quan Bai

This paper proposes a novel framework for unsupervised audio source separation using a deep autoencoder. The characteristics of unknown source signals mixed in the mixed input is automatically by properly configured autoencoders implemented…

Sound · Computer Science 2014-12-24 Giljin Jang , Han-Gyu Kim , Yung-Hwan Oh

Binaural audio remains underexplored within the music information retrieval community. Motivated by the rising popularity of virtual and augmented reality experiences as well as potential applications to accessibility, we investigate how…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-02 Richa Namballa , Agnieszka Roginska , Magdalena Fuentes

With the proliferation of video platforms on the internet, recording musical performances by mobile devices has become commonplace. However, these recordings often suffer from degradation such as noise and reverberation, which negatively…

Sound · Computer Science 2023-08-25 Yunkee Chae , Junghyun Koo , Sungho Lee , Kyogu Lee

Music foundation models possess impressive music generation capabilities. When people compose music, they may infuse their understanding of music into their work, by using notes and intervals to craft melodies, chords to build progressions,…

Sound · Computer Science 2024-10-02 Megan Wei , Michael Freeman , Chris Donahue , Chen Sun

Generalized impedance boundary conditions are effective, approximate boundary conditions that describe scattering of waves in situations where the wave interaction with the material involves multiple scales. In particular, this includes…

Numerical Analysis · Mathematics 2020-05-29 Lehel Banjai , Christian Lubich , Joerg Nick

Extracting individual elements from music mixtures is a valuable tool for music production and practice. While neural networks optimized to mask or transform mixture spectrograms into the individual source(s) have been the leading approach,…

Sound · Computer Science 2025-11-26 Genís Plaja-Roglans , Yun-Ning Hung , Xavier Serra , Igor Pereira

Learning how to localize and separate individual object sounds in the audio channel of the video is a difficult task. Current state-of-the-art methods predict audio masks from artificially mixed spectrograms, known as Mix-and-Separate…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Tanzila Rahman , Leonid Sigal

Music source separation (MSS) is the task of separating a music piece into individual sources, such as vocals and accompaniment. Recently, neural network based methods have been applied to address the MSS problem, and can be categorized…

Sound · Computer Science 2021-02-22 Xuchen Song , Qiuqiang Kong , Xingjian Du , Yuxuan Wang

We study lightweight, elastic metamaterials consisting of tensegrity-inspired prisms, which present wide, low-frequency band gaps. For their realization, we alternate tensegrity elements with solid discs in periodic arrangements that we…

Recently, symbolic music generation has become a focus of numerous deep learning research. Structure as an important part of music, contributes to improving the quality of music, and an increasing number of works start to study the…

Sound · Computer Science 2024-10-16 Yishan Lv , Jing Luo , Boyuan Ju , Xinyu Yang

Audio-visual multi-modal modeling has been demonstrated to be effective in many speech related tasks, such as speech recognition and speech enhancement. This paper introduces a new time-domain audio-visual architecture for target speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-24 Jian Wu , Yong Xu , Shi-Xiong Zhang , Lian-Wu Chen , Meng Yu , Lei Xie , Dong Yu

Discovering and exploring the underlying structure of multi-instrumental music using learning-based approaches remains an open problem. We extend the recent MusicVAE model to represent multitrack polyphonic measures as vectors in a latent…

Machine Learning · Statistics 2018-06-04 Ian Simon , Adam Roberts , Colin Raffel , Jesse Engel , Curtis Hawthorne , Douglas Eck

At present, neural network-based models, including transformers, struggle to generate memorable and readily comprehensible music from unified and repetitive musical material due to a lack of understanding of musical structure. Consequently,…

Sound · Computer Science 2026-01-21 Shangxuan Luo , Joshua Reiss

In recent years, self-supervised learning has amassed significant interest for training deep neural representations without labeled data. One such self-supervised learning approach is masked spectrogram modeling, where the objective is to…

Sound · Computer Science 2025-09-24 Sarthak Yadav , Sergios Theodoridis , Zheng-Hua Tan

This work develops a dynamic homogenization approach for metamaterials. It finds an approximate macroscopic homogenized equation with constant coefficients posed in space and time; however, the resulting homogenized equation is higher order…

Analysis of PDEs · Mathematics 2022-06-23 Kshiteej Deshmukh , Timothy Breitzman , Kaushik Dayal

To achieve a flexible recommendation and retrieval system, it is desirable to calculate music similarity by focusing on multiple partial elements of musical pieces and allowing the users to select the element they want to focus on. A…

Sound · Computer Science 2024-04-11 Yuka Hashizume , Li Li , Atsushi Miyashita , Tomoki Toda

Temporal fluctuations in the phase of waves transmitted through a dynamic, strongly scattering, mesoscopic sample are investigated using ultrasonic waves, and compared with theoretical predictions based on circular Gaussian statistics. The…

Mesoscale and Nanoscale Physics · Physics 2009-11-13 M. L. Cowan , D. Anache-Ménier , W. K. Hildebrand , J. H. Page , B. A. van Tiggelen

Style transfer is a technique for combining two images based on the activations and feature statistics in a deep learning neural network architecture. This paper studies the analogous task in the audio domain and takes a critical look at…

Sound · Computer Science 2020-08-10 M. Huzaifah , L. Wyse