English
Related papers

Related papers: Untangling Phase and Time in Monophonic Sounds

200 papers

The advent of novel nonlinear materials has stirred unprecedented interest in exploring the use of temporal inhomogeneities to achieve novel forms of wave control, amidst the greater vision of engineering metamaterials across both space and…

Optics · Physics 2022-11-24 Emanuele Galiffi , Shixiong Yin , Andrea Alù

Synthetic magnetism has been recently realized using spatiotemporal modulation patterns, producing non-reciprocal steering of charge-neutral particles such as photons and phonons. Here, we design and experimentally demonstrate a…

Applied Physics · Physics 2022-10-03 Zhaoxian Chen , Zhengwei Li , Jingkai Weng , Bin Liang , Yanqing Lu , Jianchun Cheng , Andrea Alu

We present a novel model designed for resource-efficient multichannel speech enhancement in the time domain, with a focus on low latency, lightweight, and low computational requirements. The proposed model incorporates explicit spatial and…

Sound · Computer Science 2024-01-17 Ashutosh Pandey , Buye Xu

Acoustic holograms have promising applications in sound-field reconstruction, particle manipulation, ultrasonic haptics and therapy. This paper reports on the theoretical, numerical, and experimental investigation of multiplexed acoustic…

FM Synthesis is a well-known algorithm used to generate complex timbre from a compact set of design primitives. Typically featuring a MIDI interface, it is usually impractical to control it from an audio source. On the other hand,…

Sound · Computer Science 2022-08-15 Franco Caspe , Andrew McPherson , Mark Sandler

Acoustic metasurfaces manipulate waves with specially designed structures and achieve properties that natural materials cannot offer. Similar surfaces work in audio frequency range as well and lead to marvelous acoustic phenomena that can…

Applied Physics · Physics 2017-09-13 Shuping Wang , Jiancheng Tao , Xiaojun Qiu , Jianchun Cheng

Mandarin Chinese is characterized by being a tonal language; the pitch (or $F_0$) of its utterances carries considerable linguistic information. However, speech samples from different individuals are subject to changes in amplitude and…

Convolutional layers with 1-D filters are often used as frontend to encode audio signals. Unlike fixed time-frequency representations, they can adapt to the local characteristics of input data. However, 1-D filters on raw audio are hard to…

Sound · Computer Science 2024-09-02 Daniel Haider , Felix Perfler , Vincent Lostanlen , Martin Ehler , Peter Balazs

Video to sound generation aims to generate realistic and natural sound given a video input. However, previous video-to-sound generation methods can only generate a random or average timbre without any controls or specializations of the…

Multimedia · Computer Science 2022-11-22 Chenye Cui , Yi Ren , Jinglin Liu , Rongjie Huang , Zhou Zhao

Quantum manipulation of individual phonons could offer new resources for studying fundamental physics and creating an innovative platform in quantum information science. Here, we propose to generate quantum states of strongly correlated…

Quantum Physics · Physics 2021-08-02 Yuangang Deng , Tao Shi , Su Yi

Phase retrieval is in general a non-convex and non-linear task and the corresponding algorithms struggle with the issue of local minima. We consider the case where the measurement samples within typically very small and disconnected subsets…

Signal Processing · Electrical Eng. & Systems 2022-06-28 Jonas Kornprobst , Alexander Paulus , Josef Knapp , Thomas F. Eibert

We introduce VampNet, a masked acoustic token modeling approach to music synthesis, compression, inpainting, and variation. We use a variable masking schedule during training which allows us to sample coherent music from the model by…

Sound · Computer Science 2023-07-13 Hugo Flores Garcia , Prem Seetharaman , Rithesh Kumar , Bryan Pardo

This paper introduces an extendable modular system that compiles a range of music feature extraction models to aid music information retrieval research. The features include musical elements like key, downbeats, and genre, as well as audio…

Sound · Computer Science 2025-08-08 Anuradha Chopra , Abhinaba Roy , Dorien Herremans

The unique conduction properties of condensed matter systems with topological order have recently inspired a quest for similar effects in classical wave phenomena. Acoustic topological insulators, in particular, hold the promise to…

Mesoscale and Nanoscale Physics · Physics 2016-07-13 Romain Fleury , Alex Khanikaev , Andrea Alu

We consider the resonance and scattering properties of a composite medium containing scatterers whose properties are modulated in time. When excited with an incident wave of a single frequency, the scattered field consists of a family of…

Mesoscale and Nanoscale Physics · Physics 2024-08-06 Erik Orvehed Hiltunen , Bryn Davies

Speech enhancement is crucial for ubiquitous human-computer interaction. Recently, ultrasound-based acoustic sensing has emerged as an attractive choice for speech enhancement because of its superior ubiquity and performance. However, due…

Sound · Computer Science 2025-05-20 Luca Jiang-Tao Yu , Running Zhao , Sijie Ji , Edith C. H. Ngai , Chenshu Wu

Historically, most speech models in machine-learning have used the mel-spectrogram as a speech representation. Recently, discrete audio tokens produced by neural audio codecs have become a popular alternate speech representation for speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-05 Ryan Langman , Ante Jukić , Kunal Dhawan , Nithin Rao Koluguri , Jason Li

In this paper, a deep-learning-based method for sound field reconstruction is proposed. It is shown the possibility to reconstruct the magnitude of the sound pressure in the frequency band 30-300 Hz for an entire room by using a very low…

Sound · Computer Science 2020-08-07 Francesc Lluís , Pablo Martínez-Nuevo , Martin Bo Møller , Sven Ewan Shepstone

This study investigates phase reconstruction for deep learning based monaural talker-independent speaker separation in the short-time Fourier transform (STFT) domain. The key observation is that, for a mixture of two sources, with their…

Sound · Computer Science 2018-11-26 Zhong-Qiu Wang , Ke Tan , DeLiang Wang

Voice Conversion (VC) converts the voice of a source speech to that of a target while maintaining the source's content. Speech can be mainly decomposed into four components: content, timbre, rhythm and pitch. Unfortunately, most related…

Sound · Computer Science 2023-06-22 Zhonghua Liu , Shijun Wang , Ning Chen
‹ Prev 1 4 5 6 7 8 10 Next ›