English
Related papers

Related papers: SELEBI: Percussion-aware Time Stretching via Selec…

200 papers

Vocal dereverberation remains a challenging task in audio processing, particularly for real-time applications where both accuracy and efficiency are crucial. Traditional deep learning approaches often struggle to suppress reverberation…

Sound · Computer Science 2025-10-02 Daniel G. Williams

Speaker extraction aims to mimic humans' selective auditory attention by extracting a target speaker's voice from a multi-talker environment. It is common to perform the extraction in frequency-domain, and reconstruct the time-domain signal…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-20 Chenglin Xu , Wei Rao , Eng Siong Chng , Haizhou Li

Automatic speech emotion recognition (SER) is a challenging task that plays a crucial role in natural human-computer interaction. One of the main challenges in SER is data scarcity, i.e., insufficient amounts of carefully labeled data to…

Sound · Computer Science 2021-08-17 Sarala Padi , Seyed Omid Sadjadi , Dinesh Manocha , Ram D. Sriram

Piezoelectric Micromachined Ultrasonic Transducers (PMUTs) are essential for next-generation ultrasonic sensing and imaging due to their bidirectional electromechanical behavior, compact design, and compatibility with low-voltage…

Inference-time scaling offers a versatile paradigm for aligning visual generative models with downstream objectives without parameter updates. However, existing approaches that optimize the high-dimensional initial noise suffer from severe…

Machine Learning · Computer Science 2026-02-04 Jinyan Ye , Zhongjie Duan , Zhiwen Li , Cen Chen , Daoyuan Chen , Yaliang Li , Yingda Chen

Despite the central role that melody plays in music perception, it remains an open challenge in music information retrieval to reliably detect the notes of the melody present in an arbitrary music recording. A key challenge in melody…

Sound · Computer Science 2022-12-06 Chris Donahue , John Thickstun , Percy Liang

Compositionality in knowledge and language--the ability to represent complex concepts as a combination of simpler ones--is a hallmark of human cognition and communication. Despite recent advances, deep neural networks still struggle to…

Machine Learning · Computer Science 2025-12-01 Rafael Elberg , Felipe del Rio , Mircea Petrache , Denis Parra

Many methods of sound event detection (SED) based on machine learning regard a segmented time frame as one data sample to model training. However, the sound durations of sound events vary greatly depending on the sound event class, e.g.,…

Computation of the instantaneous phase and amplitude via the Hilbert Transform is a powerful tool of data analysis. This approach finds many applications in various science and engineering branches but is not proper for causal estimation…

Neurons and Cognition · Quantitative Biology 2021-05-24 Michael Rosenblum , Arkady Pikovsky , Andrea A. Kühn , Johannes L. Busch

This paper addresses the structurally-constrained sparse decomposition of multi-dimensional signals onto overcomplete families of vectors, called dictionaries. The contribution of the paper is threefold. Firstly, a generic spatio-temporal…

Data Structures and Algorithms · Computer Science 2016-10-03 Yoann Isaac , Quentin Barthélemy , Cédric Gouy-Pailler , Michèle Sebag , Jamal Atif

In speech-related classification tasks, frequency-domain acoustic features such as logarithmic Mel-filter bank coefficients (FBANK) and cepstral-domain acoustic features such as Mel-frequency cepstral coefficients (MFCC) are often used.…

Sound · Computer Science 2022-06-20 Yikang Wang , Hiromitsu Nishizaki

Reverberation encodes spatial information regarding the acoustic source environment, yet traditional Speech Restoration (SR) usually completely removes reverberation. We propose ReverbMiipher, an SR model extending parametric resynthesis…

We propose a novel numerical approach for the optimal design of wide-area heterogeneous electromagnetic metasurfaces beyond the conventionally used unit-cell approximation. The proposed method exploits the combination of Rigorous Coupled…

Applied Physics · Physics 2017-09-19 Krupali D. Donda , Ravi S. Hegde

It has been shown that analog-to-information con- version (AIC) is an efficient scheme to perform sub-Nyquist sampling of pulsed radar echoes. However, it is often impractical, if not infeasible, to reconstruct full-range Nyquist samples…

Information Theory · Computer Science 2015-03-03 Suling Zhang , Feng Xi , Shengyao Chen , Yimin D. Zhang , Zhong Liu

Selective auditory attention decoding aims to identify the speaker of interest from listeners' neural signals, such as electroencephalography (EEG), in the presence of multiple concurrent speakers. Most existing methods operate at the…

Signal Processing · Electrical Eng. & Systems 2026-02-17 Yuanyuan Yao , Simon Geirnaert , Tinne Tuytelaars , Alexander Bertrand

The motion of a mechanical resonator is intrinsically decomposed over a collection of normal modes of vibration. When the resonator is used as a sensor, its multimode nature often deteriorates or limits its performance and sensitivity. This…

Applied Physics · Physics 2022-02-23 Giada La Gala , John P. Mathew , Pascal Neveu , Ewold Verhagen

While log-amplitude mel-spectrogram has widely been used as the feature representation for processing speech based on deep learning, the effectiveness of another aspect of speech spectrum, i.e., phase information, was shown recently for…

Sound · Computer Science 2022-05-02 Shunsuke Hidaka , Kohei Wakamiya , Tokihiko Kaburagi

We present a novel learning-based modal sound synthesis approach that includes a mixed vibration solver for modal analysis and an end-to-end sound radiation network for acoustic transfer. Our mixed vibration solver consists of a 3D sparse…

Sound · Computer Science 2022-05-31 Xutong Jin , Sheng Li , Guoping Wang , Dinesh Manocha

An initial real-time speech enhancement method is presented to reduce the effects of additive noise. The method operates in the frequency domain and is a form of spectral subtraction. Initially, minimum statistics are used to generate an…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-22 Georgios Ioannides , Vasilios Rallis

In this paper, we present a statistical beamforming algorithm as a pre-processing step for robust automatic speech recognition (ASR). By modeling the target speech as a non-stationary Laplacian distribution, a mask-based statistical…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-08 Ui-Hyeop Shin , Hyung-Min Park
‹ Prev 1 8 9 10 Next ›