English
Related papers

Related papers: Non-locally averaged pruned reassigned spectrogram…

200 papers

Speech enhancement in hearing aids remains a difficult task in nonstationary acoustic environments, mainly because current signal processing algorithms rely on fixed, manually tuned parameters that cannot adapt in situ to different users or…

In this paper, we report on the observation of nonlinear effects in a nanostrip phononic metasurface (NPM) that enable the tuning of resonance frequencies at 1.42 GHz. The NPM resonator made of periodic nanostrip array is fabricated on a…

Applied Physics · Physics 2021-03-31 Feng Gao , Amine Bermak , Sarah Benchabane , Marina Raschetti , Abdelkrim Khelif

Deep speaker embedding models have been commonly used as a building block for speaker diarization systems; however, the speaker embedding model is usually trained according to a global loss defined on the training data, which could be…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-26 Jixuan Wang , Xiong Xiao , Jian Wu , Ranjani Ramamurthy , Frank Rudzicz , Michael Brudno

There exist many scenarios where pixel information is available only on a non-regular subset of pixel positions. For further processing, however, it is required to reconstruct such images on a regular grid. Besides many other algorithms,…

Image and Video Processing · Electrical Eng. & Systems 2022-04-08 Markus Jonscher , Jürgen Seiler , André Kaup

Language-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches, such as time-frequency masking, to separate target sounds…

Sound · Computer Science 2025-01-10 Yi Yuan , Xubo Liu , Haohe Liu , Mark D. Plumbley , Wenwu Wang

In this paper, we are concerned with the recovery of the geometric shapes of inhomogeneous inclusions from the associated far field data in electrostatics and acoustic scattering. We present a local resolution analysis and show that the…

Analysis of PDEs · Mathematics 2021-08-17 Habib Ammari , Yat Tin Chow , Hongyu Liu

Accurately representing the sound field with the high spatial resolution is critical for immersive and interactive sound field reproduction technology. To minimize experimental effort, data-driven methods have been proposed to estimate…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-10 Zining Liang , Wen Zhang , Thushara D. Abhayapala

Gaussian Processes (GPs) have been widely used in machine learning to model distributions over functions, with applications including multi-modal regression, time-series prediction, and few-shot learning. GPs are particularly useful in the…

A Prompt-based Text-To-Speech model allows a user to control different aspects of speech, such as speaking rate and perceived gender, through natural language instruction. Although user-friendly, such approaches are on one hand constrained:…

Computation and Language · Computer Science 2025-07-14 Atli Sigurgeirsson , Simon King

Speaker Identification refers to the process of identifying a person using one's voice from a collection of known speakers. Environmental noise, reverberation and distortion make the task of automatic speaker identification challenging as…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-01 Sabbir Ahmed , Nursadul Mamun , Md Azad Hossain

Recently, graph neural networks (GNNs) have shown prominent performance in graph representation learning by leveraging knowledge from both graph structure and node features. However, most of them have two major limitations. First, GNNs can…

Machine Learning · Computer Science 2022-06-20 Wentao Zhang , Zeang Sheng , Mingyu Yang , Yang Li , Yu Shen , Zhi Yang , Bin Cui

We propose to use Gaussian process regression to accurately estimate the diffusion MRI signal at arbitrary locations in q-space. By estimating the signal on a grid, we can do synthetic diffusion spectrum imaging: reconstructing the ensemble…

Applications · Statistics 2020-01-03 Jens Sjölund , Anders Eklund , Evren Özarslan , Hans Knutsson

3D Gaussian Splatting (3D-GS) enables efficient novel view synthesis, but treats all frequencies uniformly, making it difficult to separate coarse structure from fine detail. Recent works have started to exploit frequency signals, but lack…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Yishai Lavi , Leo Segre , Shai Avidan

Personalized TTS is an exciting and highly desired application that allows users to train their TTS voice using only a few recordings. However, TTS training typically requires many hours of recording and a large model, making it unsuitable…

Sound · Computer Science 2023-03-22 Sung-Feng Huang , Chia-ping Chen , Zhi-Sheng Chen , Yu-Pao Tsai , Hung-yi Lee

Recent years have witnessed the rapid emergence of 3D Gaussian splatting (3DGS) as a powerful approach for 3D reconstruction and novel view synthesis. Its explicit representation with Gaussian primitives enables fast training, real-time…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Haato Watanabe , Nobuyuki Umetani

We propose UnitSpeech, a speaker-adaptive speech synthesis method that fine-tunes a diffusion-based text-to-speech (TTS) model using minimal untranscribed data. To achieve this, we use the self-supervised unit representation as a pseudo…

Sound · Computer Science 2023-06-29 Heeseung Kim , Sungwon Kim , Jiheum Yeom , Sungroh Yoon

Previous methods utilize the Neural Radiance Field (NeRF) for panoptic lifting, while their training and rendering speed are unsatisfactory. In contrast, 3D Gaussian Splatting (3DGS) has emerged as a prominent technique due to its rapid…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Yu Wang , Xiaobao Wei , Ming Lu , Guoliang Kang

Dual-band plasmonic nanoantennas, exhibiting two widely separated user-defined resonances, are fundamental building blocks for the investigation and optimization of plasmon-enhanced optical phenomena, including photoluminescence, Raman…

Optics · Physics 2026-05-04 Huatian Hu , Zhiwei Hu , Christophe Galland , Wen Chen

Restoring degraded music signals is essential to enhance audio quality for downstream music manipulation. Recent diffusion-based music restoration methods have demonstrated impressive performance, and among them, diffusion posterior…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-14 Carlos Hernandez-Olivan , Koichi Saito , Naoki Murata , Chieh-Hsin Lai , Marco A. Martínez-Ramirez , Wei-Hsiang Liao , Yuki Mitsufuji

Miniatured computational spectrometers, distinguished by their compact size and lightweight, have shown great promise for on-chip and portable applications in the fields of healthcare, environmental monitoring, food safety, and industrial…

Optics · Physics 2025-08-19 Linjun Zhai , Baolei Liu , Muchen Zhu , Yao Wang , Chaohao Chen , Zhaohua Yang , Lan Fu , Fan Wang