Related papers: Tunable add-drop filter using an active whispering…
Automatic drum transcription (ADT) is traditionally formulated as a discriminative task to predict drum events from audio spectrograms. In this work, we redefine ADT as a conditional generative task and introduce Noise-to-Notes (N2N), a…
Whispering gallery mode (WGM) microcavities strongly enhance nonlinear optical processes like optical frequency comb, Raman scattering and optomechanics, which nowadays enable cutting-edge applications in microwave synthesis, optical…
We report on fabrication of a microtoroid resonator of a high-quality factor (i. e., Q-factor of ~3.24x10^6 measured under the critical coupling condition) integrated in a microfluidic channel using femtosecond laser three-dimensional (3D)…
Materials' microstructure strongly influences its performance and is thus a critical aspect in design of functional materials. Previous efforts on microstructure mediated design mostly assume isotropy, which is not ideal when material…
We present our experiments on refractometric sensing with ultrahigh-Q, crystalline, birefringent magnesium fluoride (MgF$_2$) whispering gallery mode resonators. The difference to fused silica which is most commonly used for sensing…
Spotforming is a target-speaker extraction technique that uses multiple microphone arrays. This method applies beamforming (BF) to each microphone array, and the common components among the BF outputs are estimated as the target source.…
This paper introduces a novel low-latency online beamforming (BF) algorithm, named Modified Parametric Multichannel Wiener Filter (Mod-PMWF), for enhancing speech mixtures with unknown and varying number of speakers. Although conventional…
We present a design and implementation of frequency-tunable superconducting resonator. The resonance frequency tunability is achieved by flux-coupling a superconducting LC-loop to a current-biased feedline; the resulting screening current…
This paper presents a flexible thin-film underwater transducer based on a mesoporous PVDF membrane embedded with piezoelectrical-actuated microdomes. To enhance piezoelectric performance, ZnO nanoparticles were used as a sacrificial…
We demonstrate an integrated photonic circuit based on feed forward photonic meshes that can be programmed and reconfigured to perform arbitrary spectral filter functions. We investigate a subset of the available filter functions,…
In recent years, self-supervised learning (SSL) models have made significant progress in audio deepfake detection (ADD) tasks. However, existing SSL models mainly rely on large-scale real speech for pre-training and lack the learning of…
Automatic resonance alignment tuning is performed in high-order series coupled microring filters using a feedback system. By inputting only a reference wavelength, a filter is tuned such that passband ripples are dramatically reduced…
Adaptive filters (AFs) are vital for enhancing the performance of downstream tasks, such as speech recognition, sound event detection, and keyword spotting. However, traditional AF design prioritizes isolated signal-level objectives, often…
Real-world multimodal systems routinely face missing-input scenarios, and in reality, robots lose audio in a factory or a clinical record omits lab tests at inference time. Standard fusion layers either preserve robustness or calibration…
The electroacoustic resonator is an effcient electro-active device for noise attenuation in enclosed cavities or acoustic waveguides. It is made of a loudspeaker (the actuator) and one or more microphones (the sensors). So far, the desired…
There are some issues with traditional whispering gallery mode (WGM) resonators like poor light extraction and a dense mode spectrum. In this paper we introduce a solution to these limitations by proposing open WGM (OWGM) resonators that…
Automatic speaker verification (ASV) systems are vulnerable to spoofing attacks. We propose a spoofing-robust ASV system optimized directly for the recently introduced architecture-agnostic detection cost function (a-DCF), which allows…
Achieving single-mode operation and highly directional (preferably unidirectional) in-plane light output from whispering-gallery (WG) mode semiconductor microdisk resonators without seriously degrading the mode Q-factor challenges designers…
Speech enhancement is designed to enhance the intelligibility and quality of speech across diverse noise conditions. Recently, diffusion model has gained lots of attention in speech enhancement area, achieving competitive results. Current…
Diffusion-based Generative Models (DGMs) have achieved unparalleled performance in synthesizing high-quality visual content, opening up the opportunity to improve image super-resolution (SR) tasks. Recent solutions for these tasks often…