Related papers: Valley-Peak Modulation in Phase Space: an Exposure…
Voice activity detection (VAD) is a challenging task in low signal-to-noise ratio (SNR) environment, especially in non-stationary noise. To deal with this issue, we propose a novel attention module that can be integrated in Long Short-Term…
Visual Prompt Tuning (VPT) is a parameter-efficient fune-tuning technique that adapts a pre-trained vision Transformer (ViT) by learning a small set of parameters in the input space, known as prompts. In VPT, we uncover ``burstiness'' in…
Traditional pulsar polarization sweep analysis starts from the point dipole rotating vector model (RVM) approximation. If augmented by a measurement of the sweep phase shift, one obtains an estimate of the emission altitude (Blaskiewicz,…
Readout error models for noisy quantum devices almost universally assume that measurement noise is classical: the measurement statistics are obtained from the ideal computational-basis populations by a column-stochastic assignment matrix…
Weakly-supervised audio-visual video parsing (WS-AVVP) aims to localize the temporal extents of audio, visual and audio-visual event instances as well as identify the corresponding event categories with only video-level category labels for…
We calculate the void probability function (VPF) in simulations of Lyman-$\alpha$ emitters (LAEs) across a wide redshift range ($z=3.1,\ 4.5,\ 5.7,\ 6.6$). The VPF measures the zero-point correlation function (i.e. places devoid of…
Channel estimation at the receiver side is essential to adaptive modulation schemes, prohibiting low complexity systems from using variable rate and/or variable power transmissions. Towards providing a solution to this problem, we introduce…
Recent advances in point cloud object detection have increasingly adopted Transformer-based and State Space Models (SSMs) to capture long-range dependencies. However, these serialized frameworks strictly maintain the consistency of input…
Medical image segmentation is critical for diagnosing and treating spinal disorders. However, the presence of high noise, ambiguity, and uncertainty makes this task highly challenging. Factors such as unclear anatomical boundaries,…
Background noise considerably reduces the accuracy and reliability of speaker verification (SV) systems. These challenges can be addressed using a speech enhancement system as a front-end module. Recently, diffusion probabilistic models…
While autoregressive Large Vision-Language Models (LVLMs) demonstrate remarkable proficiency in multimodal tasks, they face a "Visual Signal Dilution" phenomenon, where the accumulation of textual history expands the attention partition…
Structural identification and damage detection can be generalized as the simultaneous estimation of input forces, physical parameters, and dynamical states. Although Kalman-type filters are efficient tools to address this problem, the…
Free-space modulation of light is crucial for many applications, from light detection and ranging to virtual or augmented reality. Traditional means of modulating free-space light involves spatial light modulators based on liquid crystals…
Generalized Vector Approximate Message Passing (GVAMP) is an efficient iterative algorithm for approximately minimum-mean-squared-error estimation of a random vector $\mathbf{x}\sim p_{\mathbf{x}}(\mathbf{x})$ from generalized linear…
Orbital angular momentum (OAM), a topological degree of freedom of light, is theoretically invariant under continuous deformations; yet, its physical observability degrades precipitously in complex media, creating a fundamental…
Statistical ensemble formalism of Kim, Mandel and Wolf (J. Opt. Soc. Am. A 4, 433 (1987)) offers a realistic model for characterizing the effect of stochastic non-image forming optical media on the state of polarization of transmittedlight.…
Audio-visual speech recognition (AVSR) provides a promising solution to ameliorate the noise-robustness of audio-only speech recognition with visual information. However, most existing efforts still focus on audio modality to improve…
Neural vocoder using denoising diffusion probabilistic model (DDPM) has been improved by adaptation of the diffusion noise distribution to given acoustic features. In this study, we propose SpecGrad that adapts the diffusion noise so that…
Phaseless diffraction measurements recorded by a CCD detector are often affected by Poisson noise. In this paper, we propose a dictionary learning model by employing patches based sparsity to denoise Poisson phaseless measurement. The model…
In dynamically varying optical wireless communication (OWC) links, conventional quadrature amplitude modulation (QAM) in optical orthogonal frequency-division multiplexing (OFDM) requires frequent channel estimation and equalization,…