Related papers: The Maximal Overlap Discrete Wavelet Scattering Tr…
We present an empirical validation of the directional non-commutative monoidal embedding framework recently introduced in prior work~\cite{Godavarti2025monoidal}. This framework defines learnable compositional embeddings using distinct…
Multi-scale Vision Transformer (ViT) has emerged as a powerful backbone for computer vision tasks, while the self-attention computation in Transformer scales quadratically w.r.t. the input patch number. Thus, existing solutions commonly…
Recent advances in the chirplet transform and wavelet-chirplet transform (WCT) have enabled the estimation of instantaneous frequencies (IFs) and chirprates, as well as mode retrieval from multicomponent signals with crossover IF curves.…
Although spatio-temporal graph neural networks have achieved great empirical success in handling multiple correlated time series, they may be impractical in some real-world scenarios due to a lack of sufficient high-quality training data.…
Wave-based analog computing in the forms of inverse-designed metastructures and the meshes of Mach-Zehnder interferometers (MZI) have recently received considerable attention due to their capability in emulating linear operators, performing…
Time-frequency methods for vibration-based gearbox faults detection have been considered the most efficient method. Among these methods, continuous wavelet transform (CWT) as one of the best time-frequency method has been used for both…
Discrete diffusion models are a new class of text generators that offer advantages such as bidirectional context use, parallelizable generation, and flexible prompting compared to autoregressive models. However, a critical limitation of…
This paper explores the design of constant modulus Matched-Illumination (MI) waveforms using the Multi-Tone Sinusoidal Frequency Modulation (MTSFM) waveform model. MI waveforms are optimized for detecting targets in known noise and clutter…
The Wavelength-shifting Optical Module (WOM) is a novel photosensor concept for the instrumentation of large detector volumes with single-photon sensitivity. The key objective is to improve the signal-to-noise ratio which is achieved by…
As large language models (LLMs) continue to scale up, mixture-of-experts (MoE) has become a common technology in SOTA models. MoE models rely on expert parallelism (EP) to alleviate memory bottleneck, which introduces all-to-all…
Most existing ultra-high resolution (UHR) segmentation methods always struggle in the dilemma of balancing memory cost and local characterization accuracy, which are both taken into account in our proposed Guided Patch-Grouping Wavelet…
A fast algorithm for Antoine and Vandergheynst's (1998) directional continuous spherical wavelet transform (CSWT) is presented. Computational requirements are reduced by a factor of O(\sqrt{N}), when N is the number of pixels on the sphere.…
One of the common and promising deep learning approaches used for medical image segmentation is transformers, as they can capture long-range dependencies among the pixels by utilizing self-attention. Despite being successful in medical…
Continuous wavelet transform (CWT) based time-scale and multi-fractal analyses have been carried out on the anode glow related nonlinear floating potential fluctuations in a hollow cathode glow discharge plasma. CWT has been used to obtain…
An original multiplex scheme is introduced, which is based on Mallat's multiresolution formulation of wavelet systems. This system is adaptable and its implementation is well matched to digital signal processors and computers. The approach…
All metamaterial applications are based upon the idea that extreme material properties can be achieved through appropriate dynamic homogenization of composites. This homogenization is almost always done for infinite domains and the results…
Augmenting vision-language-action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predicting states and actions due to the modality gap. To address this, we propose DUal-STream…
Non-overlapping patch-wise convolution is the default image tokenizer for all state-of-the-art vision Transformer (ViT) models. Even though many ViT variants have been proposed to improve its efficiency and accuracy, little research on…
As attitude and motion sensing components, inertial sensors are widely used in various portable devices. But the severe errors of inertial sensors restrain their function, especially the trajectory recovery and semantic recognition. As a…
In this paper, we design a new class of high-efficiency deep joint source-channel coding methods to achieve end-to-end video transmission over wireless channels. The proposed methods exploit nonlinear transform and conditional coding…