Related papers: Context encoding enables machine learning-based qu…
X-ray free electron laser (XFEL) experiments have brought unique capabilities and opened new directions in research, such as creating new states of matter or directly measuring atomic motion. One such area is the ability to use finely…
A universal unanswered question in neuroscience and machine learning is whether computers can decode the patterns of the human brain. Multi-Voxels Pattern Analysis (MVPA) is a critical tool for addressing this question. However, there are…
Optoacoustic (OA) imaging is based on excitation of biological tissues with nanosecond-duration laser pulses followed by subsequent detection of ultrasound waves generated via light-absorption-mediated thermoelastic expansion. OA imaging…
Prompt-based methods, which encode medical priors through descriptive text, have been only minimally explored for CT Image Quality Assessment (IQA). While such prompts can embed prior knowledge about diagnostic quality, they often introduce…
Photoacoustic imaging (PAI) and ultrasound imaging (USI) are important biomedical imaging techniques, due to their unique and complementary advantages in tissue's structure and function visualization. In this Letter, we proposed a coaxial…
Photoacoustic tomography (PAT) is an emerging imaging modality that aims at measuring the high-contrast optical properties of tissues by means of high-resolution ultrasonic measurements. The interaction between these two types of waves is…
Transformer-based language models rely on positional encoding (PE) to handle token order and support context length extrapolation. However, existing PE methods lack theoretical clarity and rely on limited evaluation metrics to substantiate…
Human face exhibits an inherent hierarchy in its representations (i.e., holistic facial expressions can be encoded via a set of facial action units (AUs) and their intensity). Variational (deep) auto-encoders (VAE) have shown great results…
Photoacoustic imaging and sensing have been studied extensively to probe the optical absorption of biological tissue in multiple scales ranging from large organs to small molecules. However, its elastic oscillation characterization is…
Visual Question Answering (VQA) presents a unique challenge as it requires the ability to understand and encode the multi-modal inputs - in terms of image processing and natural language processing. The algorithm further needs to learn how…
Image classification is a core task of intelligent sensing, conventionally follows a sequential imaging then processing pipeline. However, redundant high-dimensional image reconstruction is inherently inefficient, especially in photon…
Machine learning methods for computational imaging require uncertainty estimation to be reliable in real settings. While Bayesian models offer a computationally tractable way of recovering uncertainty, they need large data volumes to be…
The task of word-level quality estimation (QE) consists of taking a source sentence and machine-generated translation, and predicting which words in the output are correct and which are wrong. In this paper, propose a method to effectively…
Background: A universal unanswered question in neuroscience and machine learning is whether computers can decode the patterns of the human brain. Multi-Voxels Pattern Analysis (MVPA) is a critical tool for addressing this question. However,…
The resolution of photoacoustic imaging deep inside scattering media is limited by the acoustic diffraction limit. In this work, taking inspiration from super-resolution imaging techniques developed to beat the optical diffraction limit, we…
Modern photon-counting sensors are increasingly dominated by Poisson noise, yet conventional feature-specific imaging (FSI), based on principal component analysis (PCA), is optimized for additive Gaussian noise and variance preservation…
This paper studies a simple extension of image-based Masked Autoencoders (MAE) to self-supervised representation learning from audio spectrograms. Following the Transformer encoder-decoder design in MAE, our Audio-MAE first encodes audio…
As a widely recognized approach to deep generative modeling, Variational Auto-Encoders (VAEs) still face challenges with the quality of generated images, often presenting noticeable blurriness. This issue stems from the unrealistic…
The previous advancements in pathology image understanding primarily involved developing models tailored to specific tasks. Recent studies has demonstrated that the large vision-language model can enhance the performance of various…
Multiwavelength photoacoustic images encode information about a tissue's optical absorption distribution. This can be used to estimate its blood oxygen saturation distribution (sO2), an important physiological indicator of tissue health and…