Related papers: Context encoding enables machine learning-based qu…
In this article we use the Natural Scenes Dataset (NSD) to train a family of feature-weighted receptive field neural encoding models. These models use a pre-trained vision or text backbone and map extracted features to the voxel space via…
Computer-aided pathology detection algorithms for video-based imaging modalities must accurately interpret complex spatiotemporal information by integrating findings across multiple frames. Current state-of-the-art methods operate by…
In many clinical conditions, such as head trauma, stroke, and low cardiac output states, the brain is at risk for hypoxic-ischemic injury. The metabolic rate and oxygen consumption of the brain are reflected in internal jugular venous…
We propose Parabolic Position Encoding (PaPE), a parabola-based position encoding for vision modalities in attention-based architectures. Given a set of vision tokens-such as from videos, event camera streams, images, or point clouds-our…
Drag-based image editing enables intuitive visual manipulation through point-based drag operations. Existing methods mainly rely on diffusion inversion or pixel-space warping with inpainting. However, inversion inherently introduces…
Visual question answering (VQA) refers to the problem where, given an image and a natural language question about the image, a correct natural language answer has to be generated. A VQA model has to demonstrate both the visual understanding…
Visual Question Answering (VQA) models aim to answer natural language questions about given images. Due to its ability to ask questions that differ from those used when training the model, medical VQA has received substantial attention in…
Due to its specificity, fluorescence microscopy (FM) has become a quintessential imaging tool in cell biology. However, photobleaching, phototoxicity, and related artifacts continue to limit FM's utility. Recently, it has been shown that…
Visual Question Answering (VQA) models employ attention mechanisms to discover image locations that are most relevant for answering a specific question. For this purpose, several multimodal fusion strategies have been proposed, ranging from…
Visual question answering (VQA) usesimage processing algorithms to process the image and natural language processing methods to understand and answer the question. VQA is helpful to a visually impaired person, can be used for the security…
Recent works in medical image segmentation have actively explored various deep learning architectures or objective functions to encode high-level features from volumetric data owing to limited image annotations. However, most existing…
Photoacoustic imaging (PAI) is rapidly moving from the laboratory to the clinic, increasing the need to understand confounders which might adversely affect patient care. Over the past five years, landmark studies have shown the clinical…
Quantum sensing exploits quantum phenomena to enhance the detection and estimation of classical parameters of physical systems and biological entities, particularly so as to overcome the inefficiencies of its classical counterparts. A…
Visual Question Answering (VQA) is a complex semantic task requiring both natural language processing and visual recognition. In this paper, we explore whether VQA is solvable when images are captured in a sub-Nyquist compressive paradigm.…
Classical machine learning often struggles with complex, high-dimensional data. Quantum machine learning offers a potential solution, promising more efficient processing. The quantum convolutional neural network (QCNN), a hybrid algorithm,…
Two-photon absorption (TPA) is a nonlinear optical process with wide-ranging applications from spectroscopy to super-resolution imaging. Despite this, the precise measurement and characterisation of TPA parameters are challenging due to…
The ability to perform high-precision optical measurements is paramount to science and engineering. Laser interferometry enables interaction-free sensing with a precision ultimately limited by shot noise. Quantum optical sensors can surpass…
Interpretable machine learning is rapidly becoming a crucial tool for scientific discovery. Among existing approaches, variational autoencoders (VAEs) have shown promise in extracting the hidden physical features of some input data, with no…
The current models of image representation based on Convolutional Neural Networks (CNN) have shown tremendous performance in image retrieval. Such models are inspired by the information flow along the visual pathway in the human visual…
Biomedical photoacoustic tomography, which can provide high resolution 3D soft tissue images based on the optical absorption, has advanced to the stage at which translation from the laboratory to clinical settings is becoming possible. The…