Related papers: Predicting Chroma from Luma in AV1
Large Vision-Language Models (LVLMs) are an extension of Large Language Models (LLMs) that facilitate processing both image and text inputs, expanding AI capabilities. However, LVLMs struggle with object hallucinations due to their reliance…
Correlation plenoptic imaging (CPI) is a scanning-free diffraction-limited 3D optical imaging technique exploiting the peculiar properties of correlated light sources. CPI has been further extended to samples of interest to microscopy, such…
Hallucination remains a major challenge in multimodal large language models (MLLMs). To address this, various contrastive decoding (CD) methods have been proposed that contrasts original logits with hallucinated logits generated from…
We present new predictions for the galaxy three-point correlation function (3PCF) using high-resolution dissipationless cosmological simulations of a flat LCDM Universe which resolve galaxy-size halos and subhalos. We create realistic mock…
Currently, the most dominant approach to establishing language-image alignment is to pre-train text and image encoders jointly through contrastive learning, such as CLIP and its variants. In this work, we question whether such a costly…
The luminosity function of galaxies is derived from a cosmological hydrodynamic simulation of a Lambda cold dark matter (CDM) universe with the aid of a stellar population synthesis model. At z=0, the resulting B band luminosity function…
This paper presents a comprehensive analysis of motion vectors extracted from AV1-encoded video streams and their application in accelerating optical flow estimation. We demonstrate that motion vectors from AV1 video codec can serve as a…
We consider the capabilities of ALMA and the ngVLA to detect and image the[CII] 158\,$\mu$m line from galaxies into the cosmic `dark ages' ($z \sim 10$ to 20). The [CII] line may prove to be a powerful tool in determining spectroscopic…
In this paper, we present CLCC, a novel contrastive learning framework for color constancy. Contrastive learning has been applied for learning high-quality visual representations for image classification. One key aspect to yield useful…
Weakly supervised audio-visual video parsing (AVVP) methods aim to detect audible-only, visible-only, and audible-visible events using only video-level labels. Existing approaches tackle this by leveraging unimodal and cross-modal contexts.…
Recent advancements in image understanding have benefited from the extensive use of web image-text pairs. However, video understanding remains a challenge despite the availability of substantial web video-text data. This difficulty…
Vision-language models (VLMs) enable text-guided object detection but degrade severely under cross-view scenarios where ground and aerial viewpoints differ in altitude, scale, and spatial layout. These geometric changes introduce systematic…
Confocal microscopy is the backbone of cellular research labs across the world but unfortunately, the imaging is restricted to a single plane. Chromatic confocal microscopy offers the possibility to image multiple planes simultaneously thus…
Recently, learned video compression has drawn lots of attention and show a rapid development trend with promising results. However, the previous works still suffer from some criticial issues and have a performance gap with traditional…
Vision-language models (VLMs), such as CLIP and ALIGN, are generally trained on datasets consisting of image-caption pairs obtained from the web. However, real-world multimodal datasets, such as healthcare data, are significantly more…
We use 80922 galaxies in the Galaxy And Mass Assembly (GAMA) survey to measure the galaxy luminosity function (LF) in different environments over the redshift range 0.04<z<0.26. The depth and size of GAMA allows us to define samples split…
Daala is a new royalty-free video codec based on perceptually-driven coding techniques. We explore using its keyframe format for still picture coding and show how it has improved over the past year. We believe the technology used in Daala…
Cross-Domain Few-Shot Learning (CDFSL) aims to adapt large-scale pretrained models to specialized target domains with limited samples, yet the few-shot fine-tuning of vision-language models like CLIP remains underexplored. By establishing…
In contrast to traditional compression techniques performing linear transforms, the latent space of popular compressive autoencoders is obtained from a learned nonlinear mapping and hard to interpret. In this paper, we explore a promising…
Large Language Model-based Vision-Language Models (LLM-based VLMs) have demonstrated impressive results in various vision-language understanding tasks. However, how well these VLMs can see image detail beyond the semantic level remains…