English
Related papers

Related papers: Predicting Chroma from Luma in AV1

200 papers

Recent years have witnessed a significant increase in the performance of Vision and Language tasks. Foundational Vision-Language Models (VLMs), such as CLIP, have been leveraged in multiple settings and demonstrated remarkable performance…

Computer Vision and Pattern Recognition · Computer Science 2024-03-04 Santiago Castro , Amir Ziai , Avneesh Saluja , Zhuoning Yuan , Rada Mihalcea

Visible and infrared image fusion (VIF) aims to combine information from visible and infrared images into a single fused image. Previous VIF methods usually employ a color space transformation to keep the hue and saturation from the…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Hesong Li , Ying Fu

Vision-language models (VLMs) such as CLIP demonstrate strong generalization in zero-shot classification but remain highly vulnerable to adversarial perturbations. Existing methods primarily focus on adversarial fine-tuning or prompt…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Xingyu Zhu , Beier Zhu , Shuo Wang , Kesen Zhao , Hanwang Zhang

Purpose: Handheld gamma cameras with coded aperture collimators are under investigation for intraoperative imaging in nuclear medicine. Coded apertures are a promising collimation technique for applications such as lymph node localization…

Medical Physics · Physics 2024-01-22 Tobias Meißner , Laura Antonia Cerbone , Paolo Russo , Werner Nahm , Jürgen Hesser

With the flourishing of social media platforms, vision-language pre-training (VLP) recently has received great attention and many remarkable progresses have been achieved. The success of VLP largely benefits from the information…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Zhiyuan Ma , Jianjun Li , Guohui Li , Kaiyan Huang

Given recent advances in learned video prediction, we investigate whether a simple video codec using a pre-trained deep model for next frame prediction based on previously encoded/decoded frames without sending any motion side information…

Image and Video Processing · Electrical Eng. & Systems 2020-07-21 Serkan Sulun , A. Murat Tekalp

We present "Cross-Camera Convolutional Color Constancy" (C5), a learning-based method, trained on images from multiple cameras, that accurately estimates a scene's illuminant color from raw images captured by a new camera previously unseen…

Computer Vision and Pattern Recognition · Computer Science 2022-02-14 Mahmoud Afifi , Jonathan T. Barron , Chloe LeGendre , Yun-Ta Tsai , Francois Bleibel

Motion compensation is a fundamental technology in video coding to remove the temporal redundancy between video frames. To further improve the coding efficiency, sub-pel motion compensation has been utilized, which requires interpolation of…

Multimedia · Computer Science 2018-03-30 Ning Yan , Dong Liu , Houqiang Li , Feng Wu

Depth estimation from a single image of a conventional camera is a challenging task since depth cues are lost during the acquisition process. State-of-the-art approaches improve the discrimination between different depths by introducing a…

Computer Vision and Pattern Recognition · Computer Science 2023-09-18 Jhon Lopez , Edwin Vargas , Henry Arguello

Low-Light Image Enhancement (LLIE) task aims at improving contrast while restoring details and textures for images captured in low-light conditions. HVI color space has made significant progress in this task by enabling precise decoupling…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Xin Xu , Hao Liu , Wei Liu , Wei Wang , Jiayi Wu , Kui Jiang

Color filter array is spatial multiplexing of pixel-sized filters placed over pixel detectors in camera sensors. The state-of-the-art lossless coding techniques of raw sensor data captured by such sensors leverage spatial or cross-color…

Image and Video Processing · Electrical Eng. & Systems 2020-09-22 Yeejin Lee , Keigo Hirakawa

The UV (2000 A) luminosity function (hereafter UV LF) of Coma cluster galaxies, based on more than 120 members, is computed as the statistical difference between counts in the Coma direction and in the field. Our UV LF is an up-date of a…

Astrophysics · Physics 2007-05-23 S. Andreon

Cutting out an object and estimating its opacity mask, known as image matting, is a key task in many image editing applications. Deep learning approaches have made significant progress by adapting the encoder-decoder architecture of…

Computer Vision and Pattern Recognition · Computer Science 2020-03-18 Marco Forte , François Pitié

Computational Colour Constancy (CCC) consists of estimating the colour of one or more illuminants in a scene and using them to remove unwanted chromatic distortions. Much research has focused on illuminant estimation for CCC on single…

Computer Vision and Pattern Recognition · Computer Science 2023-03-03 Matteo Rizzo , Cristina Conati , Daesik Jang , Hui Hu

For video frame interpolation (VFI), existing deep-learning-based approaches strongly rely on the ground-truth (GT) intermediate frames, which sometimes ignore the non-unique nature of motion judging from the given adjacent frames. As a…

Computer Vision and Pattern Recognition · Computer Science 2022-03-22 Kun Zhou , Wenbo Li , Xiaoguang Han , Jiangbo Lu

Vision-language models (VLMs) unify computer vision and natural language processing in a single architecture capable of interpreting and describing images. Most state-of-the-art systems rely on two computationally intensive components:…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Andrew Kiruluta , Priscilla Burity

Recent advances in training-free video editing have enabled lightweight and precise cross-frame generation by leveraging pre-trained text-to-image diffusion models. However, existing methods often rely on heuristic frame selection to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Zhangkai Wu , Xuhui Fan , Zhongyuan Xie , Kaize Shi , Longbing Cao

We present a simple yet effective technique to estimate lighting in a single input image. Current techniques rely heavily on HDR panorama datasets to train neural networks to regress an input with limited field-of-view to a full environment…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Pakkapon Phongthawee , Worameth Chinchuthakun , Nontaphat Sinsunthithet , Amit Raj , Varun Jampani , Pramook Khungurn , Supasorn Suwajanakorn

Most Video-Large Language Models (Video-LLMs) adopt an encoder-decoder framework, where a vision encoder extracts frame-wise features for processing by a language model. However, this approach incurs high computational costs, introduces…

Computer Vision and Pattern Recognition · Computer Science 2025-11-06 Handong Li , Yiyuan Zhang , Longteng Guo , Xiangyu Yue , Jing Liu

Medical vision-language models (VLMs) are strong zero-shot recognizers for medical imaging, but their reliability under domain shift hinges on calibrated uncertainty with guarantees. Split conformal prediction (SCP) offers finite-sample…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Behzad Bozorgtabar , Dwarikanath Mahapatra , Sudipta Roy , Muzammal Naseer , Imran Razzak , Zongyuan Ge