English
Related papers

Related papers: Predicting Chroma from Luma in AV1

200 papers

This paper introduces FlowMAC, a novel neural audio codec for high-quality general audio compression at low bit rates based on conditional flow matching (CFM). FlowMAC jointly learns a mel spectrogram encoder, quantizer and decoder. At…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-08 Nicola Pia , Martin Strauss , Markus Multrus , Bernd Edler

[Abridged] The Coma cluster luminosity function (LF) from ultraviolet (2000 AA) to the near-infrared (H band) is summarized. In the UV the LF is very steep, much steeper than in the optical. The steep Coma UV LF implies that faint and…

Astrophysics · Physics 2007-05-23 S. Andreon , J. -C. Cuillandre , R. Pello

Prior to encoding color images for RGB full-color, Bayer color filter array (CFA), and digital time delay integration (DTDI) CFA images, performing chroma subsampling on their converted chroma images is necessary and important. In this…

Image and Video Processing · Electrical Eng. & Systems 2020-09-24 Kuo-Liang Chung , Szu-Ni Chen , Yu-Ling Lee , Chao-Liang Yu

Temporal feature extraction is an important issue in video-based action recognition. Optical flow is a popular method to extract temporal feature, which produces excellent performance thanks to its capacity of capturing pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2020-05-07 Yuecong Xu , Jianfei Yang , Kezhi Mao , Jianxiong Yin , Simon See

Pre-training has been proven to be effective in boosting the performance of Isolated Sign Language Recognition (ISLR). Existing pre-training methods solely focus on the compact pose data, which eliminates background perturbation but…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Kepeng Wu , Zecheng Li , Hezhen Hu , Wengang Zhou , Houqiang Li

This work examines whether decoder-only Transformers such as LLaMA, which were originally designed for large language models (LLMs), can be adapted to the computer vision field. We first "LLaMAfy" a standard ViT step-by-step to align with…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Jiahao Wang , Wenqi Shao , Mengzhao Chen , Chengyue Wu , Yong Liu , Taiqiang Wu , Kaipeng Zhang , Songyang Zhang , Kai Chen , Ping Luo

We introduce here a predictive coding based model that aims to generate accurate and sharp future frames. Inspired by the predictive coding hypothesis and related works, the total model is updated through a combination of bottom-up and…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Chaofan Ling , Weihua Li , Junpei Zhong

Despite the remarkable performance of vision language models (VLMs) such as Contrastive Language Image Pre-training (CLIP), the large size of these models is a considerable obstacle to their use in federated learning (FL) systems where the…

Machine Learning · Computer Science 2025-03-11 Yihang Wu , Ahmad Chaddad , Christian Desrosiers , Tareef Daqqaq , Reem Kateb

We present the first clustering results of X-ray selected AGN at z~3. Using Chandra X-ray imaging and UVR optical colors from MUSYC photometry in the ECDF-S field, we selected a sample of 58 z~3 AGN candidates. From the optical data we also…

CFHTLS optical photometry has been used to study the galaxy luminosity functions of 14 X-ray selected clusters from the XMM-LSS survey. These are mostly groups and poor clusters, with masses (M_{500}) in the range 0.6 to 19x10 ^{13} M_solar…

Cosmology and Nongalactic Astrophysics · Physics 2015-05-14 Abdulmonem Alshino , Habib Khosroshahi , Trevor Ponman , Jon Willis , Marguerite Pierre , Florian Pacaud , Graham P. Smith

We constructed the composite Luminosity Function (LF) of cluster galaxies in the g,r and i bands from the photometry of a mixed (Abell and X-ray selected) sample of the cores of 65 clusters, ranging in redshift from 0.05 to 0.25. The…

Astrophysics · Physics 2010-11-04 Bianca Garilli , Dario Maccagni , Stefano Andreon

Continual learning with vision-language models like CLIP offers a pathway toward scalable machine learning systems by leveraging its transferable representations. Existing CLIP-based methods adapt the pre-trained image encoder by adding…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Mao-Lin Luo , Zi-Hao Zhou , Tong Wei , Min-Ling Zhang

In recent years, large-scale vision-language models (VLMs) have demonstrated remarkable performance on multimodal understanding and reasoning tasks. However, handling high-dimensional visual features often incurs substantial computational…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Xiaoyang Guo , Keze Wang

HEVC includes a Coding Unit (CU) level luminance-based perceptual quantization technique known as AdaptiveQP. AdaptiveQP perceptually adjusts the Quantization Parameter (QP) at the CU level based on the spatial activity of raw input video…

Multimedia · Computer Science 2018-02-13 Lee Prangnell , Miguel Hernández-Cabronero , Victor Sanchez

Most sRGB-based LLIE methods suffer from entangled luminance and color, while the HSV color space offers insufficient decoupling at the cost of introducing significant red and black noise artifacts. Recently, the HVI color space has been…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Zhixin Cheng , Fangwen Zhang , Xiaotian Yin , Baoqun Yin , Haodian Wang

The color - stellar mass-to-light ratio relation (CMLR) is a widely accepted tool to estimate the stellar mass (M*) of a galaxy. However, an individual CMLR tends to give distinct M* for a same galaxy when it is applied in different bands.…

Astrophysics of Galaxies · Physics 2020-07-22 Wei Du , Stacy S. McGaugh

Semantic location prediction from multimodal social media posts is a critical task with applications in personalized services and human mobility analysis. This paper introduces \textit{Contextualized Vision-Language Alignment (CoVLA)}, a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Liu Jing , Amirul Rahman

Reliable estimation of illuminant chromaticity is crucial for simulating color constancy and for white balancing digital images. However, estimating illuminant chromaticity from a single image is an ill-posed task, in general, and existing…

Computer Vision and Pattern Recognition · Computer Science 2019-06-14 Eytan Lifshitz , Dani Lischinski

Visual-Language Alignment (VLA) has gained a lot of attention since CLIP's groundbreaking work. Although CLIP performs well, the typical direct latent feature alignment lacks clarity in its representation and similarity scores. On the other…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Yifan Li , Yikai Wang , Yanwei Fu , Dongyu Ru , Zheng Zhang , Tong He

Graph contrastive learning (GCL) is the most representative and prevalent self-supervised learning approach for graph-structured data. Despite its remarkable success, existing GCL methods highly rely on an augmentation scheme to learn the…

Machine Learning · Computer Science 2022-06-07 Haonan Wang , Jieyu Zhang , Qi Zhu , Wei Huang