English
Related papers

Related papers: Daala: Building A Next-Generation Video Codec From…

200 papers

The exponential surge in video traffic has intensified the imperative for Video Quality Assessment (VQA). Leveraging cutting-edge architectures, current VQA models have achieved human-comparable accuracy. However, recent studies have…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Ao-Xiang Zhang , Yuan-Gen Wang , Yu Ran , Weixuan Tang , Qingxiao Guan , Chunsheng Yang

In this paper, we present a novel low-light image enhancement method called dark region-aware low-light image enhancement (DALE), where dark regions are accurately recognized by the proposed visual attention module and their brightness are…

Image and Video Processing · Electrical Eng. & Systems 2020-08-31 Dokyeong Kwon , Guisik Kim , Junseok Kwon

Vision-Language-Action (VLA) models offer a compelling framework for tackling complex robotic manipulation tasks, but they are often expensive to train. In this paper, we propose a novel VLA approach that leverages the competitive…

Robotics · Computer Science 2025-12-23 Max Argus , Jelena Bratulic , Houman Masnavi , Maxim Velikanov , Nick Heppert , Abhinav Valada , Thomas Brox

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it…

Vision-language-action (VLA) models remain constrained by the scarcity of action-labeled robot data, whereas action-free videos provide abundant evidence of how the physical world changes. Latent action models offer a promising way to…

Existing mezzanine image codecs lack specialized screen content coding tools and therefore struggle to maintain high image quality under bandwidth constraints, especially in areas with dense text. Although distribution codecs offer advanced…

Hardware Architecture · Computer Science 2026-04-01 Chenlong He , Leilei Huang , Wei Li , Hanyang Cui , Zhijian Hao , Xiaoyang Zeng , Yibo Fan

This document describes the Data Access Layer Interface (DALI). DALI defines the base web service interface common to all Data Access Layer (DAL) services. This standard defines the behaviour of common resources, the meaning and use of…

Instrumentation and Methods for Astrophysics · Physics 2019-05-22 Patrick Dowler , Markus Demleitner , Mark Taylor , Doug Tody

Substantial research has been done in saliency modeling to develop intelligent machines that can perceive and interpret their surroundings. But existing models treat videos as merely image sequences excluding any audio information, unable…

Image and Video Processing · Electrical Eng. & Systems 2023-02-27 Maryam Qamar Butt , Anis Ur Rahman

This paper develops a new video compression approach based on underdetermined blind source separation. Underdetermined blind source separation, which can be used to efficiently enhance the video compression ratio, is combined with various…

Multimedia · Computer Science 2012-05-22 Jing Liu , Fei Qiao , Qi Wei , Huazhong Yang

Vision-Language-Action (VLA) models have demonstrated strong multi-modal reasoning capabilities, enabling direct action generation from visual perception and language instructions in an end-to-end manner. However, their substantial…

Robotics · Computer Science 2025-10-22 Siyu Xu , Yunke Wang , Chenghao Xia , Dihao Zhu , Tao Huang , Chang Xu

End-to-end learning-based video compression has made steady progress over the last several years. However, unlike learning-based image coding, which has already surpassed its handcrafted counterparts, learning-based video coding still has…

Image and Video Processing · Electrical Eng. & Systems 2023-04-20 Hadi Hadizadeh , Ivan V. Bajić

With the rapid growth of User-Generated Content (UGC) exchanged between users and sharing platforms, the need for video quality assessment in the wild is increasingly evident. UGC is typically acquired using consumer devices and undergoes…

Image and Video Processing · Electrical Eng. & Systems 2025-03-14 Xinyi Wang , Angeliki Katsenou , David Bull

Deep generative models, and particularly facial animation schemes, can be used in video conferencing applications to efficiently compress a video through a sparse set of keypoints, without the need to transmit dense motion vectors. While…

Multimedia · Computer Science 2022-07-28 Goluck Konuko , Stéphane Lathuilière , Giuseppe Valenzise

The rapid advancement of artificial intelligence (AI) technology has led to the prioritization of standardizing the processing, coding, and transmission of video using neural networks. To address this priority area, the Moving Picture,…

Multimedia · Computer Science 2023-09-15 Chuanmin Jia , Feng Ye , Fanke Dong , Kai Lin , Leonardo Chiariglione , Siwei Ma , Huifang Sun , Wen Gao

Latent Action Models (LAMs) enable the learning of world models from unlabeled video by inferring abstract actions between consecutive frames. However, LAMs face a fundamental trade-off between action abstraction and generation fidelity.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Tianqiu Zhang , Muyang Lyu , Yufan Zhang , Fang Fang , Si Wu

The growth in video Internet traffic and advancements in video attributes such as framerate, resolution, and bit-depth boost the demand to devise a large-scale, highly efficient video encoding environment. This is even more essential for…

The Bj{\o}ntegaard Delta (BD) method proposed in 2001 has become a popular tool for comparing video codec compression efficiency. It was initially proposed to compute bitrate and quality differences between two Rate-Distortion curves using…

Multimedia · Computer Science 2024-01-09 Nabajeet Barman , Maria G. Martini , Yuriy Reznik

Video generation is rapidly evolving towards unified audio-video generation. In this paper, we present ALIVE, a generation model that adapts a pretrained Text-to-Video (T2V) model to Sora-style audio-video generation and animation. In…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Ying Guo , Qijun Gan , Yifu Zhang , Jinlai Liu , Yifei Hu , Pan Xie , Dongjun Qian , Yu Zhang , Ruiqi Li , Yuqi Zhang , Ruibiao Lu , Xiaofeng Mei , Bo Han , Xiang Yin , Bingyue Peng , Zehuan Yuan

RaceVLA presents an innovative approach for autonomous racing drone navigation by leveraging Visual-Language-Action (VLA) to emulate human-like behavior. This research explores the integration of advanced algorithms that enable drones to…

Vision-language models (VLMs) are vulnerable to adversarial image perturbations. Existing works based on adversarial training against task-specific adversarial examples are computationally expensive and often fail to generalize to unseen…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Jingning Xu , Haochen Luo , Chen Liu