English
Related papers

Related papers: Self-Supervised Learning of Perceptually Optimized…

200 papers

Cross-modal video retrieval aims to retrieve the semantically relevant videos given a text as a query, and is one of the fundamental tasks in Multimedia. Most of top-performing methods primarily leverage Visual Transformer (ViT) to extract…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Ning Han , Xun Yang , Ee-Peng Lim , Hao Chen , Qianru Sun

The ubiquity of vision transformers (ViTs) for various edge applications, including personalized learning, has created the demand for on-device fine-tuning. However, training with the limited memory and computation power of edge devices…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Sreetama Sarkar , Souvik Kundu , Kai Zheng , Peter A. Beerel

There has been a growing trend in compressing and transmitting videos from terminals for machine vision tasks. Nevertheless, most video coding optimization method focus on minimizing distortion according to human perceptual metrics,…

Multimedia · Computer Science 2025-12-18 Fei Zhao , Mengxi Guo , Shijie Zhao , Junlin Li , Li Zhang , Xiaodong Xie

Recent years have witnessed the significant development of learning-based video compression methods, which aim at optimizing objective or perceptual quality and bit rates. In this paper, we introduce deep video compression with perceptual…

Image and Video Processing · Electrical Eng. & Systems 2021-10-11 Saiping Zhang , Marta Mrak , Luis Herranz , Marc Górriz , Shuai Wan , Fuzheng Yang

Classical video coding for satisfying humans as the final user is a widely investigated field of studies for visual content, and common video codecs are all optimized for the human visual system (HVS). But are the assumptions and…

Image and Video Processing · Electrical Eng. & Systems 2022-03-14 Kristian Fischer , Christian Herglotz , André Kaup

The upcoming video coding standard, Versatile Video Coding (VVC), has shown great improvement compared to its predecessor, High Efficiency Video Coding (HEVC), in terms of bitrate saving. Despite its substantial performance, compressed…

Image and Video Processing · Electrical Eng. & Systems 2021-12-09 Fatemeh Nasiri , Wassim Hamidouche , Luce Morin , Nicolas Dhollande , Gildas Cocherel

Recent advances in end-to-end video compression have shown promising results owing to their unified end-to-end learning optimization. However, such generalized frameworks often lack content-specific adaptation, leading to suboptimal…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Tiange Zhang , Xiandong Meng , Siwei Ma

Cross-component linear model (CCLM) prediction has been repeatedly proven to be effective in reducing the inter-channel redundancies in video compression. Essentially speaking, the linear model is identically trained by employing accessible…

Multimedia · Computer Science 2021-09-01 Junru Li , Meng Wang , Li Zhang , Shiqi Wang , Kai Zhang , Shanshe Wang , Siwei Ma , Wen Gao

With the rapid development of Vision-Language Models (VLMs) and the growing demand for their applications, efficient compression of the image inputs has become increasingly important. Existing VLMs predominantly digest and understand…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Zifu Zhang , Tongda Xu , Siqi Li , Shengxi Li , Yue Zhang , Mai Xu , Yan Wang

One key challenge to learning-based video compression is that motion predictive coding, a very effective tool for video compression, can hardly be trained into a neural network. In this paper we propose the concept of PixelMotionCNN (PMCNN)…

Multimedia · Computer Science 2019-01-15 Zhibo Chen , Tianyu He , Xin Jin , Feng Wu

Efficient compression of 360-degree video content requires the application of advanced motion models for interframe prediction. The Motion Plane Adaptive (MPA) motion model projects the frames on multiple perspective planes in the 3D space.…

Image and Video Processing · Electrical Eng. & Systems 2025-04-01 Marina Ritthaler , Andy Regensky , André Kaup

The interactions between different tools added successively to a block-based video codec are critical to its rate-distortion efficiency. In particular, when deep neural network-based intra prediction modes are inserted into a block-based…

Image and Video Processing · Electrical Eng. & Systems 2021-08-19 Thierry Dumas , Franck Galpin , Philippe Bordes

In recent years, the field of learned video compression has witnessed rapid advancement, exemplified by the latest neural video codecs DCVC-DC that has outperformed the upcoming next-generation codec ECM in terms of compression ratio.…

Image and Video Processing · Electrical Eng. & Systems 2024-07-24 Zidian Qiu , Zongyao He , Zhi Jin

In this paper, we study a simplified affine motion model based coding framework to overcome the limitation of translational motion model and maintain low computational complexity. The proposed framework mainly has three key contributions.…

Multimedia · Computer Science 2017-02-22 Li Li , Houqiang Li , Dong Liu , Haitao Yang , Sixin Lin , Huanbang Chen , Feng Wu

We present a novel simple yet effective algorithm for motion-based video frame interpolation. Existing motion-based interpolation methods typically rely on a pre-trained optical flow model or a U-Net based pyramid network for motion…

Computer Vision and Pattern Recognition · Computer Science 2022-11-08 Xin Jin , Longhai Wu , Guotao Shen , Youxin Chen , Jie Chen , Jayoon Koo , Cheul-hee Hahm

Vision Transformers (ViT) become widely-adopted architectures for various vision tasks. Masked auto-encoding for feature pretraining and multi-scale hybrid convolution-transformer architectures can further unleash the potentials of ViT,…

Computer Vision and Pattern Recognition · Computer Science 2022-05-20 Peng Gao , Teli Ma , Hongsheng Li , Ziyi Lin , Jifeng Dai , Yu Qiao

Given recent advances in learned video prediction, we investigate whether a simple video codec using a pre-trained deep model for next frame prediction based on previously encoded/decoded frames without sending any motion side information…

Image and Video Processing · Electrical Eng. & Systems 2020-07-21 Serkan Sulun , A. Murat Tekalp

This paper proposes a novel modulation and coding scheme (MCS) selection framework that integrates mutual information (MI) prediction based on vector similarity search (VSS) for massive multi-user multiple-input multiple-output orthogonal…

Signal Processing · Electrical Eng. & Systems 2026-03-30 Fuga Kobayashi , Takumi Takahashi , Shinsuke Ibi , Takanobu Doi , Kazushi Muraoka , Hideki Ochiai

Recent work has demonstrated that using a carefully designed sensing matrix rather than a random one, can improve the performance of compressed sensing. In particular, a well-designed sensing matrix can reduce the coherence between the…

Information Theory · Computer Science 2010-09-09 Kevin Rosenblum , Lihi Zelnik-Manor , Yonina C. Eldar

Among the new techniques of Versatile Video Coding (VVC), the quadtree with nested multi-type tree (QT+MTT) block structure yields significant coding gains by providing more flexible block partitioning patterns. However, the recursive…

Image and Video Processing · Electrical Eng. & Systems 2025-07-16 Xinmin Feng , Zhuoyuan Li , Li Li , Dong Liu , Feng Wu