English
Related papers

Related papers: A multi-level approach with visual information for…

200 papers

Long-video understanding with multimodal language models suffers from three compounding bottlenecks: heavy decode cost to obtain dense RGB frames, quadratic token growth with frame count, and weak motion perception under sparse keyframe…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Haopeng Jin , Hongzhu Yi , Wenlong Zhao , Jinwen Luo , Shani Ye , Zhenyu Guan , Shiquan Dong , Tiankun Yang , Tao Yu

Long-video understanding has emerged as a crucial capability in real-world applications such as video surveillance, meeting summarization, educational lecture analysis, and sports broadcasting. However, it remains computationally…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Benjamin Schneider , Dongfu Jiang , Chao Du , Tianyu Pang , Wenhu Chen

The Large Vision-Language Model (LVLM) has enhanced the performance of various downstream tasks in visual-language understanding. Most existing approaches encode images and videos into separate feature spaces, which are then fed as inputs…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Bin Lin , Yang Ye , Bin Zhu , Jiaxi Cui , Munan Ning , Peng Jin , Li Yuan

Understanding long-form egocentric videos remains challenging for multimodal large language models (MLLMs) due to limited context length and insufficient grounding of fine-grained visual details. The recently proposed HD-EPIC benchmark…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Yinsong Xu , Wei Jing , Liuxin Zhang , Wanjun Lv , Hui Li

As the latest video coding standard, versatile video coding (VVC) has shown its ability in retaining pixel quality. To excavate more compression potential for video conference scenarios under ultra-low bitrate, this paper proposes a bitrate…

Image and Video Processing · Electrical Eng. & Systems 2023-03-21 Anni Tang , Yan Huang , Jun Ling , Zhiyu Zhang , Yiwei Zhang , Rong Xie , Li Song

Feature coding has been recently considered to facilitate intelligent video analysis for urban computing. Instead of raw videos, extracted features in the front-end are encoded and transmitted to the back-end for further processing. In this…

Multimedia · Computer Science 2020-09-11 Weiyao Lin , Xiaoyi He , Wenrui Dai , John See , Tushar Shinde , Hongkai Xiong , Lingyu Duan

Despite the widespread adoption of Shannon's confusion-diffusion architecture in image encryption, the implementation of diffusion to sequentially establish inter-pixel dependencies for attaining plaintext sensitivity constrains algorithmic…

Cryptography and Security · Computer Science 2026-04-01 Dong Jiang , Hui-ran Luo , Zi-jian Cui , Xi-jue Zhao , Lin-sheng Huang , Liang-liang Lu

Due to the strong correlation between adjacent pixels, most image encryption schemes perform multiple rounds of confusion and diffusion to protect the image against attacks. Such operations, however, are time-consuming, cannot meet the…

Cryptography and Security · Computer Science 2024-08-28 Dong Jiang , Zhen Yuan , Wen-xin Li , Liang-liang Lu

With recent advancements in video backbone architectures, combined with the remarkable achievements of large language models (LLMs), the analysis of long-form videos spanning tens of minutes has become both feasible and increasingly…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Yuxiao Chen , Jue Wang , Zhikang Zhang , Jingru Yi , Xu Zhang , Yang Zou , Zhaowei Cai , Jianbo Yuan , Xinyu Li , Hao Yang , Davide Modolo

Video-text retrieval is an important yet challenging task in vision-language understanding, which aims to learn a joint embedding space where related video and text instances are close to each other. Most current works simply measure the…

Computer Vision and Pattern Recognition · Computer Science 2021-08-02 Peng Wu , Xiangteng He , Mingqian Tang , Yiliang Lv , Jing Liu

Screen content coding (SCC) is becoming increasingly important in various applications, such as desktop sharing, video conferencing, and remote education. When compared to natural camera- captured content, screen content has different…

Graphics · Computer Science 2015-11-12 Haoming Chen , Ankur Saxena , Felix Fernandes

Today, visual data is often analyzed by a neural network without any human being involved, which demands for specialized codecs. For standard-compliant codec adaptations towards certain information sinks, HEVC or VVC provide the possibility…

Image and Video Processing · Electrical Eng. & Systems 2023-01-23 Kristian Fischer , Fabian Brand , Christian Herglotz , André Kaup

With the advancement of multi-modal Large Language Models (LLMs), Video LLMs have been further developed to perform on holistic and specialized video understanding. However, existing works are limited to specialized video understanding…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Hewen Pan , Cong Wei , Dashuang Liang , Zepeng Huang , Pengfei Gao , Ziqi Zhou , Lulu Xue , Pengfei Yan , Xiaoming Wei , Minghui Li , Shengshan Hu

Video data, especially long-form video, is extremely dense and high-dimensional. Text-based summaries of video content offer a way to represent query-relevant content in a much more compact manner than raw video. In addition, textual…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Kuleen Sasse , Efsun Sarioglu Kayi , Arun Reddy

The video technology scenery has been very vivid over the past years, with novel video coding technologies introduced that promise improved compression performance over state-of-the-art technologies. Despite the fact that a lot of video…

Image and Video Processing · Electrical Eng. & Systems 2022-04-13 Angeliki V. Katsenou , Fan Zhang , Mariana Afonso , Goce Dimitrov , David R. Bull

In recent years, screen content (SC) video including computer generated text, graphics and animations, have drawn more attention than ever, as many related applications become very popular. To address the need for efficient coding of such…

Multimedia · Computer Science 2020-12-01 Xiaozhong Xu , Shan Liu

In previous research, it was shown that the software decoding energy demand of High Efficiency Video Coding (HEVC) can be reduced by 15$\%$ by using a decoding-energy-rate-distortion optimization algorithm. To achieve this, the energy…

Image and Video Processing · Electrical Eng. & Systems 2022-09-22 Matthias Kränzler , Christian Herglotz , André Kaup

Video compression is a basic requirement for consumer and professional video applications alike. Video coding standards such as H.264/AVC and H.265/HEVC are widely deployed in the market to enable efficient use of bandwidth and storage for…

Image and Video Processing · Electrical Eng. & Systems 2021-04-28 Zhao Wang , Changyue Ma , Yan Ye

This technical report summarizes our method for the Video-And-Language Understanding Evaluation (VALUE) challenge (https://value-benchmark.github.io/challenge\_2021.html). We propose a CLIP-Enhanced method to incorporate the image-text…

Computer Vision and Pattern Recognition · Computer Science 2021-10-15 Guohao Li , Feng He , Zhifan Feng

Encryption on the internet with the shift to HTTPS has been an important step to improve the privacy of internet users. However, there is an increasing body of work about extracting information from encrypted internet traffic without having…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Arwin Gansekoele , Tycho Bot , Rob van der Mei , Sandjai Bhulai , Mark Hoogendoorn