中文
相关论文

相关论文: Video in 10 Bits: Few-Bit VideoQA for Efficiency a…

200 篇论文

Talking head video compression has advanced with neural rendering and keypoint-based methods, but challenges remain, especially at low bit rates, including handling large head movements, suboptimal lip synchronization, and distorted facial…

图像与视频处理 · 电气工程与系统科学 2025-06-17 Riku Takahashi , Ryugo Morita , Jinjia Zhou

A main goal in developing video-compression algorithms is to enhance human-perceived visual quality while maintaining file size. But modern video-analysis efforts such as detection and recognition, which are integral to video surveillance…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Mikhail Dremin , Konstantin Kozhemyakov , Ivan Molodetskikh , Malakhov Kirill , Artur Sagitov , Dmitriy Vatolin

Despite recent progress on computer vision and natural language processing, developing a machine that can understand video story is still hard to achieve due to the intrinsic difficulty of video story. Moreover, researches on how to…

计算与语言 · 计算机科学 2020-12-18 Seongho Choi , Kyoung-Woon On , Yu-Jung Heo , Ahjeong Seo , Youwon Jang , Minsu Lee , Byoung-Tak Zhang

This paper tackles the intricate challenge of video question-answering (VideoQA). Despite notable progress, current methods fall short of effectively integrating questions with video frames and semantic object-level abstractions to create…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Sai Bhargav Rongali , Mohamad Hassan N C , Ankit Jha , Neha Bhargava , Saurabh Prasad , Biplab Banerjee

In recent years, model quantization for face recognition has gained prominence. Traditionally, compressing models involved vast datasets like the 5.8 million-image MS1M dataset as well as extensive training times, raising the question of…

计算机视觉与模式识别 · 计算机科学 2024-02-29 William Gazali , Jocelyn Michelle Kho , Joshua Santoso , Williem

Accurate and efficient discrete video tokenization is essential for long video sequences processing. Yet, the inherent complexity and variable information density of videos present a significant bottleneck for current tokenizers, which…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Haotian Ye , Qiyuan He , Jiaqi Han , Puheng Li , Jiaojiao Fan , Zekun Hao , Fitsum Reda , Yogesh Balaji , Huayu Chen , Sheng Liu , Angela Yao , James Zou , Stefano Ermon , Haoxiang Wang , Ming-Yu Liu

Videos often capture objects, their visible properties, their motion, and the interactions between different objects. Objects also have physical properties such as mass, which the imaging pipeline is unable to directly capture. However,…

计算机视觉与模式识别 · 计算机科学 2022-11-08 Maitreya Patel , Tejas Gokhale , Chitta Baral , Yezhou Yang

Blind video quality assessment (BVQA) is a highly challenging task due to the intrinsic complexity of video content and visual distortions, especially given the high popularity of social media videos, which originate from a wide range of…

图像与视频处理 · 电气工程与系统科学 2026-01-06 Wei Sun , Linhan Cao , Jun Jia , Zhichao Zhang , Zicheng Zhang , Xiongkuo Min , Guangtao Zhai

Mathematical reasoning in real-world video settings presents a fundamentally different challenge than in static images or text. It requires interpreting fine-grained visual information, accurately reading handwritten or digital text, and…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Hanoona Rasheed , Abdelrahman Shaker , Anqi Tang , Muhammad Maaz , Ming-Hsuan Yang , Salman Khan , Fahad Shahbaz Khan

We present a framework to analyze various aspects of models for video question answering (VideoQA) using customizable synthetic datasets, which are constructed automatically from gameplay videos. Our work is motivated by the fact that…

计算机视觉与模式识别 · 计算机科学 2017-08-15 Jonghwan Mun , Paul Hongsuck Seo , Ilchae Jung , Bohyung Han

We propose Compressed Video Aggregator (CVA), a lightweight micro-video recommendation module that decouples video information from preference learning. It aggregates frozen VFM embeddings, and uses latent reasoning without cross-attention…

机器学习 · 计算机科学 2026-05-12 Yang Xiao , Huiyuan Chen , Kaiyuan Deng , Chao Jiang , Zinan Ling , Ruimeng Ye , Xiaolong Ma , Bo Hui

Recent incremental learning for action recognition usually stores representative videos to mitigate catastrophic forgetting. However, only a few bulky videos can be stored due to the limited memory. To address this problem, we propose…

计算机视觉与模式识别 · 计算机科学 2022-11-03 Yixuan Pei , Zhiwu Qing , Jun Cen , Xiang Wang , Shiwei Zhang , Yaxiong Wang , Mingqian Tang , Nong Sang , Xueming Qian

Video summarization plays an important role in selecting keyframe for understanding a video. Traditionally, it aims to find the most representative and diverse contents (or frames) in a video for short summaries. Recently, query-conditioned…

计算机视觉与模式识别 · 计算机科学 2020-09-14 Neeraj Baghel , Suresh C. Raikwar , Charul Bhatnagar

Instructional videos provide detailed how-to guides for various tasks, with viewers often posing questions regarding the content. Addressing these questions is vital for comprehending the content, yet receiving immediate answers is…

计算机视觉与模式识别 · 计算机科学 2024-02-01 Saelyne Yang , Sunghyun Park , Yunseok Jang , Moontae Lee

Small object-centric spatial understanding in indoor videos remains a significant challenge for multimodal large language models (MLLMs), despite its practical value for object search and assistive applications. Although existing benchmarks…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Zhiyu Zhou , Peilin Liu , Ruoxuan Zhang , Luyang Zhang , Cheng Zhang , Hongxia Xie , Wen-Huang Cheng

Video question-answering (QA) is a core task in video understanding. Evaluating the quality of video QA and video caption data quality for training video large language models (VideoLLMs) is an essential challenge. Although various methods…

计算机视觉与模式识别 · 计算机科学 2025-02-07 Hao Liang , Zirong Chen , Hejun Dong , Wentao Zhang

Human perception is at the core of lossy video compression and yet, it is challenging to collect data that is sufficiently dense to drive compression. In perceptual quality assessment, human feedback is typically collected as a single…

图像与视频处理 · 电气工程与系统科学 2022-05-10 Evgenya Pergament , Pulkit Tandon , Kedar Tatwawadi , Oren Rippel , Lubomir Bourdev , Bruno Olshausen , Tsachy Weissman , Sachin Katti , Alexander G. Anderson

Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of shortcut solutions based on superficial visual or textual cues.…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Benno Krojer , Mojtaba Komeili , Candace Ross , Quentin Garrido , Koustuv Sinha , Nicolas Ballas , Mahmoud Assran

The proliferation of single-photon image sensors has opened the door to a plethora of high-speed and low-light imaging applications. However, data collected by these sensors are often 1-bit or few-bit, and corrupted by noise and strong…

图像与视频处理 · 电气工程与系统科学 2024-11-18 Prateek Chennuri , Yiheng Chi , Enze Jiang , G. M. Dilshan Godaliyadda , Abhiram Gnanasambandam , Hamid R. Sheikh , Istvan Gyongy , Stanley H. Chan

The canonical approach to video-and-language learning (e.g., video question answering) dictates a neural model to learn from offline-extracted dense video features from vision models and text features from language models. These feature…

计算机视觉与模式识别 · 计算机科学 2021-02-12 Jie Lei , Linjie Li , Luowei Zhou , Zhe Gan , Tamara L. Berg , Mohit Bansal , Jingjing Liu