English
Related papers

Related papers: UGC-VQA: Benchmarking Blind Video Quality Assessme…

200 papers

The security risks of AI-driven video editing have garnered significant attention. Although recent studies indicate that adding perturbations to images can protect them from malicious edits, directly applying image-based methods to perturb…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 KaiZhou Li , Jindong Gu , Xinchun Yu , Junjie Cao , Yansong Tang , Xiao-Ping Zhang

User generated content (UGC) refers to videos that are uploaded by users and shared over the Internet. UGC may have low quality due to noise and previous compression. When re-encoding UGC for streaming or downloading, a traditional video…

Image and Video Processing · Electrical Eng. & Systems 2023-03-14 Xin Xiong , Eduardo Pavez , Antonio Ortega , Balu Adsumilli

Recent advancements in Large Video-Language Models (LVLMs) have led to promising results in multimodal video understanding. However, it remains unclear whether these models possess the cognitive capabilities required for high-level tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Chenglin Li , Qianglong Chen , Zhi Li , Feng Tao , Yin Zhang

The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot showcase the model's capabilities, fixed evaluation operators…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Yuhang Yang , Ke Fan , Shangkun Sun , Hongxiang Li , Ailing Zeng , FeiLin Han , Wei Zhai , Wei Liu , Yang Cao , Zheng-Jun Zha

Advances in multimodal large language models enable automatic video narration and question answering (VQA), offering scalable alternatives to labor-intensive, human-authored audio descriptions (ADs) for blind and low vision (BLV) viewers.…

Human-Computer Interaction · Computer Science 2026-03-17 Maryam Cheema , Sina Elahimanesh , Pooyan Fazli , Hasti Seifi

Understanding surveillance video content remains a critical yet underexplored challenge in vision-language research, particularly due to its real-world complexity, irregular event dynamics, and safety-critical implications. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Bo Liu , Pengfei Qiao , Minhan Ma , Xuange Zhang , Yinan Tang , Peng Xu , Kun Liu , Tongtong Yuan

The applications of short-term user-generated video (UGV), such as Snapchat, and Youtube short-term videos, booms recently, raising lots of multimodal machine learning tasks. Among them, learning the correspondence between audio and visual…

Artificial Intelligence · Computer Science 2020-10-20 Runze Su , Fei Tao , Xudong Liu , Haoran Wei , Xiaorong Mei , Zhiyao Duan , Lei Yuan , Ji Liu , Yuying Xie

Visual Question Answering (VQA) is a multi-modal task that involves answering questions from an input image, semantically understanding the contents of the image and answering it in natural language. Using VQA for disaster management is an…

Computer Vision and Pattern Recognition · Computer Science 2022-11-14 Aditya Kane , V Manushree , Sahil Khose

This paper presents a state-of-the-art model for visual question answering (VQA), which won the first place in the 2017 VQA Challenge. VQA is a task of significant importance for research in artificial intelligence, given its multimodal…

Computer Vision and Pattern Recognition · Computer Science 2017-08-10 Damien Teney , Peter Anderson , Xiaodong He , Anton van den Hengel

Visual Question Answering (VQA) presents a unique challenge by requiring models to understand and reason about visual content to answer questions accurately. Existing VQA models often struggle with biases introduced by the training data,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Zhifei Li , Feng Qiu , Yiran Wang , Yujing Xia , Kui Xiao , Miao Zhang , Yan Zhang

In recent years, visual question answering (VQA) has become topical. The premise of VQA's significance as a benchmark in AI, is that both the image and textual question need to be well understood and mutually grounded in order to infer the…

Computer Vision and Pattern Recognition · Computer Science 2018-03-20 Feng Liu , Tao Xiang , Timothy M. Hospedales , Wankou Yang , Changyin Sun

Deep learning-based video quality assessment (deep VQA) has demonstrated significant potential in surpassing conventional metrics, with promising improvements in terms of correlation with human perception. However, the practical deployment…

Image and Video Processing · Electrical Eng. & Systems 2025-06-03 Chen Feng , Duolikun Danier , Haoran Wang , Fan Zhang , Benoit Vallade , Alex Mackin , David Bull

Visual Question-Answering (VQA) has become key to user experience, particularly after improved generalization capabilities of Vision-Language Models (VLMs). But evaluating VLMs for an application requirement using a standardized framework…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Neelabh Sinha , Vinija Jain , Aman Chadha

Existing approaches to complaint analysis largely rely on unimodal, short-form content such as tweets or product reviews. This work advances the field by leveraging multimodal, multi-turn customer support dialogues, where users often share…

Computation and Language · Computer Science 2025-11-19 Rishu Kumar Singh , Navneet Shreya , Sarmistha Das , Apoorva Singh , Sriparna Saha

Blind image quality assessment (BIQA) aims at automatically and accurately forecasting objective scores for visual signals, which has been widely used to monitor product and service quality in low-light applications, covering smartphone…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Miaohui Wang , Zhuowei Xu , Mai Xu , Weisi Lin

Omnidirectional videos (ODVs) play an increasingly important role in the application fields of medical, education, advertising, tourism, etc. Assessing the quality of ODVs is significant for service-providers to improve the user's Quality…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Xilei Zhu , Huiyu Duan , Yuqin Cao , Yuxin Zhu , Yucheng Zhu , Jing Liu , Li Chen , Xiongkuo Min , Guangtao Zhai

The growing capabilities of AI in generating video content have brought forward significant challenges in effectively evaluating these videos. Unlike static images or text, video content involves complex spatial and temporal dynamics which…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Xiao Liu , Xinhao Xiang , Zizhong Li , Yongheng Wang , Zhuoheng Li , Zhuosheng Liu , Weidi Zhang , Weiqi Ye , Jiawei Zhang

Recent advances in image generation models have expanded their applications beyond aesthetic imagery toward practical visual content creation. However, existing benchmarks mainly focus on natural image synthesis and fail to systematically…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yan Li , Zezi Zeng , Ziwei Zhou , Xin Gao , Muzhao Tian , Yifan Yang , Mingxi Cheng , Qi Dai , Yuqing Yang , Lili Qiu , Zhendong Wang , Zhengyuan Yang , Xue Yang , Lijuan Wang , Ji Li , Chong Luo

In recent years, user generated content (UGC) has become the dominant force in internet traffic. However, UGC videos exhibit a higher degree of variability and diverse characteristics compared to traditional encoding test videos. This…

Multimedia · Computer Science 2025-12-19 Fei Zhao , Mengxi Guo , Shijie Zhao , Junlin Li , Li Zhang , Xiaodong Xie

Human-centric Video Anomaly Detection (VAD) aims to identify human behaviors that deviate from normal. At its core, human-centric VAD faces substantial challenges, such as the complexity of diverse human behaviors, the rarity of anomalies,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Armin Danesh Pazho , Shanle Yao , Ghazal Alinezhad Noghre , Babak Rahimi Ardabili , Vinit Katariya , Hamed Tabkhi
‹ Prev 1 8 9 10 Next ›