中文
相关论文

相关论文: Subjective and Objective Quality Assessment of Ren…

200 篇论文

The Visual Question Answering (VQA) task combines challenges for processing data with both Visual and Linguistic processing, to answer basic `common sense' questions about given images. Given an image and a question in natural language, the…

计算机视觉与模式识别 · 计算机科学 2020-12-24 Yash Srivastava , Vaishnav Murali , Shiv Ram Dubey , Snehasis Mukherjee

Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feeding frame-level…

In recent years, AI generative models have made remarkable progress across various domains, including text generation, image generation, and video generation. However, assessing the quality of text-to-video generation is still in its…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Xinli Yue , Jianhui Sun , Han Kong , Liangchao Yao , Tianyi Wang , Lei Li , Fengyun Rao , Jing Lv , Fan Xia , Yuetang Deng , Qian Wang , Lingchen Zhao

Self-attention based Transformer has achieved great success in many computer vision tasks. However, its application to video quality assessment (VQA) has not been satisfactory so far. Evaluating the quality of in-the-wild videos is…

计算机视觉与模式识别 · 计算机科学 2023-06-22 Fengchuang Xing , Yuan-Gen Wang , Weixuan Tang , Guopu Zhu , Sam Kwong

Depth perception plays an essential role in the viewer experience for immersive virtual reality (VR) visual environments. However, previous research investigations in the depth quality of 3D/stereoscopic images are rather limited, and in…

计算机视觉与模式识别 · 计算机科学 2024-08-20 Wei Zhou , Zhou Wang

In learning vision-language representations from web-scale data, the contrastive language-image pre-training (CLIP) mechanism has demonstrated a remarkable performance in many vision tasks. However, its application to the widely studied…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Fengchuang Xing , Mingjie Li , Yuan-Gen Wang , Guopu Zhu , Xiaochun Cao

Visual question answering (VQA) is an important and challenging multimodal task in computer vision. Recently, a few efforts have been made to bring VQA task to aerial images, due to its potential real-world applications in disaster…

计算机视觉与模式识别 · 计算机科学 2023-01-24 Kun Li , George Vosselman , Michael Ying Yang

Audio descriptions (ADs) narrate important visual details in movies, enabling Blind and Low Vision (BLV) users to understand narratives and appreciate visual details. Existing works in automatic AD generation mostly focus on few-second…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Divy Kala , Eshika Khandelwal , Makarand Tapaswi

Building photorealistic, animatable full-body digital humans remains a longstanding challenge in computer graphics and vision. Recent advances in animatable avatar modeling have largely progressed along two directions: improving the…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Heming Zhu , Guoxing Sun , Marc Habermann

Digital humans are attracting more and more research interest during the last decade, the generation, representation, rendering, and animation of which have been put into large amounts of effort. However, the quality assessment of digital…

计算机视觉与模式识别 · 计算机科学 2023-03-01 Zicheng Zhang , Yingjie Zhou , Wei Sun , Xiongkuo Min , Yuzhe Wu , Guangtao Zhai

Video quality assessment (VQA) aims to simulate the human perception of video quality, which is influenced by factors ranging from low-level color and texture details to high-level semantic content. To effectively model these complicated…

计算机视觉与模式识别 · 计算机科学 2023-04-14 Kai Zhao , Kun Yuan , Ming Sun , Xing Wen

Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 question-answer pairs, with privately held ground-truth…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Paul Gavrikov , Wei Lin , M. Jehanzeb Mirza , Soumya Jahagirdar , Muhammad Huzaifa , Sivan Doveh , Serena Yeung-Levy , James Glass , Hilde Kuehne

The rapid growth of user-generated content (UGC) videos has produced an urgent need for effective video quality assessment (VQA) algorithms to monitor video quality and guide optimization and recommendation procedures. However, current VQA…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Huiyu Duan , Qiang Hu , Jiarui Wang , Liu Yang , Zitong Xu , Lu Liu , Xiongkuo Min , Chunlei Cai , Tianxiao Ye , Xiaoyun Zhang , Guangtao Zhai

Visual Question and Answering (VQA) problems are attracting increasing interest from multiple research disciplines. Solving VQA problems requires techniques from both computer vision for understanding the visual contents of a presented…

计算机视觉与模式识别 · 计算机科学 2016-04-07 Ilija Ilievski , Shuicheng Yan , Jiashi Feng

Video dimensions are continuously increasing to provide more realistic and immersive experiences to global streaming and social media viewers. However, increments in video parameters such as spatial resolution and frame rate are inevitably…

图像与视频处理 · 电气工程与系统科学 2022-01-19 Dae Yeol Lee , Somdyuti Paul , Christos G. Bampis , Hyunsuk Ko , Jongho Kim , Se Yoon Jeong , Blake Homan , Alan C. Bovik

The recent progress in Vision-Language Models (VLMs) has broadened the scope of multimodal applications. However, evaluations often remain limited to functional tasks, neglecting abstract dimensions such as personality traits and human…

计算与语言 · 计算机科学 2025-06-04 Jingxuan Li , Yuning Yang , Shengqi Yang , Linfan Zhang , Ying Nian Wu

Action Quality Assessment (AQA), which aims at automatic and fair evaluation of athletic performance, has gained increasing attention in recent years. However, athletes are often in rapid movement and the corresponding visual appearance…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Mengshi Qi , Hao Ye , Jiaxuan Peng , Huadong Ma

The rise in video streaming applications has increased the demand for video quality assessment (VQA). In 2016, Netflix introduced Video Multi-Method Assessment Fusion (VMAF), a full reference VQA metric that strongly correlates with…

多媒体 · 计算机科学 2024-01-25 Vignesh V Menon , Prajit T Rajendran , Reza Farahani , Klaus Schoeffmann , Christian Timmerer

We present Neural Head Avatars, a novel neural representation that explicitly models the surface geometry and appearance of an animatable human avatar that can be used for teleconferencing in AR/VR or other applications in the movie or…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Philip-William Grassal , Malte Prinzler , Titus Leistner , Carsten Rother , Matthias Nießner , Justus Thies

Virtual reality (VR) is making waves around the world recently. However, traditional video streaming is not suitable for VR video because of the huge size and view switch requirements of VR videos. Since the view of each user is limited, it…

网络与互联网体系结构 · 计算机科学 2019-06-28 Jie Feng , Yongpeng Wu , Guangtao Zhai , Ning Liu , Wenjun Zhang