中文
相关论文

相关论文: VTONQA: A Multi-Dimensional Quality Assessment Dat…

200 篇论文

In this paper, we propose a novel virtual try-on from unconstrained designs (ucVTON) task to enable photorealistic synthesis of personalized composite clothing on input human images. Unlike prior arts constrained by specific input types,…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Shuliang Ning , Duomin Wang , Yipeng Qin , Zirong Jin , Baoyuan Wang , Xiaoguang Han

In recent years, artificial intelligence (AI)-driven video generation has gained significant attention. Consequently, there is a growing need for accurate video quality assessment (VQA) metrics to evaluate the perceptual quality of…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Zhichao Zhang , Wei Sun , Xinyue Li , Jun Jia , Xiongkuo Min , Zicheng Zhang , Chunyi Li , Zijian Chen , Puyi Wang , Fengyu Sun , Shangling Jui , Guangtao Zhai

Visual reasoning over structured data such as tables is a critical capability for modern vision-language models (VLMs), yet current benchmarks remain limited in scale, diversity, or reasoning depth, especially when it comes to rendered…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Boammani Aser Lompo , Marc Haraoui

In recent years, people have increasingly used AI to help them with their problems by asking questions on different topics. One of these topics can be software-related and programming questions. In this work, we focus on the questions which…

计算机视觉与模式识别 · 计算机科学 2024-05-20 Motahhare Mirzaei , Mohammad Javad Pirhadi , Sauleh Eetemadi

Objective visual quality assessment of 3D models is a fundamental issue in computer graphics. Quality assessment metrics may allow a wide range of processes to be guided and evaluated, such as level of detail creation, compression,…

计算机视觉与模式识别 · 计算机科学 2021-02-09 Jinjiang Guo , Vincent Vidal , Irene Cheng , Anup Basu , Atilla Baskurt , Guillaume Lavoue

Integrating external tools into Large Foundation Models (LFMs) has emerged as a promising approach to enhance their problem-solving capabilities. While existing studies have demonstrated strong performance in tool-augmented Visual Question…

人工智能 · 计算机科学 2026-03-05 Shaofeng Yin , Ting Lei , Yang Liu

Studies of virtual try-on (VITON) have been shown their effectiveness in utilizing the generative neural network for virtually exploring fashion products, and some of recent researches of VITON attempted to synthesize human image wearing…

计算机视觉与模式识别 · 计算机科学 2022-05-11 Soonchan Park , Jinah Park

Generating natural, diverse, and meaningful questions from images is an essential task for multimodal assistants as it confirms whether they have understood the object and scene in the images properly. The research in visual question…

计算机视觉与模式识别 · 计算机科学 2020-12-08 Alkesh Patel , Akanksha Bindal , Hadas Kotek , Christopher Klein , Jason Williams

Document image quality assessment (DIQA) is an important component for various applications, including optical character recognition (OCR), document restoration, and the evaluation of document image processing systems. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Zhichao Ma , Fan Huang , Lu Zhao , Fengjun Guo , Guangtao Zhai , Xiongkuo Min

Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 question-answer pairs, with privately held ground-truth…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Paul Gavrikov , Wei Lin , M. Jehanzeb Mirza , Soumya Jahagirdar , Muhammad Huzaifa , Sivan Doveh , Serena Yeung-Levy , James Glass , Hilde Kuehne

Research on image quality assessment (IQA) remains limited mainly due to our incomplete knowledge about human visual perception. Existing IQA algorithms have been designed or trained with insufficient subjective data with a small degree of…

计算机视觉与模式识别 · 计算机科学 2020-09-29 Lucie Lévêque , Ji Yang , Xiaohan Yang , Pengfei Guo , Kenneth Dasalla , Leida Li , Yingying Wu , Hantao Liu

Visual Question Answering (VQA) is an evolving research field aimed at enabling machines to answer questions about visual content by integrating image and language processing techniques such as feature extraction, object detection, text…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Ngoc Dung Huynh , Mohamed Reda Bouadjenek , Sunil Aryal , Imran Razzak , Hakim Hacid

The Diffusion model has a strong ability to generate wild images. However, the model can just generate inaccurate images with the guidance of text, which makes it very challenging to directly apply the text-guided generative model for…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Shufang Zhang , Minxue Ni , Lei Wang , Wenxin Ding , Shuai Chen , Yuhong Liu

Perceptual quality assessment of the videos acquired in the wilds is of vital importance for quality assurance of video services. The inaccessibility of reference videos with pristine quality and the complexity of authentic distortions pose…

图像与视频处理 · 电气工程与系统科学 2022-04-06 Bowen Li , Weixia Zhang , Meng Tian , Guangtao Zhai , Xianpei Wang

Face Image Quality Assessment (FIQA) aims to predict the utility of a face image for face recognition (FR) systems. State-of-the-art FIQA methods mainly rely on convolutional neural networks (CNNs), leaving the potential of Vision…

计算机视觉与模式识别 · 计算机科学 2025-08-25 Andrea Atzori , Fadi Boutros , Naser Damer

Image-based Virtual Try-ON aims to transfer an in-shop garment onto a specific person. Existing methods employ a global warping module to model the anisotropic deformation for different garment parts, which fails to preserve the semantic…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Zhenyu Xie , Zaiyu Huang , Xin Dong , Fuwei Zhao , Haoye Dong , Xijin Zhang , Feida Zhu , Xiaodan Liang

Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models can not read! Our paper takes a first step towards addressing…

计算与语言 · 计算机科学 2019-05-15 Amanpreet Singh , Vivek Natarajan , Meet Shah , Yu Jiang , Xinlei Chen , Dhruv Batra , Devi Parikh , Marcus Rohrbach

Is it possible to develop an "AI Pathologist" to pass the board-certified examination of the American Board of Pathology? To achieve this goal, the first step is to create a visual question answering (VQA) dataset where the AI agent is…

计算与语言 · 计算机科学 2020-03-24 Xuehai He , Yichen Zhang , Luntian Mou , Eric Xing , Pengtao Xie

The conventional image-based virtual try-on method cannot generate fitting images that correspond to the clothing size because the system cannot accurately reflect the body information of a person. In this study, an image-based virtual…

计算机视觉与模式识别 · 计算机科学 2023-03-01 Minoru Kuribayashi , Koki Nakai , Nobuo Funabiki

Medical Visual Question Answering (MedVQA) presents a significant opportunity to enhance diagnostic accuracy and healthcare delivery by leveraging artificial intelligence to interpret and answer questions based on medical images. In this…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Xiaoman Zhang , Chaoyi Wu , Ziheng Zhao , Weixiong Lin , Ya Zhang , Yanfeng Wang , Weidi Xie
‹ 上一页 1 8 9 10 下一页 ›