English
Related papers

Related papers: Pythia v0.1: the Winning Entry to the VQA Challeng…

200 papers

Foundation models in digital pathology use massive datasets to learn useful compact feature representations of complex histology images. However, there is limited transparency into what drives the correlation between dataset size and…

Existing blind image quality assessment (BIQA) methods are mostly designed in a disposable way and cannot evolve with unseen distortions adaptively, which greatly limits the deployment and application of BIQA models in real-world scenarios.…

Multimedia · Computer Science 2021-04-30 Jianzhao Liu , Wei Zhou , Jiahua Xu , Xin Li , Shukun An , Zhibo Chen

Introduction: We describe the foundation of PETRIC, an image reconstruction challenge to minimise the computational runtime of related algorithms for Positron Emission Tomography (PET). Purpose: Although several similar challenges are…

We propose a novel framework that leverages Visual Question Answering (VQA) models to automate the evaluation of LLM-generated data visualizations. Traditional evaluation methods often rely on human judgment, which is costly and unscalable,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-30 James Ford , Xingmeng Zhao , Dan Schumacher , Anthony Rios

Large Language Models (LLMs) are trained on vast amounts of data, most of which is automatically scraped from the internet. This data includes encyclopedic documents that harbor a vast amount of general knowledge (e.g., Wikipedia) but also…

This paper addresses the problem of determining the best answer in Community-based Question Answering (CQA) websites by focussing on the content. In particular, we present a system, ACQUA [http://acqua.kmi.open.ac.uk], that can be installed…

Computation and Language · Computer Science 2015-06-18 George Gkotsis , Maria Liakata , Carlos Pedrinaci , John Domingue

Humans gather information by engaging in conversations involving a series of interconnected questions and answers. For machines to assist in information gathering, it is therefore essential to enable them to answer conversational questions.…

Computation and Language · Computer Science 2019-04-02 Siva Reddy , Danqi Chen , Christopher D. Manning

Several datasets have recently been constructed to expose brittleness in models trained on existing benchmarks. While model performance on these challenge datasets is significantly lower compared to the original benchmark, it is unclear…

Computation and Language · Computer Science 2019-04-30 Nelson F. Liu , Roy Schwartz , Noah A. Smith

When a model is trying to gather information in an interactive setting, it benefits from asking informative questions. However, in the case of a grounded multi-turn image identification task, previous studies have been constrained to polar…

Computation and Language · Computer Science 2023-11-16 Sedrick Keh , Justin T. Chiu , Daniel Fried

Inference-Time Scaling has been critical to the success of recent models such as OpenAI o1 and DeepSeek R1. However, many techniques used to train models for inference-time scaling require tasks to have answers that can be verified,…

Computation and Language · Computer Science 2025-06-02 Zhilin Wang , Jiaqi Zeng , Olivier Delalleau , Daniel Egert , Ellie Evans , Hoo-Chang Shin , Felipe Soares , Yi Dong , Oleksii Kuchaiev

The proliferation of generative AI tools has rendered traditional modular assessments in computing and data-centric education increasingly ineffective, creating a disconnect between academic evaluation and authentic skill measurement. This…

Computers and Society · Computer Science 2026-01-22 Kaihua Ding

In order to increase the effectiveness of model training, data reduction is essential to data-centric Artificial Intelligence (AI). It achieves this by locating the most instructive examples in massive datasets. To increase data quality and…

Machine Learning · Computer Science 2025-08-11 Fei Chen , Wenchi Zhou

Retrieval-augmented generation (RAG) systems are widely used in question-answering (QA) tasks, but current benchmarks lack metadata integration, limiting their evaluation in scenarios requiring both textual data and external information. To…

Information Retrieval · Computer Science 2026-02-13 Davide Bruni , Marco Avvenuti , Nicola Tonellotto , Maurizio Tesconi

Vision Language Models (VLMs) are increasingly adopted as central reasoning modules for embodied agents. Existing benchmarks evaluate their capabilities under ideal, well-lit conditions, yet robust 24/7 operation demands performance under a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Yohan Park , Hyunwoo Ha , Wonjun Jo , Tae-Hyun Oh

Existing literature on Question Answering (QA) mostly focuses on algorithmic novelty, data augmentation, or increasingly large pre-trained language models like XLNet and RoBERTa. Additionally, a lot of systems on the QA leaderboards do not…

Computation and Language · Computer Science 2019-09-13 Lin Pan , Rishav Chakravarti , Anthony Ferritto , Michael Glass , Alfio Gliozzo , Salim Roukos , Radu Florian , Avirup Sil

Recent advances in Large Language Models (LLMs) have driven the adoption of copilots in complex technical scenarios, underscoring the growing need for specialized information retrieval solutions. In this paper, we introduce FLAIR, a…

Information Retrieval · Computer Science 2025-08-20 William Zhang , Yiwen Zhu , Yunlei Lu , Mathieu Demarne , Wenjing Wang , Kai Deng , Nutan Sahoo , Katherine Lin , Miso Cilimdzic , Subru Krishnan

An accurate computational model for image quality assessment (IQA) benefits many vision applications, such as image filtering, image processing, and image generation. Although the study of face images is an important subfield in computer…

Computer Vision and Pattern Recognition · Computer Science 2023-08-01 Shaolin Su , Hanhe Lin , Vlad Hosu , Oliver Wiedemann , Jinqiu Sun , Yu Zhu , Hantao Liu , Yanning Zhang , Dietmar Saupe

This paper presents a strong baseline for real-world visual reasoning (GQA), which achieves 60.93% in GQA 2019 challenge and won the sixth place. GQA is a large dataset with 22M questions involving spatial understanding and multi-step…

Computer Vision and Pattern Recognition · Computer Science 2019-06-18 Chenfei Wu , Yanzhao Zhou , Gen Li , Nan Duan , Duyu Tang , Xiaojie Wang

Due to the diversity of assessment requirements in various application scenarios for the IQA task, existing IQA methods struggle to directly adapt to these varied requirements after training. Thus, when facing new requirements, a typical…

Computer Vision and Pattern Recognition · Computer Science 2024-03-11 Zewen Chen , Haina Qin , Juan Wang , Chunfeng Yuan , Bing Li , Weiming Hu , Liang Wang

Artificial Intelligence (AI) is redefining the frontiers of scientific domains, ranging from drug discovery to meteorological modeling, yet its integration within industrial manufacturing remains nascent and fraught with operational…

Computational Engineering, Finance, and Science · Computer Science 2025-11-18 Yingjie Shi , Yiru Gong , Yiqun Su , Suya Xiong , Jiale Han , Runtian Miao