English
Related papers

Related papers: DeepSurvey-Bench: Evaluating Academic Value of Aut…

200 papers

While deep learning models have greatly improved the performance of most artificial intelligence tasks, they are often criticized to be untrustworthy due to the black-box problem. Consequently, many works have been proposed to study the…

Computation and Language · Computer Science 2021-09-08 Lijie Wang , Hao Liu , Shuyuan Peng , Hongxuan Tang , Xinyan Xiao , Ying Chen , Hua Wu , Haifeng Wang

Machine learning on graphs has made substantial progress across domains such as molecular property prediction and chip design. Yet benchmarking practices remain fragmented, often relying on narrow, task-specific datasets and inconsistent…

The rapid proliferation of AI-generated content, driven by advances in generative adversarial networks, diffusion models, and multimodal large language models, has made the creation and dissemination of synthetic media effortless,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Guangyu Lin , Li Lin , Christina P. Walker , Daniel S. Schiff , Shu Hu

Recent advances in Text-to-3D (T23D) generative models have enabled the synthesis of diverse, high-fidelity 3D assets from textual prompts. However, existing challenges restrict the development of reliable T23D quality assessment (T23DQA).…

Computer Vision and Pattern Recognition · Computer Science 2025-11-06 Bingyang Cui , Yujie Zhang , Qi Yang , Zhu Li , Yiling Xu

Deep research agents have attracted growing attention for their potential to orchestrate multi-stage research workflows, spanning literature synthesis, methodological design, and empirical verification. Despite these strides, evaluating…

Artificial Intelligence · Computer Science 2025-11-11 Haiyuan Wan , Chen Yang , Junchi Yu , Meiqi Tu , Jiaxuan Lu , Di Yu , Jianbao Cao , Ben Gao , Jiaqing Xie , Aoran Wang , Wenlong Zhang , Philip Torr , Dongzhan Zhou

Recently, there has been a growing interest among large language model (LLM) developers in LLM-based document reading systems, which enable users to upload their own documents and pose questions related to the document contents, going…

Computation and Language · Computer Science 2024-07-16 Anni Zou , Wenhao Yu , Hongming Zhang , Kaixin Ma , Deng Cai , Zhuosheng Zhang , Hai Zhao , Dong Yu

As the mathematical capabilities of large language models (LLMs) improve, it becomes increasingly important to evaluate their performance on research-level tasks at the frontier of mathematical knowledge. However, existing benchmarks are…

We introduce OVERTONBENCH, a novel framework for measuring Overton pluralism in LLMs--the extent to which diverse viewpoints are represented in model outputs. We (i) formalize Overton pluralism as a set coverage metric (OVERTONSCORE), (ii)…

Artificial Intelligence · Computer Science 2026-03-03 Elinor Poole-Dayan , Jiayi Wu , Taylor Sorensen , Jiaxin Pei , Michiel A. Bakker

We present FormalProofBench, a private benchmark designed to evaluate whether AI models can produce formally verified mathematical proofs at the graduate level. Each task pairs a natural-language problem with a Lean~4 formal statement, and…

Artificial Intelligence · Computer Science 2026-03-31 Nikil Ravi , Kexing Ying , Vasilii Nesterov , Rayan Krishnan , Elif Uskuplu , Bingyu Xia , Janitha Aswedige , Langston Nashold

Autonomous agents are increasingly expected to support scientific research, and recent benchmarks report progress in code repair and autonomous experimentation. However, these evaluations typically assume a pre-configured execution…

Software Engineering · Computer Science 2026-03-12 Yubang Wang , Chenxi Zhang , Bowen Chen , Zezheng Huai , Zihao Dai , Xinchi Chen , Yuxin Wang , Yining Zheng , Jingjing Gong , Xipeng Qiu

Explainable AI (XAI) has gained significant attention for providing insights into the decision-making processes of deep learning models, particularly for image classification tasks through visual explanations visualized by saliency maps.…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Yifei Zhang , James Song , Siyi Gu , Tianxu Jiang , Bo Pan , Guangji Bai , Liang Zhao

Currently, there is no consistent benchmarking across multi-disciplines. Even no previous work tries to relate different categories of benchmarks in multi-disciplines. This article investigates the origin and evolution of the benchmark…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-04-12 Jianfeng Zhan

Survey research is a fundamental empirical method in software engineering, enabling the systematic collection of data on professional practices, perceptions, and experiences. However, recent advances in large language models (LLMs) have…

Software Engineering · Computer Science 2025-12-22 Ronnie de Souza Santos , Italo Santos , Maria Teresa Baldassarre , Cleyton Magalhaes , Mairieli Wessel

Existing metrics for assessing question generation not only require costly human reference but also fail to take into account the input context of generation, rendering the lack of deep understanding of the relevance between the generated…

Computation and Language · Computer Science 2022-05-02 Xiaoqiang Wang , Bang Liu , Siliang Tang , Lingfei Wu

Data-driven science is an emerging paradigm where scientific discoveries depend on the execution of computational AI models against rich, discipline-specific datasets. With modern machine learning frameworks, anyone can develop and execute…

Machine Learning · Computer Science 2022-08-09 Seth Ockerman , John Wu , Christopher Stewart

Much of the focus in the design of deep neural networks has been on improving accuracy, leading to more powerful yet highly complex network architectures that are difficult to deploy in practical scenarios, particularly on edge devices such…

Computer Vision and Pattern Recognition · Computer Science 2018-08-28 Alexander Wong

Scientific fact-checking aims to determine the veracity of scientific claims by retrieving and analysing evidence from research literature. The problem is inherently more complex than general fact-checking since it must accommodate the…

Information Retrieval · Computer Science 2025-08-18 Xingyu Deng , Xi Wang , Mark Stevenson

Cultural AI benchmarks often rely on implicit assumptions about measured constructs, leading to vague formulations with poor validity and unclear interrelations. We propose exposing these assumptions using explicit cognitive models…

Artificial Intelligence · Computer Science 2024-09-26 Jonathan H. Rystrøm , Kenneth C. Enevoldsen

Aligning AI with human values is a pressing unsolved problem. To address the lack of quantitative metrics for value alignment, we propose EigenBench: a black-box method for comparatively benchmarking language models' values. Given an…

Artificial Intelligence · Computer Science 2026-03-03 Jonathn Chang , Leonhard Piff , Suvadip Sana , Jasmine X. Li , Lionel Levine

Leveraging Multi-modal Large Language Models (MLLMs) to accelerate frontier scientific research is promising, yet how to rigorously evaluate such systems remains unclear. Existing benchmarks mainly focus on single-document understanding,…

Artificial Intelligence · Computer Science 2026-04-14 Lei Xiong , Huaying Yuan , Zheng Liu , Zhao Cao , Zhicheng Dou