中文
相关论文

相关论文: Paper Reconstruction Evaluation: Evaluating Presen…

200 篇论文

We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research. Agents must replicate 20 ICML 2024 Spotlight and Oral papers from scratch, including understanding paper contributions,…

Efficient reproduction of research papers is pivotal to accelerating scientific progress. However, the increasing complexity of proposed methods often renders reproduction a labor-intensive endeavor, necessitating profound domain expertise.…

人工智能 · 计算机科学 2026-04-27 Xuanle Zhao , Zilin Sang , Yuxuan Li , Qi Shi , Weilun Zhao , Shuo Wang , Duzhen Zhang , Xu Han , Zhiyuan Liu , Maosong Sun

Synthesizing unstructured research materials into manuscripts is an essential yet under-explored challenge in AI-driven scientific discovery. Existing autonomous writers are rigidly coupled to specific experimental pipelines, and produce…

人工智能 · 计算机科学 2026-04-08 Yiwen Song , Yale Song , Tomas Pfister , Jinsung Yoon

This paper introduces TrueGradeAI, an AI-driven digital examination framework designed to overcome the shortcomings of traditional paper-based assessments, including excessive paper usage, logistical complexity, grading delays, and…

人工智能 · 计算机科学 2025-09-29 Rakesh Thakur , Shivaansh Kaushik , Gauri Chopra , Harsh Rohilla

Frontier AI agents show increasing promise as scientific research assistants, and may eventually be useful for extended, open-ended research workflows. However, in order to use agents for novel research, we must first assess the underlying…

Computational reproducibility is essential for the credibility of scientific findings, particularly in the social sciences, where findings often inform real-world decisions. Manual reproducibility assessment is costly and time-consuming, as…

计算机与社会 · 计算机科学 2026-03-03 Linhao Zhang , Tong Xia , Jinghua Piao , Lizhen Cui , Yong Li

Reproducing machine learning papers is essential for scientific progress but remains challenging for both humans and automated agents. Existing agent-based methods often struggle to fully and accurately reproduce implementation details such…

软件工程 · 计算机科学 2025-08-26 Mingyang Zhou , Quanming Yao , Lun Du , Lanning Wei , Da Zheng

Assessing the reproducibility of social science papers is essential for promoting rigor in research processes, but manual assessment is costly. With recent advances in agentic AI systems (i.e., AI agents), we seek to evaluate their…

计算与语言 · 计算机科学 2025-07-28 Chuxuan Hu , Liyun Zhang , Yeji Lim , Aum Wadhwani , Austin Peters , Daniel Kang

Despite the rapid development of AI reviewers, evaluating such systems remains challenging: metrics favor overlap with human reviews over correctness. However, since human reviews often cover only a subset of salient issues and sometimes…

计算与语言 · 计算机科学 2026-05-19 Hexuan Deng , Xiaopeng Ke , Yichen Li , Ruina Hu , Dehao Huang , Derek F. Wong , Yue Wang , Xuebo Liu , Min Zhang

Autonomous AI systems can now generate complete economics research papers, but they substantially underperform human-authored publications in head-to-head comparisons. This paper decomposes the quality gap into two independent components:…

综合经济学 · 经济学 2026-04-07 Ning Li

Large language models (LLMs) are increasingly used in academic writing workflows, yet they frequently hallucinate by generating citations to sources that do not exist. This study analyzes 100 AI-generated hallucinated citations that…

数字图书馆 · 计算机科学 2026-02-06 Samar Ansari

Knowledge syntheses (literature reviews) are essential to health professions education (HPE), consolidating findings to advance theory and practice. However, they are labor-intensive, especially during data extraction. Artificial…

AI agents powered by large language models exhibit strong reasoning and problem-solving capabilities, enabling them to assist scientific research tasks such as formula derivation and code generation. However, whether these agents can…

AI research pipelines can now generate academic work that may satisfy existing peer review standards for quality, novelty, and methodological rigor. However, the publication system was built around the assumption that research is produced…

人工智能 · 计算机科学 2026-05-13 Yang Lu , Rabimba Karanjai , Lei Xu , Weidong Shi

Recent progress in autonomous code generation has fueled excitement around AI agents capable of accelerating scientific discovery by running experiments. However, there is currently no benchmark that evaluates whether such agents can…

人工智能 · 计算机科学 2025-06-25 Gyeongwon James Kim , Alex Wilf , Louis-Philippe Morency , Daniel Fried

Hallucination in generative AI is often treated as a technical failure to produce factually correct output. Yet this framing underrepresents the broader significance of hallucinated content in language models, which may appear fluent,…

计算机与社会 · 计算机科学 2025-10-27 Zihao Li , Weiwei Yi , Jiahong Chen

As AI agents integrate into enterprise applications, their evaluation demands benchmarks that reflect the complexity of real-world operations. Instead, existing benchmarks overemphasize open-domains such as code, use narrow accuracy…

人工智能 · 计算机科学 2026-02-03 Amanda Dsouza , Ramya Ramakrishnan , Charles Dickens , Bhavishya Pohani , Christopher M Glaze

Recent advances in generative AI technologies like large language models have boosted the incorporation of AI assistance in writing workflows, leading to the rise of a new paradigm of human-AI co-creation in writing. To understand how…

计算与语言 · 计算机科学 2024-10-08 Zhuoyan Li , Chen Liang , Jing Peng , Ming Yin

Artificial intelligence (AI) has transformed imaging inverse problems, from medical diagnostics to Earth observation. Yet deep neural networks can produce hallucinations, realistic-looking but incorrect details, undermining their…

机器学习 · 统计学 2026-05-14 David Iagaru , Nina M. Gottschling , Anders C. Hansen , Josselin Garnier

Recent advances in AI enable the automatic generation of visualizations directly from textual prompts using agentic workflows. However, visualizations produced via one-shot generative methods often suffer from insufficient quality,…

人机交互 · 计算机科学 2026-03-19 Roxana Bujack , Li-Ta Lo , Ethan Stam , Ayan Biswas , David Rogers
‹ 上一页 1 2 3 10 下一页 ›