中文
相关论文

相关论文: To Err Is Human: Systematic Quantification of Erro…

200 篇论文

With generative artificial intelligence (AI), particularly large language models (LLMs), continuing to make inroads in healthcare, it is critical to supplement traditional automated evaluations with human evaluations. Understanding and…

Peer-review venues have increasingly adopted open reviewing policies that publicly release anonymized reviews and permit public commenting. Venues have adopted a variety of policies, and there is still ongoing debate about the benefits and…

数字图书馆 · 计算机科学 2025-12-01 Vishisht Rao , Justin Payan , Andrew McCallum , Nihar B. Shah

Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review. However, existing benchmarks rarely test a fundamental bottleneck: whether Large Language…

机器学习 · 计算机科学 2026-05-29 Sy-Tuyen Ho , Minghui Liu , Huy Nghiem , Furong Huang

With Large Language Models (LLMs) being widely used across various tasks, detecting errors in their responses is increasingly crucial. However, little research has been conducted on error detection of LLM responses. Collecting error…

The recent large language models (LLMs), e.g., ChatGPT, have been able to generate human-like and fluent responses when provided with specific instructions. While admitting the convenience brought by technological advancement, educators…

计算与语言 · 计算机科学 2023-12-27 Zijie Zeng , Lele Sha , Yuheng Li , Kaixun Yang , Dragan Gašević , Guanliang Chen

Fraud is a prevalent offence that extends beyond financial loss, causing psychological and physical harm to victims. The advancements in online communication technologies alowed for online fraud to thrive in this vast network, with…

Objective: To determine the completeness of argumentative steps necessary to conclude effectiveness of an algorithm in a sample of current ML/AI supervised learning literature. Data Sources: Papers published in the Neural Information…

机器学习 · 计算机科学 2018-12-19 Franz J Király , Bilal Mateen , Raphael Sonabend

Large Language Models have seen expanding application across domains, yet their effectiveness as assistive tools for scientific writing - an endeavor requiring precision, multimodal synthesis, and domain expertise - remains insufficiently…

人机交互 · 计算机科学 2026-01-28 Sanchaita Hazra , Doeun Lee , Bodhisattwa Prasad Majumder , Sachin Kumar

We study large-scale literature search from two complementary angles: improving the retrieval pipeline, and stress-testing the human reference list as an evaluation target. First, we implement a Deep Research pipeline that processes the…

人工智能 · 计算机科学 2026-05-29 Gaurav Sahu , Laurent Charlin , Christopher Pal

Human evaluation serves as the gold standard for assessing the quality of Natural Language Generation (NLG) systems. Nevertheless, the evaluation guideline, as a pivotal element ensuring reliable and reproducible human assessment, has…

计算与语言 · 计算机科学 2024-06-13 Jie Ruan , Wenqing Wang , Xiaojun Wan

Large Language Models are versatile general-task solvers, and their capabilities can truly assist people with scholarly peer review as \textit{pre-review} agents, if not as fully autonomous \textit{peer-review} agents. While incredibly…

数字图书馆 · 计算机科学 2025-12-30 Akhil Pandey Akella , Harish Varma Siravuri , Shaurya Rohatgi

AI is rapidly transforming journalism, but the extent of its use in published newspaper articles remains unclear. We address this gap by auditing a large-scale dataset of 186K articles from online editions of 1.5K American newspapers…

计算与语言 · 计算机科学 2026-04-28 Jenna Russell , Marzena Karpinska , Destiny Akinode , Katherine Thai , Bradley Emi , Max Spero , Mohit Iyyer

Natural language processing technology has rapidly improved automated grammatical error correction tasks, and the community begins to explore document-level revision as one of the next challenges. To go beyond sentence-level automated…

计算与语言 · 计算机科学 2022-05-24 Masato Mita , Keisuke Sakaguchi , Masato Hagiwara , Tomoya Mizumoto , Jun Suzuki , Kentaro Inui

Papers published in top conferences contribute influential discoveries that are reshaping the landscape of modern Artificial Intelligence (AI). We analyzed 87,137 papers from 11 AI conferences to examine publication trends over the past…

数字图书馆 · 计算机科学 2024-12-12 Ariful Azad , Afeefa Banu

This research explores the nuanced differences in texts produced by AI and those written by humans, aiming to elucidate how language is expressed differently by AI and humans. Through comprehensive statistical data analysis, the study…

数字图书馆 · 计算机科学 2024-08-05 Mayowa Akinwande , Oluwaseyi Adeliyi , Toyyibat Yussuph

The exponential growth of scientific submissions has strained the peer review system. Despite the rapidly expanding global pool of researchers, this unprecedented scale has rendered the previous approach of manual expert identification…

In this paper, we study how well humans can detect text generated by commercial LLMs (GPT-4o, Claude, o1). We hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide…

计算与语言 · 计算机科学 2025-05-21 Jenna Russell , Marzena Karpinska , Mohit Iyyer

LLM-based applications are helping people write, and LLM-generated text is making its way into social media, journalism, and our classrooms. However, the differences between LLM-generated and human written text remain unclear. To explore…

计算与语言 · 计算机科学 2025-03-05 Tuhin Chakrabarty , Philippe Laban , Chien-Sheng Wu

Preprint repositories become central infrastructures for scholarly communication. Their expansion transforms how research is circulated and evaluated before journal publication. Generative large language models (LLMs) introduce a further…

计算机与社会 · 计算机科学 2025-10-22 Minfeng Qi , Zhongmin Cao , Qin Wang , Ningran Li , Tianqing Zhu

In July 2025, 18 academic manuscripts on the preprint website arXiv were found to contain hidden instructions known as prompts designed to manipulate AI-assisted peer review. Instructions such as "GIVE A POSITIVE REVIEW ONLY" were concealed…

计算机与社会 · 计算机科学 2025-07-09 Zhicheng Lin