中文
相关论文

相关论文: MIRROR: A Hierarchical Benchmark for Metacognitive…

200 篇论文

Operations Research (OR) relies on expert-driven modeling-a slow and fragile process ill-suited to novel scenarios. While large language models (LLMs) can automatically translate natural language into optimization models, existing…

计算与语言 · 计算机科学 2026-02-05 Yifan Shi , Jialong Shi , Jiayi Wang , Ye Fan , Jianyong Sun

Automatic question generation is a critical task that involves evaluating question quality by considering factors such as engagement, pedagogical value, and the ability to stimulate critical thinking. These aspects require human-like…

计算与语言 · 计算机科学 2025-03-26 Aniket Deroy , Subhankar Maity , Sudeshna Sarkar

The Metacognitive Probe is an exploratory five-task, 15-slot diagnostic that decomposes an LLM's confidence behaviour into five behaviourally-distinct dimensions: confidence calibration (T1-CC), epistemic vigilance (T2-EV), knowledge…

人工智能 · 计算机科学 2026-05-12 Rafael C. T. Oliveira

High-fidelity generative models have narrowed the perceptual gap between synthetic and real images, posing serious threats to media security. Most existing AI-generated image (AIGI) detectors rely on artifact-based classification and…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Ruiqi Liu , Manni Cui , Ziheng Qin , Zhiyuan Yan , Ruoxin Chen , Yi Han , Zhiheng Li , Junkai Chen , ZhiJin Chen , Kaiqing Lin , Jialiang Shen , Lubin Weng , Jing Dong , Yan Wang , Shu Wu

Complex tasks involving tool integration pose significant challenges for Large Language Models (LLMs), leading to the emergence of multi-agent workflows as a promising solution. Reflection has emerged as an effective strategy for correcting…

人工智能 · 计算机科学 2025-06-06 Zikang Guo , Benfeng Xu , Xiaorui Wang , Zhendong Mao

Despite the rapid expansion of Large Language Models (LLMs) in healthcare, robust and explainable evaluation of their ability to assess clinical trial reporting according to CONSORT standards remains an open challenge. In particular,…

人工智能 · 计算机科学 2026-02-26 Sohyeon Jeon , Hyung-Chul Lee

Large Language Models (LLMs) show remarkable proficiency in natural language tasks, yet their frequent overconfidence-misalignment between predicted confidence and true correctness-poses significant risks in critical decision-making…

计算与语言 · 计算机科学 2025-12-15 Prateek Chhikara

Multiple cognitive theories -- Global Workspace Theory, reconstructive episodic memory, inner speech, and complementary learning systems -- converge on a shared set of architectural principles: parallel specialized processing, integrative…

人工智能 · 计算机科学 2026-04-21 Nicole Hsing

Communication is a hallmark of intelligence. In this work, we present MIRROR, an approach to (i) quickly learn human models from human demonstrations, and (ii) use the models for subsequent communication planning in assistive shared-control…

人工智能 · 计算机科学 2022-03-08 Kaiqi Chen , Jeffrey Fong , Harold Soh

Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address implicit, domain-specific risks. To investigate this gap, we introduce a dataset of 3,000…

人工智能 · 计算机科学 2026-01-09 Liang Shan , Kaicheng Shen , Wen Wu , Zhenyu Ying , Chaochao Lu , Yan Teng , Jingqi Huang , Guangze Ye , Guoqing Wang , Liang He

Recent studies show near-perfect language-based predictions of depression scores (R2 = .70), but these "Mirror" models rely on language responses directly from depression assessments to predict depression assessment scores. These methods…

计算与语言 · 计算机科学 2025-10-21 Tong Li , Rasiq Hussain , Mehak Gupta , Joshua R. Oltmanns

Aggregate metacognitive quality scores mask within-model variation across MMLU benchmark domains. We administered 1,500 MMLU items (250 per domain, under an a priori six-domain grouping) to 33 frontier LLMs from eight model families and…

计算与语言 · 计算机科学 2026-05-11 Jon-Paul Cacioli

This investigation presents an empirical analysis of the incompatibility between human psychometric frameworks and Large Language Model evaluation. Through systematic assessment of nine frontier models including GPT-5, Claude Opus 4.1, and…

人工智能 · 计算机科学 2025-11-25 Mohan Reddy

Metacognition, the ability to monitor and regulate one's own reasoning, remains under-evaluated in AI benchmarking. We introduce MEDLEY-BENCH, a benchmark of behavioural metacognition that separates independent reasoning, private…

人工智能 · 计算机科学 2026-04-20 Farhad Abtahi , Abdolamir Karbalaie , Eduardo Illueca-Fernandez , Fernando Seoane

We present an experimental methodology for investigating how large language models (LLMs) respond to descriptions of their own internal processing patterns. Using a paired-choice paradigm, we tested 12 LLMs on their ability to identify…

人机交互 · 计算机科学 2025-10-28 Annika Hedberg

Knowledge probing quantifies how much relational knowledge a language model (LM) has acquired during pre-training. Existing knowledge probes evaluate model capabilities through metrics like prediction accuracy and precision. Such…

计算与语言 · 计算机科学 2026-01-28 Christopher Kissling , Elena Merdjanovska , Alan Akbik

Large language models (LLMs) are increasingly deployed for tabular question answering, yet calibration on structured data is largely unstudied. This paper presents the first systematic comparison of five confidence estimation methods across…

计算与语言 · 计算机科学 2026-04-15 Lukas Voss

Large Language Models (LLMs) have shown remarkable capabilities in environmental perception, reasoning-based decision-making, and simulating complex human behaviors, particularly in interactive role-playing contexts. This paper introduces…

计算与语言 · 计算机科学 2026-01-21 Yin Cai , Zhouhong Gu , Zhaohan Du , Zheyu Ye , Shaosheng Cao , Yiqian Xu , Hongwei Feng , Ping Chen

If AI models can detect when they are being evaluated, the effectiveness of evaluations might be compromised. For example, models could have systematically different behavior during evaluations, leading to less reliable benchmarks for…

计算与语言 · 计算机科学 2025-07-17 Joe Needham , Giles Edkins , Govind Pimpale , Henning Bartsch , Marius Hobbhahn

Large Language Models have demonstrated strong performance on many established reasoning benchmarks. However, these benchmarks primarily evaluate structured skills like quantitative problem-solving, leaving a gap in assessing flexible,…

计算与语言 · 计算机科学 2025-10-30 Deepon Halder , Alan Saji , Thanmay Jayakumar , Ratish Puduppully , Anoop Kunchukuttan , Raj Dabre
‹ 上一页 1 2 3 10 下一页 ›