English
Related papers

Related papers: MEDLEY-BENCH: Scale Buys Evaluation but Not Contro…

200 papers

As AI systems progress, we rely more on them to make decisions with us and for us. To ensure that such decisions are aligned with human values, it is imperative for us to understand not only what decisions they make but also how they come…

People navigate complex environments using cues, heuristics, and other strategies, which are often adaptive in stable settings. However, as AI increasingly permeates society's information environments, those become more adaptive and…

Human-Computer Interaction · Computer Science 2026-02-03 Ezequiel Lopez-Lopez , Christoph M. Abels , Philipp Lorenz-Spreen , Stephan Lewandowsky , Stefan M. Herzog

Automated ASD screening tools remain limited by single-architecture evaluations, axis-restricted assessment, and near-exclusive focus on adult cohorts, obscuring age-specific diagnostic patterns critical for early intervention. We introduce…

Machine Learning · Computer Science 2026-05-13 Shubhankit Singh , Hassan Shaikh , Kuldeep Raghuwanshi , Keshav Bulia

With large language models (LLMs) increasingly deployed as cognitive engines for AI agents, the reliability and effectiveness critically hinge on their intrinsic epistemic agency, which remains understudied. Epistemic agency, the ability to…

Artificial Intelligence · Computer Science 2025-06-05 Lingyu Li , Yixu Wang , Haiquan Zhao , Shuqi Kong , Yan Teng , Chunbo Li , Yingchun Wang

We introduce a cross-domain behavioural assay of monitoring-control coupling in LLMs, grounded in the Nelson and Narens (1990) metacognitive framework and applying human psychometric methodology to LLM evaluation. The battery comprises 524…

Computation and Language · Computer Science 2026-04-22 Jon-Paul Cacioli

Large Language Models (LLMs) have demonstrated remarkable capabilities on general text; however, their proficiency in specialized scientific domains that require deep, interconnected knowledge remains largely uncharacterized. Metabolomics…

Computation and Language · Computer Science 2025-10-17 Yuxing Lu , Xukai Zhao , J. Ben Tamo , Micky C. Nnamdi , Rui Peng , Shuang Zeng , Xingyu Hu , Jinzhuo Wang , May D. Wang

Artificial intelligence has advanced rapidly across perception, language, reasoning, and multimodal domains. Yet despite these achievements, modern AI systems remain fundamentally limited in their ability to self-monitor, self-correct, and…

Artificial Intelligence · Computer Science 2025-12-03 Noorbakhsh Amiri Golilarz , Sindhuja Penchala , Shahram Rahimi

Existing AI evaluation practices often fail to capture how systems actually perform in low-resource environments, where operational constraints shape usability as much as model quality. Through a structured analysis of existing benchmark…

Artificial Intelligence · Computer Science 2026-05-28 Aakash Pant , Kavya Shah , Apoorv Agnihotri , Sneha Nikam , Prasaanth Balraj , Nakul Jain

If an AI agent makes decisions on a person's behalf, those decisions must align with its user. We introduce representational accuracy to measure how faithfully a system captures a person's interpretation. An interpretive layer is…

Computation and Language · Computer Science 2026-05-29 Aarik Gulaya

Computational metacognition represents a cognitive systems perspective on high-order reasoning in integrated artificial systems that seeks to leverage ideas from human metacognition and from metareasoning approaches in artificial…

Artificial Intelligence · Computer Science 2022-02-01 Michael Cox , Zahiduddin Mohammad , Sravya Kondrakunta , Ventaksamapth Raja Gogineni , Dustin Dannenhauer , Othalia Larue

Fine-grained manipulation marks a regime where global scene context no longer suffices, and success hinges on the tight coupling of local attribute grounding, high-fidelity spatial perception, and constraint-respecting motor execution.…

Robotics · Computer Science 2026-05-20 He-Yang Xu , Pengyuan Zhang , Zongyuan Ge , Xiaoshuai Hao , Serge Belongie , Xin Geng , Yuxin Peng , Xiu-Shen Wei

Emotional intelligence (EI), the ability to perceive, understand, and respond appropriately to others' emotional states, is central to human communication, and increasingly important to assess as LLMs assume conversational roles in everyday…

Artificial Intelligence · Computer Science 2026-05-29 Kate M. Lubrano , Faisal Sayed , Ankita Rathod , Akshansh , Craver Corbyn Thomas-Smith , Mark E. Whiting , Karina Nguyen

Beyond representing the external world, humans also represent their own cognitive processes. In the context of perception, this metacognition helps us identify unreliable percepts, such as when we recognize that we are seeing an illusion.…

Artificial Intelligence · Computer Science 2020-12-01 Marlene Berke , Mario Belledonne , Julian Jara-Ettinger

Generative artificial intelligence (AI) is a promising direction for augmenting clinical diagnostic decision support and reducing diagnostic errors, a leading contributor to medical errors. To further the development of clinical AI systems,…

Computation and Language · Computer Science 2023-06-14 Brihat Sharma , Yanjun Gao , Timothy Miller , Matthew M. Churpek , Majid Afshar , Dmitriy Dligach

Metacognition, defined as the awareness and regulation of one's cognitive processes, is central to human adaptability in unknown situations. In contrast, current autonomous agents often struggle in novel environments due to their limited…

Machine Learning · Computer Science 2025-11-18 Rodolfo Valiente , Praveen K. Pilly

We introduce Cube Bench, a Rubik's-cube benchmark for evaluating spatial and sequential reasoning in multimodal large language models (MLLMs). The benchmark decomposes performance into five skills: (i) reconstructing cube faces from images…

Computation and Language · Computer Science 2025-12-24 Dhruv Anand , Ehsan Shareghi

Recent benchmark studies have claimed that AI has approached or even surpassed human-level performances on various cognitive tasks. However, this position paper argues that current AI evaluation paradigms are insufficient for assessing…

This study examines the relationship between speech representations and the hierarchical structure of cognitive assessment in mild cognitive impairment. Utilizing 5,754 German neuropsychological assessment recordings, we evaluate six…

More than one hundred benchmarks have been developed to test the commonsense knowledge and commonsense reasoning abilities of artificial intelligence (AI) systems. However, these benchmarks are often flawed and many aspects of common sense…

Artificial Intelligence · Computer Science 2023-02-24 Ernest Davis

This position paper argues for metacognition as a general design principle for creating more accurate, secure, and efficient AI. The metacognitive solution involves systems monitoring their own states and judiciously allocating resources…

Artificial Intelligence · Computer Science 2026-05-18 Sergei Chuprov , Richard D. Lange , Leon Reznik , Paulo Shakarian , Raman Zatsarenko , Dmitrii Korobeinikov