中文
相关论文

相关论文: The Role of the Availability Heuristic in Multiple…

200 篇论文

The availability heuristic is a strategy that people use to make quick decisions but often lead to systematic errors. We propose three ways that visualization could facilitate unbiased decision-making. First, visualizations can alter the…

人机交互 · 计算机科学 2016-10-11 Evanthia Dimara , Pierre Dragicevic , Anastasia Bezerianos

Questions involving commonsense reasoning about everyday situations often admit many $\textit{possible}$ or $\textit{plausible}$ answers. In contrast, multiple-choice question (MCQ) benchmarks for commonsense reasoning require a hard…

计算与语言 · 计算机科学 2024-10-16 Shramay Palta , Nishant Balepur , Peter Rankel , Sarah Wiegreffe , Marine Carpuat , Rachel Rudinger

Multiple-choice questions (MCQ) are frequently used to assess large language models (LLMs). Typically, an LLM is given a question and selects the answer deemed most probable after adjustments for factors like length. Unfortunately, LLMs may…

计算与语言 · 计算机科学 2024-06-12 Aidar Myrzakhan , Sondos Mahmoud Bsharat , Zhiqiang Shen

The difficulty of multiple-choice questions (MCQs) is a crucial factor for educational assessments. Predicting MCQ difficulty is challenging since it requires understanding both the complexity of reaching the correct option and the…

人工智能 · 计算机科学 2025-03-12 Wanyong Feng , Peter Tran , Stephen Sireci , Andrew Lan

In this work, we consider how preference models in interactive recommendation systems determine the availability of content and users' opportunities for discovery. We propose an evaluation procedure based on stochastic reachability to…

信息检索 · 计算机科学 2021-07-05 Mihaela Curmei , Sarah Dean , Benjamin Recht

The Availability bias, manifested in the over-representation of extreme eventualities in decision-making, is a well-known cognitive bias, and is generally taken as evidence of human irrationality. In this work, we present the first…

神经元与认知 · 定量生物学 2018-01-31 Ardavan S. Nobandegani , Kevin da Silva Castanheira , A. Ross Otto , Thomas R. Shultz

Decision-making AI agents are often faced with two important challenges: the depth of the planning horizon, and the branching factor due to having many choices. Hierarchical reinforcement learning methods aim to solve the first problem, by…

机器学习 · 计算机科学 2022-01-25 Andrei Nica , Khimya Khetarpal , Doina Precup

Reasoning models represent a significant advance in LLM capabilities, particularly for complex reasoning tasks such as mathematics and coding. Previous studies confirm that parallel test-time compute-sampling multiple solutions and…

机器学习 · 计算机科学 2025-10-27 Raul Cavalcante Dinardi , Bruno Yamamoto , Anna Helena Reali Costa , Artur Jordao

Predicting the difficulty of multiple-choice questions (MCQs) is important for effective assessment, yet current methods typically assume a unimodal student ability distribution, overlooking the heterogeneous nature of student…

计算机与社会 · 计算机科学 2026-05-19 Dhriti Krishnan , Jaromir Savelka

In this article, I show that a recent family of quantum algorithms, based on the quantum amplitude amplification algorithm, can be used to describe a cognitive heuristic called availability bias. The amplitude amplification algorithm is…

量子物理 · 物理学 2008-10-07 Riccardo Franco

A popular strategy for active learning is to specifically target a reduction in epistemic uncertainty, since aleatoric uncertainty is often considered as being intrinsic to the system of interest and therefore not reducible. Yet,…

统计方法学 · 统计学 2024-12-12 Jake Thomas , Jeremie Houssineau

Despite the prevalence of voting systems in the real world there is no consensus among researchers of how people vote strategically, even in simple voting settings. This paper addresses this gap by comparing different approaches that have…

计算机科学与博弈论 · 计算机科学 2019-09-24 Roy Fairstein , Adam Lauz , Kobi Gal , Reshef Meir

In an educational setting, an estimate of the difficulty of multiple-choice questions (MCQs), a commonly used strategy to assess learning progress, constitutes very useful information for both teachers and students. Since human assessment…

计算与语言 · 计算机科学 2025-04-21 Leonidas Zotos , Hedderik van Rijn , Malvina Nissim

The handling of probabilities in the form of uncertainty or partial information is an essential task for LLMs in many settings and applications. A common approach to evaluate an LLM's probabilistic reasoning capabilities is to assess its…

人工智能 · 计算机科学 2026-02-12 Manuel Mondal , Ljiljana Dolamic , Gérôme Bovet , Philippe Cudré-Mauroux , Julien Audiffren

The widespread adoption of Large Language Models (LLMs) has become commonplace, particularly with the emergence of open-source models. More importantly, smaller models are well-suited for integration into consumer devices and are frequently…

计算与语言 · 计算机科学 2024-08-16 Aisha Khatun , Daniel G. Brown

While large language models (LLMs) like GPT-3 have achieved impressive results on multiple choice question answering (MCQA) tasks in the zero, one, and few-shot settings, they generally lag behind the MCQA state of the art (SOTA). MCQA…

计算与语言 · 计算机科学 2023-03-20 Joshua Robinson , Christopher Michael Rytting , David Wingate

In this paper, we study the possibility of almost unsupervised Multiple Choices Question Answering (MCQA). Starting from very basic knowledge, MCQA model knows that some choices have higher probabilities of being correct than the others.…

计算与语言 · 计算机科学 2021-11-02 Chi-Liang Liu , Hung-yi Lee

This paper reconsiders the problem of the absent-minded driver who must choose between alternatives with different payoff with imperfect recall and varying degrees of knowledge of the system. The classical absent-minded driver problem…

人工智能 · 计算机科学 2017-02-21 Subhash Kak

Hallucination is often viewed as a direct consequence of missing knowledge: a model answers incorrectly when the correct answer is absent from its generation-time distribution, and correctly when it is present. We test this assumption by…

计算与语言 · 计算机科学 2026-05-22 Jewon Yeom , Jaewon Sok , Heejun Kim , Seonghyeon Park , Jeongjae Park , Taesup Kim

Multiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing, but we argue for its reform. We first reveal flaws in MCQA's format, as it struggles to: 1) test generation/subjectivity;…

计算与语言 · 计算机科学 2025-06-03 Nishant Balepur , Rachel Rudinger , Jordan Lee Boyd-Graber
‹ 上一页 1 2 3 10 下一页 ›