English
Related papers

Related papers: Evaluating Superhuman Models with Consistency Chec…

200 papers

Understanding emotions is fundamental to human interaction and experience. Humans easily infer emotions from situations or facial expressions, situations from emotions, and do a variety of other affective cognition. How adept is modern AI…

Computation and Language · Computer Science 2026-02-18 Kanishk Gandhi , Zoe Lynch , Jan-Philipp Fränken , Kayla Patterson , Sharon Wambu , Tobias Gerstenberg , Desmond C. Ong , Noah D. Goodman

Agent-based modelling is a powerful tool when simulating human systems, yet when human behaviour cannot be described by simple rules or maximising one's own profit, we quickly reach the limits of this methodology. Machine learning has the…

Multiagent Systems · Computer Science 2022-01-21 Georg Jäger , Daniel Reisinger

Biased human decisions have consequential impacts across various domains, yielding unfair treatment of individuals and resulting in suboptimal outcomes for organizations and society. In recognition of this fact, organizations regularly…

Machine Learning · Computer Science 2024-12-11 Wanxue Dong , Maria De-Arteaga , Maytal Saar-Tsechansky

When we test a theory using data, it is common to focus on correctness: do the predictions of the theory match what we see in the data? But we also care about completeness: how much of the predictable variation in the data is captured by…

Machine Learning · Computer Science 2017-06-22 Jon Kleinberg , Annie Liang , Sendhil Mullainathan

Trustworthy machine learning is of primary importance to the practical deployment of deep learning models. While state-of-the-art models achieve astonishingly good performance in terms of accuracy, recent literature reveals that their…

Machine Learning · Computer Science 2023-02-07 Ailin Deng , Shen Li , Miao Xiong , Zhirui Chen , Bryan Hooi

With the rise of AI systems in real-world applications comes the need for reliable and trustworthy AI. An essential aspect of this are explainable AI systems. However, there is no agreed standard on how explainable AI systems should be…

Artificial Intelligence · Computer Science 2022-07-04 Sascha Saralajew , Ammar Shaker , Zhao Xu , Kiril Gashteovski , Bhushan Kotnis , Wiem Ben Rim , Jürgen Quittek , Carolin Lawrence

Although neural models have performed impressively well on various tasks such as image recognition and question answering, their reasoning ability has been measured in only few studies. In this work, we focus on spatial reasoning and…

Artificial Intelligence · Computer Science 2021-08-19 Hyunjae Kim , Yookyung Koh , Jinheon Baek , Jaewoo Kang

For a newcomer, paraconsistent logics can be difficult to grasp. Even experts in logic can find the concept of paraconsistency to be suspicious or misguided, if not actually wrong. The problem is that although they usually have much in…

Logic · Mathematics 2013-12-17 Jesse Alama

As large-scale, pre-trained language models achieve human-level and superhuman accuracy on existing language understanding tasks, statistical bias in benchmark data and probing studies have recently called into question their true…

Computation and Language · Computer Science 2021-09-13 Shane Storks , Joyce Chai

Whether and how data scientists, statisticians and modellers should be accountable for the AI systems they develop remains a controversial and highly debated topic, especially given the complexity of AI systems and the difficulties in…

Artificial Intelligence · Computer Science 2023-09-12 Cassandra Bird , Daniel Williamson , Sabina Leonelli

In an era increasingly dominated by digital platforms, the spread of misinformation poses a significant challenge, highlighting the need for solutions capable of assessing information veracity. Our research contributes to the field of…

Computation and Language · Computer Science 2024-10-22 Darius Feher , Abdullah Khered , Hao Zhang , Riza Batista-Navarro , Viktor Schlegel

As AI becomes more capable, we entrust it with more general and consequential tasks. The risks from failure grow more severe with increasing task scope. It is therefore important to understand how extremely capable AI models will fail: Will…

Artificial Intelligence · Computer Science 2026-04-13 Alexander Hägele , Aryo Pradipta Gema , Henry Sleight , Ethan Perez , Jascha Sohl-Dickstein

Do AI systems truly understand human concepts or merely mimic surface patterns? We investigate this through chess, where human creativity meets precise strategic concepts. Analyzing a 270M-parameter transformer that achieves…

Machine Learning · Computer Science 2025-11-05 Semyon Lomasov , Judah Goldfeder , Mehmet Hamza Erol , Matthew So , Yao Yan , Addison Howard , Nathan Kutz , Ravid Shwartz Ziv

Facing the current debate on whether Large Language Models (LLMs) attain near-human intelligence levels (Mitchell & Krakauer, 2023; Bubeck et al., 2023; Kosinski, 2023; Shiffrin & Mitchell, 2023; Ullman, 2023), the current study introduces…

Artificial Intelligence · Computer Science 2024-05-21 Junqi Wang , Chunhui Zhang , Jiapeng Li , Yuxi Ma , Lixing Niu , Jiaheng Han , Yujia Peng , Yixin Zhu , Lifeng Fan

Understanding human perceptions of robot performance is crucial for designing socially intelligent robots that can adapt to human expectations. Current approaches often rely on surveys, which can disrupt ongoing human-robot interactions. As…

The evaluation of supervised machine learning models is a critical stage in the development of reliable predictive systems. Despite the widespread availability of machine learning libraries and automated workflows, model assessment is often…

Machine Learning · Computer Science 2026-04-16 Xuanyan Liu , Ignacio Cabrera Martin , Marcello Trovati , Xiaolong Xu , Nikolaos Polatidis

Understanding how humans revise their beliefs in light of new information is crucial for developing AI systems which can effectively model, and thus align with, human reasoning. While theoretical belief revision frameworks rely on a set of…

Artificial Intelligence · Computer Science 2025-06-12 Stylianos Loukas Vasileiou , Antonio Rago , Maria Vanina Martinez , William Yeoh

Human intuition has been simulated by several research projects using artificial intelligence techniques. Most of these algorithms or models lack the ability to handle complications or diversions. Moreover, they also do not explain the…

Artificial Intelligence · Computer Science 2011-06-30 Jitesh Dundas , David Chik

Understanding how people behave in strategic settings--where they make decisions based on their expectations about the behavior of others--is a long-standing problem in the behavioral sciences. We conduct the largest study to date of…

General Economics · Economics 2024-08-16 Jian-Qiao Zhu , Joshua C. Peterson , Benjamin Enke , Thomas L. Griffiths

ML decision-aid systems are increasingly common on the web, but their successful integration relies on people trusting them appropriately: they should use the system to fill in gaps in their ability, but recognize signals that the system…

Human-Computer Interaction · Computer Science 2020-05-25 Harini Suresh , Natalie Lao , Ilaria Liccardi
‹ Prev 1 3 4 5 6 7 10 Next ›