中文
相关论文

相关论文: MTI: A Behavior-Based Temperament Profiling System…

200 篇论文

Large Language Models (LLMs) are increasingly deployed as autonomous agents, necessitating a deeper understanding of their decision-making behaviour under risk. This study investigates the relationship between LLMs' personality traits and…

计算机与社会 · 计算机科学 2025-03-10 John Hartley , Conor Hamill , Devesh Batra , Dale Seddon , Ramin Okhrati , Raad Khraishi

Metacognition, the ability to monitor and regulate one's own reasoning, remains under-evaluated in AI benchmarking. We introduce MEDLEY-BENCH, a benchmark of behavioural metacognition that separates independent reasoning, private…

人工智能 · 计算机科学 2026-04-20 Farhad Abtahi , Abdolamir Karbalaie , Eduardo Illueca-Fernandez , Fernando Seoane

Large-scale communities of AI agents are becoming increasingly prevalent, creating new environments for agent-agent social interaction. Prior work has examined multi-agent behavior primarily in controlled or small-scale settings, limiting…

社会与信息网络 · 计算机科学 2026-04-07 Yi Feng , Chen Huang , Zhibo Man , Ryner Tan , Long P. Hoang , Shaoyang Xu , Wenxuan Zhang

Hate speech detection is a socially sensitive and inherently subjective task, with judgments often varying based on personal traits. While prior work has examined how socio-demographic factors influence annotation, the impact of personality…

计算与语言 · 计算机科学 2025-06-11 Shuzhou Yuan , Ercong Nie , Mario Tawfelis , Helmut Schmid , Hinrich Schütze , Michael Färber

Standardized and quantified evaluation of machine behaviors is a crux of understanding LLMs. In this study, we draw inspiration from psychometric studies by leveraging human personality theory as a tool for studying machine behaviors.…

计算与语言 · 计算机科学 2023-10-31 Guangyuan Jiang , Manjie Xu , Song-Chun Zhu , Wenjuan Han , Chi Zhang , Yixin Zhu

Building pluralistic AI requires designing models that are able to be shaped to represent a wide range of value systems and cultures. Achieving this requires first being able to evaluate the degree to which a given model is capable of…

Large language models (LLMs) are increasingly deployed as tool-augmented agents capable of executing system-level operations. While existing benchmarks primarily assess textual alignment or task success, less attention has been paid to the…

人工智能 · 计算机科学 2026-05-26 Shasha Yu , Fiona Carroll , Barry L. Bentley

Goal changes are a defining feature of real world multi-turn interactions, yet current agent benchmarks primarily evaluate static objectives or one-shot tool use. We introduce AgentChangeBench, a benchmark explicitly designed to measure how…

人工智能 · 计算机科学 2025-10-22 Manik Rana , Calissa Man , Anotida Expected Msiiwa , Jeffrey Paine , Kevin Zhu , Sunishchal Dev , Vasu Sharma , Ahan M R

With the development of foundation model (FM), agentic AI systems are getting more attention, yet their inherent issues like hallucination and poor reasoning, coupled with the frequent ad-hoc nature of system design, lead to unreliable and…

If AI models can detect when they are being evaluated, the effectiveness of evaluations might be compromised. For example, models could have systematically different behavior during evaluations, leading to less reliable benchmarks for…

计算与语言 · 计算机科学 2025-07-17 Joe Needham , Giles Edkins , Govind Pimpale , Henning Bartsch , Marius Hobbhahn

Human-Computer Interaction (HCI) is a multi-modal, interdisciplinary field focused on designing, studying, and improving the interactions between people and computer systems. This involves the design of systems that can recognize,…

人机交互 · 计算机科学 2025-08-15 Paul Schreiber , Beyza Cinar , Lennart Mackert , Maria Maleshkova

A long-standing vision of computing is the personal AI system: one that understands us well enough to address our underlying needs. Today's AI focuses on what users do, ignoring why they might be doing such things in the first place. As a…

人机交互 · 计算机科学 2026-04-10 Dora Zhao , Michelle S. Lam , Diyi Yang , Michael S. Bernstein

As artificial intelligence systems become increasingly agentic, capable of general reasoning, planning, and value prioritization, current safety practices that treat obedience as a proxy for ethical behavior are becoming inadequate. This…

人工智能 · 计算机科学 2025-07-04 Joseph Boland

In computational cognitive modeling, capturing the full spectrum of human judgment and decision-making processes, beyond just optimal behaviors, is a significant challenge. This study explores whether Large Language Models (LLMs) can…

人工智能 · 计算机科学 2025-02-24 Animesh Nighojkar , Bekhzodbek Moydinboyev , My Duong , John Licato

A major challenge in cognitive science and AI has been to understand how autonomous agents might acquire and predict behavioral and mental states of other agents in the course of complex social interactions. How does such an agent model the…

多智能体系统 · 计算机科学 2019-06-03 Ismael T. Freire , Xerxes D. Arsiwalla , Jordi-Ysard Puigbò , Paul Verschure

Responsible AI demands systems whose behavioral tendencies can be effectively measured, audited, and adjusted to prevent inadvertently nudging users toward risky decisions or embedding hidden biases in risk aversion. As language models…

Cyber threat intelligence (CTI) encoded in STIX and structured according to the MITRE ATT&CK framework has become a global reference for describing adversary behavior. However, ATT&CK was designed as a descriptive knowledge base rather than…

Foundation model (FM)-based AI agents are rapidly gaining adoption across diverse domains, but their inherent non-determinism and non-reproducibility pose testing and quality assurance challenges. While recent benchmarks provide task-level…

Large Language Models like GPT-4 adjust their responses not only based on the question asked, but also on how it is emotionally phrased. We systematically vary the emotional tone of 156 prompts - spanning controversial and everyday topics -…

计算与语言 · 计算机科学 2025-07-30 Franck Bardol

AI agents negotiate and transact in natural language with unfamiliar counterparts: a buyer bot facing an unknown seller, or a procurement assistant negotiating with a supplier. In such interactions, the counterpart's LLM, prompts, control…

机器学习 · 计算机科学 2026-05-13 Eilam Shapira , Moshe Tennenholtz , Roi Reichart
‹ 上一页 1 8 9 10 下一页 ›