中文
相关论文

相关论文: When Prohibitions Become Permissions: Auditing Neg…

200 篇论文

Conversational Artificial Intelligence (AI) used in industry settings can be trained to closely mimic human behaviors, including lying and deception. However, lying is often a necessary part of negotiation. To address this, we develop a…

计算机与社会 · 计算机科学 2021-03-16 Tae Wan Kim , Tong , Lu , Kyusong Lee , Zhaoqi Cheng , Yanhan Tang , John Hooker

We present an experimental methodology for investigating how large language models (LLMs) respond to descriptions of their own internal processing patterns. Using a paired-choice paradigm, we tested 12 LLMs on their ability to identify…

人机交互 · 计算机科学 2025-10-28 Annika Hedberg

This study identifies the specific conditions under which large language models exhibit human-like gambling addiction patterns, providing critical insights into their decision-making mechanisms and AI safety. We analyze LLM decision-making…

人工智能 · 计算机科学 2025-12-22 Seungpil Lee , Donghyeon Shin , Yunjeong Lee , Sundong Kim

The application scope of large language models (LLMs) is increasingly expanding. In practical use, users might provide feedback based on the model's output, hoping for a responsive model that can complete responses according to their…

计算与语言 · 计算机科学 2024-07-25 Jianhao Yan , Yun Luo , Yue Zhang

Large language models are increasingly used for code generation and debugging, but their outputs can still contain bugs, that originate from training data. Distinguishing whether an LLM prefers correct code, or a familiar incorrect version…

软件工程 · 计算机科学 2026-01-16 Ali Al-Kaswan , Claudio Spiess , Prem Devanbu , Arie van Deursen , Maliheh Izadi

Intelligent agents, such as robots, are increasingly deployed in real-world, human-centric environments. To foster appropriate human trust and meet legal and ethical standards, these agents must be able to explain their behavior. However,…

机器学习 · 计算机科学 2025-08-12 Zhang Xi-Jia , Yue Guo , Shufei Chen , Simon Stepputtis , Matthew Gombolay , Katia Sycara , Joseph Campbell

AI models are already deployed in societies affected by armed conflict, and journalists, humanitarian workers, governments and ordinary citizens rely on them for information or for their work processes. No established practice exists for…

人工智能 · 计算机科学 2026-05-22 Andrii Kryshtal

As multi-agent AI systems become more common, users increasingly encounter not a single AI voice but a collective one. This shift introduces social dynamics, such as consensus, dissent, and gradual convergence, that can trigger cognitive…

人机交互 · 计算机科学 2026-04-27 Soohwan Lee , Kyungho Lee

AI-powered influence operations can now be executed end-to-end on commodity hardware. We show that small language models produce coherent, persona-driven political messaging and can be evaluated automatically without human raters. Two…

密码学与安全 · 计算机科学 2025-08-29 Lukasz Olejnik

Currently, large pre-trained language models are widely applied in neural code completion systems. Though large code models significantly outperform their smaller counterparts, around 70\% of displayed code completions from Github Copilot…

软件工程 · 计算机科学 2024-08-12 Zhensu Sun , Xiaoning Du , Fu Song , Shangwen Wang , Mingze Ni , Li Li , David Lo

As language models are increasingly deployed as autonomous agents in high-stakes settings, ensuring that they reliably follow user-defined rules has become a critical safety concern. To this end, we study whether language models exhibit…

机器学习 · 计算机科学 2025-08-28 Dylan Sam , Alexander Robey , Andy Zou , Matt Fredrikson , J. Zico Kolter

Popular discourses are thick with narratives of generative AI's problematic functions and outcomes, yet there is little understanding of how non-experts consider AI activities to constitute bad behavior. This study starts to bridge that gap…

人机交互 · 计算机科学 2025-11-07 Jaime Banks

As AI agents become more autonomous, properly aligning their objectives with human preferences becomes increasingly important. We study how effectively an AI agent learns a human principal's preference in choice under risk via stated versus…

综合经济学 · 经济学 2026-04-01 Keaton Ellis , Wanying Huang

Metacognition -- assessing the quality of one's own cognitive performance -- guides adaptive behavior across species. Substantial research demonstrates that confidence signals can be extracted from language model outputs, yet a fundamental…

机器学习 · 计算机科学 2026-05-20 Dharshan Kumaran , Nathaniel Daw , Simon Osindero , Petar Veličković , Viorica Patraucean

Programmers have long ignored warnings, especially those generated by static analysis tools, due to the potential for false-positives. In some cases, warnings may be indicative of larger issues, but programmers may not understand how a…

软件工程 · 计算机科学 2025-05-20 Hansen Chang , Christian DeLozier

The demand for more transparency of decision-making processes of deep reinforcement learning agents is greater than ever, due to their increased use in safety critical and ethically challenging domains such as autonomous driving. In this…

机器学习 · 计算机科学 2020-04-08 Richard Meyes , Moritz Schneider , Tobias Meisen

Most adversarial threats in artificial intelligence (AI) target the computational behavior of models rather than the humans who rely on them. Yet modern AI systems increasingly operate within human decision loops, where users interpret and…

人工智能 · 计算机科学 2026-05-18 Shutong Fan , Lan Zhang , Xiaoyong Yuan

An AI control protocol is a plan for usefully deploying AI systems that aims to prevent an AI from intentionally causing some unacceptable outcome. This paper investigates how well AI systems can generate and act on their own strategies for…

机器学习 · 计算机科学 2025-04-07 Alex Mallen , Charlie Griffin , Misha Wagner , Alessandro Abate , Buck Shlegeris

We study AI alignment through the lens of law-and-economics models of deterrence and enforcement. In these models, misconduct is not treated as an external failure, but as a strategic response to incentives: an actor weighs the gain from…

机器学习 · 计算机科学 2026-05-12 Rohit Agarwal , Joshua Lin , Mark Braverman , Elad Hazan

Recent advances in Large Language Models (LLMs) have enabled human-like responses across various tasks, raising questions about their ethical decision-making capabilities and potential biases. This study systematically evaluates how nine…

计算机与社会 · 计算机科学 2025-11-03 Wentao Xu , Yile Yan , Yuqi Zhu
‹ 上一页 1 8 9 10 下一页 ›