中文
相关论文

相关论文: Addressing and Visualizing Misalignments in Human …

200 篇论文

Modern AI enables a high-level, declarative form of interaction: Users describe the intended outcome they wish an AI to produce, but do not actually create the outcome themselves. In contrast, in traditional user interfaces, users invoke…

人机交互 · 计算机科学 2024-09-18 Michael Terry , Chinmay Kulkarni , Martin Wattenberg , Lucas Dixon , Meredith Ringel Morris

In mobile robot shared control, effectively understanding human motion intention is critical for seamless human-robot collaboration. This paper presents a novel shared control framework featuring planning-level intention prediction. A path…

机器人学 · 计算机科学 2025-11-13 Jinyu Zhang , Lijun Han , Feng Jian , Lingxi Zhang , Hesheng Wang

The field of AI alignment is concerned with AI systems that pursue unintended goals. One commonly studied mechanism by which an unintended goal might arise is specification gaming, in which the designer-provided specification is flawed in a…

机器学习 · 计算机科学 2022-11-03 Rohin Shah , Vikrant Varma , Ramana Kumar , Mary Phuong , Victoria Krakovna , Jonathan Uesato , Zac Kenton

Reward functions, learned or manually specified, are rarely perfect. Instead of accurately expressing human goals, these reward functions are often distorted by human beliefs about how best to achieve those goals. Specifically, these reward…

机器学习 · 计算机科学 2025-07-16 Henrik Marklund , Alex Infanger , Benjamin Van Roy

The Bhatt Conjectures framework introduces rigorous, hierarchical benchmarks for evaluating AI reasoning and understanding, moving beyond pattern matching to assess representation invariance, robustness, and metacognitive self-awareness.…

密码学与安全 · 计算机科学 2025-06-23 Manish Bhatt

In coming years or decades, artificial general intelligence (AGI) may surpass human capabilities across many critical domains. We argue that, without substantial effort to prevent it, AGIs could learn to pursue goals that are in conflict…

人工智能 · 计算机科学 2025-05-06 Richard Ngo , Lawrence Chan , Sören Mindermann

Human-robot collaboration (HRC) relies on accurate and timely recognition of human intentions to ensure seamless interactions. Among common HRC tasks, human-to-robot object handovers have been studied extensively for planning the robot's…

机器人学 · 计算机科学 2025-02-18 Parag Khanna , Nona Rajabi , Sumeyra U. Demir Kanik , Danica Kragic , Mårten Björkman , Christian Smith

As robots across domains start collaborating with humans in shared environments, algorithms that enable them to reason over human intent are important to achieve safe interplay. In our work, we study human intent through the problem of…

机器人学 · 计算机科学 2022-09-14 Ingrid Navarro , Jean Oh

Shared mental models are critical to team success; however, in practice, team members may have misaligned models due to a variety of factors. In safety-critical domains (e.g., aviation, healthcare), lack of shared mental models can lead to…

As artificial intelligence (AI) becomes deeply integrated into critical infrastructures and everyday life, ensuring its safe deployment is one of humanity's most urgent challenges. Current AI models prioritize task optimization over safety,…

人工智能 · 计算机科学 2024-11-08 Joshua T. S. Hewson

The field of AI alignment aims to steer AI systems toward human goals, preferences, and ethical principles. Its contributions have been instrumental for improving the output quality, safety, and trustworthiness of today's AI models. This…

人工智能 · 计算机科学 2024-11-26 Robert West , Roland Aydin

As AI adoption expands across human society, the problem of aligning AI models to match human preferences remains a grand challenge. Currently, the AI alignment field is deeply divided between behavioral and representational approaches,…

计算机与社会 · 计算机科学 2025-08-12 Ben Y. Reis , William La Cava

For robots to interact socially, they must interpret human intentions and anticipate their potential outcomes accurately. This is particularly important for social robots designed for human care, which may face potentially dangerous…

机器人学 · 计算机科学 2024-07-11 Noé Zapata , Gerardo Pérez , Lucas Bonilla , Pedro Núñez , Pilar Bachiller , Pablo Bustos

How to build AI that understands human intentions, and uses this knowledge to collaborate with people? We describe a computational framework for evaluating models of goal inference in the domain of 3D motor actions, which receives as input…

人工智能 · 计算机科学 2021-12-03 Yingdong Qian , Marta Kryven , Tao Gao , Hanbyul Joo , Josh Tenenbaum

Collaboration with artificial intelligence (AI) has improved human decision-making across various domains by leveraging the complementary capabilities of humans and AI. Yet, humans systematically overrely on AI advice, even when their…

人机交互 · 计算机科学 2026-05-15 Joshua Holstein , Patrick Hemmer , Gerhard Satzger , Wei Sun

Many students in introductory programming courses fare poorly in the code writing tasks of the final summative assessment. Such tasks are designed to assess whether novices have developed the analytical skills to translate from the given…

Continual learning in robotics seeks systems that can constantly adapt to changing environments and tasks, mirroring human adaptability. A key challenge is refining dynamics models, essential for planning and control, while addressing…

机器人学 · 计算机科学 2025-09-09 Alejandro Murillo-Gonzalez , Lantao Liu

According to several empirical investigations, despite enhancing human capabilities, human-AI cooperation frequently falls short of expectations and fails to reach true synergy. We propose a task-driven framework that reverses prevalent…

计算机与社会 · 计算机科学 2026-05-26 Saleh Afroogh , Kush R. Varshney , Jason D'Cruz

The cost of error in many high-stakes settings is asymmetric: misdiagnosing pneumonia when absent is an inconvenience, but failing to detect it when present can be life-threatening. Because of this, artificial intelligence (AI) models used…

综合经济学 · 经济学 2025-11-12 David Autor , Andrew Caplin , Daniel Martin , Philip Marx

The booming air transportation industry inevitably burdens air traffic controllers' workload, causing unexpected human factor-related incidents. Current air traffic control systems fail to consider spoken instructions for traffic…

声音 · 计算机科学 2024-10-21 Dongyue Guo , Zheng Zhang , Bo Yang , Jianwei Zhang , Hongyu Yang , Yi Lin