中文
相关论文

相关论文: Brief Notes on Hard Takeoff, Value Alignment, and …

200 篇论文

Explainable AI is an emerging field providing solutions for acquiring insights into automated systems' rationale. It has been put on the AI map by suggesting ways to tackle key ethical and societal issues. Existing explanation techniques…

机器学习 · 计算机科学 2022-05-02 Ioannis Mollas , Nick Bassiliades , Grigorios Tsoumakas

Sentiment analysis is known as one of the most crucial tasks in the field of natural language processing and Convolutional Neural Network (CNN) is one of those prominent models that is commonly used for this aim. Although convolutional…

计算与语言 · 计算机科学 2021-02-24 Hossein Sadr , Mozhdeh Nazari Solimandarabi , Mir Mohsen Pedram , Mohammad Teshnehlab

As the deployment of artificial intelligence (AI) is changing many fields and industries, there are concerns about AI systems making decisions and recommendations without adequately considering various ethical aspects, such as…

计算机与社会 · 计算机科学 2023-10-02 Conrad Sanderson , Qinghua Lu , David Douglas , Xiwei Xu , Liming Zhu , Jon Whittle

Objective: This paper describes the development of hybrid artificial intelligence strategies for drone navigation. Methods: The navigation module combines a deep learning model with a rule-based engine depending on the agent state. The deep…

人工智能 · 计算机科学 2025-01-09 Rubén San-Segundo , Lucía Angulo , Manuel Gil-Martín , David Carramiñana , Ana M. Bernardos

Traditional methods for aligning Large Language Models (LLMs), such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on implicit principles, limiting interpretability. Constitutional AI…

机器学习 · 计算机科学 2025-04-01 Carl-Leander Henneking , Claas Beger

Verified artificial intelligence (AI) is the goal of designing AI-based systems that that have strong, ideally provable, assurances of correctness with respect to mathematically-specified requirements. This paper considers Verified AI from…

人工智能 · 计算机科学 2020-07-24 Sanjit A. Seshia , Dorsa Sadigh , S. Shankar Sastry

Does AI understand human values? While this remains an open philosophical question, we take a pragmatic stance by introducing VAPT, the Value-Alignment Perception Toolkit, for studying how LLMs reflect people's values and how people judge…

人机交互 · 计算机科学 2026-04-15 Bhada Yun , Renn Su , April Yi Wang

With the increased expectation of artificial intelligence, academic research face complex questions of human-centred, responsible and trustworthy technology embedded into society and culture. Several academic debates, social consultations…

其他计算机科学 · 计算机科学 2020-05-07 Katalin Feher , Asta Zelenkauskaite

Being a complex subject of major importance in AI Safety research, value alignment has been studied from various perspectives in the last years. However, no final consensus on the design of ethical utility functions facilitating AI value…

人工智能 · 计算机科学 2019-07-02 Nadisha-Marie Aliman , Leon Kester

Explainability of AI models is an important topic that can have a significant impact in all domains and applications from autonomous driving to healthcare. The existing approaches to explainable AI (XAI) are mainly limited to simple machine…

机器学习 · 计算机科学 2023-05-24 Poushali Sengupta , Yan Zhang , Sabita Maharjan , Frank Eliassen

Since the introduction of strong anticipation by D.~Dubois the numerous investigations of concrete systems have been proposed. In proposed paper the new examples of discrete dynamical systems with anticipation are considered. The…

适应与自组织系统 · 物理学 2019-08-22 Alexander Makarenko

The existence of a homogeneous decomposition for continuous and epi-translation invariant valuations on super-coercive functions is established. Continuous and epi-translation invariant valuations that are epi-homogeneous of degree $n$ are…

度量几何 · 数学 2020-05-15 A. Colesanti , M. Ludwig , F. Mussnig

The critical inquiry pervading the realm of Philosophy, and perhaps extending its influence across all Humanities disciplines, revolves around the intricacies of morality and normativity. Surprisingly, in recent years, this thematic thread…

人工智能 · 计算机科学 2024-06-19 Nicholas Kluge Corrêa

Qualitative Choice Logic (QCL) and Conjunctive Choice Logic (CCL) are formalisms for preference handling, with especially QCL being well established in the field of AI. So far, analyses of these logics need to be done on a case-by-case…

计算机科学中的逻辑 · 计算机科学 2021-06-10 Michael Bernreiter , Jan Maly , Stefan Woltran

Abstraction, counterexample-guided refinement, and interpolation are techniques that are essential to the success of predicate-based program analysis. These techniques have not yet been applied together to explicit-value program analysis.…

软件工程 · 计算机科学 2013-01-01 Dirk Beyer , Stefan Löwe

With the rapid advancement of large language models (LLMs), aligning them with human values for safety and ethics has become a critical challenge. This problem is especially challenging when multiple, potentially conflicting human values…

机器学习 · 计算机科学 2025-11-25 Hefei Xu , Le Wu , Chen Cheng , Hao Liu

The truthfulness of existing explanation methods in authentically elucidating the underlying model's decision-making process has been questioned. Existing methods have deviated from faithfully representing the model, thus susceptible to…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Sangyu Han , Yearim Kim , Nojun Kwak

The ability to interpret machine learning models has become increasingly important now that machine learning is used to inform consequential decisions. We propose an approach called model extraction for interpreting complex, blackbox…

机器学习 · 计算机科学 2018-03-14 Osbert Bastani , Carolyn Kim , Hamsa Bastani

AI Alignment research seeks to align human and AI goals to ensure independent actions by a machine are always ethical. This paper argues empathy is necessary for this task, despite being often neglected in favor of more deductive…

神经与进化计算 · 计算机科学 2023-12-14 Devin Gonier , Adrian Adduci , Cassidy LoCascio

Although AI has become increasingly smart, its wisdom has not kept pace. In this article, we examine what is known about human wisdom and sketch a vision of its AI counterpart. We analyze human wisdom as a set of strategies for solving…