中文
相关论文

相关论文: Privileged Self-Access Matters for Introspection i…

200 篇论文

The debate around bias in AI systems is central to discussions on algorithmic fairness. However, the term bias often lacks a clear definition, despite frequently being contrasted with fairness, implying that an unbiased model is inherently…

人工智能 · 计算机科学 2025-02-26 Chiara Lindloff , Ingo Siegert

Would artificial intelligence (AI) cut interest rates or adopt conservative monetary policy? Would it deregulate or opt for a more controlled economy? As AI use by economic policymakers, academics, and market participants grows…

计算机与社会 · 计算机科学 2025-12-10 Maxim Chupilkin

There are many goals for an AI that could become dangerous if the AI becomes superintelligent or otherwise powerful. Much work on the AI control problem has been focused on constructing AI goals that are safe even for such AIs. This paper…

人工智能 · 计算机科学 2017-05-31 Stuart Armstrong , Benjamin Levinstein

Artificial Intelligence (AI) is about making computers that do the sorts of things that minds can do, and as we progress towards this goal, we tend to increasingly delegate human tasks to machines. However, AI systems usually do these tasks…

人工智能 · 计算机科学 2024-06-12 Peter R. Lewis , Stefan Sarkadi

The interactive nature of Large Language Models (LLMs), which closely track user data and context, has prompted users to share personal and private information in unprecedented ways. Even when users opt out of allowing their data to be used…

密码学与安全 · 计算机科学 2025-08-26 GodsGift Uzor , Hasan Al-Qudah , Ynes Ineza , Abdul Serwadda

Privacy leakage in AI-based decision processes poses significant risks, particularly when sensitive information can be inferred. We propose a formal framework to audit privacy leakage using abductive explanations, which identifies minimal…

人工智能 · 计算机科学 2025-11-14 Belona Sonna , Alban Grastien , Claire Benn

Motivated by mitigating potentially harmful impacts of technologies, the AI community has formulated and accepted mathematical definitions for certain pillars of accountability: e.g. privacy, fairness, and model transparency. Yet, we argue…

机器学习 · 计算机科学 2022-12-16 Teresa Datta , Daniel Nissani , Max Cembalest , Akash Khanna , Haley Massa , John P. Dickerson

Self-consciousness, the introspection of one's existence and thoughts, represents a high-level cognitive process. As language models advance at an unprecedented pace, a critical question arises: Are these models becoming self-conscious?…

计算与语言 · 计算机科学 2024-10-25 Sirui Chen , Shu Yu , Shengjie Zhao , Chaochao Lu

How do large language models (LLMs) obtain their answers? The ability to explain and control an LLM's reasoning process is key for reliability, transparency, and future model developments. We propose SelfIE (Self-Interpretation of…

计算与语言 · 计算机科学 2024-03-27 Haozhe Chen , Carl Vondrick , Chengzhi Mao

This paper provides empirical concerns about post-hoc explanations of black-box ML models, one of the major trends in AI explainability (XAI), by showing its lack of interpretability and societal consequences. Using a representative…

人机交互 · 计算机科学 2021-10-01 Jean-Marie John-Mathews

Machine Learning algorithms are technological key enablers for artificial intelligence (AI). Due to the inherent complexity, these learning algorithms represent black boxes and are difficult to comprehend, therefore influencing compliance…

计算机与社会 · 计算机科学 2020-02-21 NIklas Kuhl , Jodie Lobana , Christian Meske

Attribute inference - the process of analyzing publicly available data in order to uncover hidden information - has become a major threat to privacy, given the recent technological leap in machine learning. One way to tackle this threat is…

人工智能 · 计算机科学 2023-04-25 Marcin Waniek , Navya Suri , Abdullah Zameek , Bedoor AlShebli , Talal Rahwan

This position paper argues that the AI/ML community should stop overclaiming and retire the label "positive backdoor," and instead treat trigger-activated hidden behaviors as Secret Alignment. Crucially, protective claims based on Secret…

密码学与安全 · 计算机科学 2026-05-28 Jianwei Li , Jung-Eun Kim

Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review. However, existing benchmarks rarely test a fundamental bottleneck: whether Large Language…

机器学习 · 计算机科学 2026-05-29 Sy-Tuyen Ho , Minghui Liu , Huy Nghiem , Furong Huang

Leakage of data from publicly available Machine Learning (ML) models is an area of growing significance as commercial and government applications of ML can draw on multiple sources of data, potentially including users' and clients'…

We motivate and outline a programme for a formal theory of measurement of artificial intelligence. We argue that formalising measurement for AI will allow researchers, practitioners, and regulators to: (i) make comparisons between systems…

人工智能 · 计算机科学 2025-07-09 Elija Perrier

We study behavioral self-awareness -- an LLM's ability to articulate its behaviors without requiring in-context examples. We finetune LLMs on datasets that exhibit particular behaviors, such as (a) making high-risk economic decisions, and…

计算与语言 · 计算机科学 2025-01-22 Jan Betley , Xuchan Bao , Martín Soto , Anna Sztyber-Betley , James Chua , Owain Evans

Ensuring privacy during inference stage is crucial to prevent malicious third parties from reconstructing users' private inputs from outputs of public models. Despite a large body of literature on privacy preserving learning (which ensures…

密码学与安全 · 计算机科学 2024-12-02 Fengwei Tian , Ravi Tandon

As robots aspire for long-term autonomous operations in complex dynamic environments, the ability to reliably take mission-critical decisions in ambiguous situations becomes critical. This motivates the need to build systems that have…

机器人学 · 计算机科学 2016-08-01 Shreyansh Daftry , Sam Zeng , J. Andrew Bagnell , Martial Hebert

Large language models (LLMs) are increasingly employed for decision-support across multiple domains. We investigate whether these models display a systematic preferential bias in favor of artificial intelligence (AI) itself. Across three…

计算与语言 · 计算机科学 2026-01-21 Benaya Trabelsi , Jonathan Shaki , Sarit Kraus