中文
相关论文

相关论文: Astra: AI Safety, Trust, & Risk Assessment

200 篇论文

As artificial intelligence systems become increasingly powerful and pervasive, democratic societies face unprecedented challenges in governing these technologies while preserving core democratic values and institutions. This paper presents…

计算机与社会 · 计算机科学 2025-11-18 Subramanyam Sahoo , Aditi Chhawacharia

The potential presented by Artificial Intelligence (AI) for healthcare has long been recognised by the technical community. More recently, this potential has been recognised by policymakers, resulting in considerable public and private…

人工智能 · 计算机科学 2021-04-15 Jessica Morley , Caroline Morton , Kassandra Karpathakis , Mariarosaria Taddeo , Luciano Floridi

As AI systems' sophistication and proliferation have increased, awareness of the risks has grown proportionally (Sorkin et al. 2023). In response, calls have grown for stronger emphasis on disclosure and transparency in the AI industry…

人工智能 · 计算机科学 2023-09-26 Eli Sherman , Ian W. Eisenberg

Two years after publicly launching the AI Incident Database (AIID) as a collection of harms or near harms produced by AI in the world, a backlog of "issues" that do not meet its incident ingestion criteria have accumulated in its review…

计算机与社会 · 计算机科学 2022-11-21 Sean McGregor , Kevin Paeth , Khoa Lam

As AI systems proliferate in society, the AI community is increasingly preoccupied with the concept of AI Safety, namely the prevention of failures due to accidents that arise from an unanticipated departure of a system's behavior from…

计算机与社会 · 计算机科学 2024-01-23 Inioluwa Deborah Raji , Roel Dobbe

Autonomous Vehicles (AV) are expected to bring considerable benefits to society, such as traffic optimization and accidents reduction. They rely heavily on advances in many Artificial Intelligence (AI) approaches and techniques. However,…

India generates substantial volumes of public agricultural data, yet artificial intelligence (AI) adoption in farming remains limited and largely confined to pilot initiatives. This paper examines this gap by assessing India's agricultural…

综合经济学 · 经济学 2026-03-27 K. B. Vedamurthy , Manojkumar Patil , Vaishnavi , Priyanka V , Suman L , Ajayakumar , Sagar

The rapid development of Artificial Intelligence (AI) technology has enabled the deployment of various systems based on it. However, many current AI systems are found vulnerable to imperceptible attacks, biased against underrepresented…

人工智能 · 计算机科学 2022-05-27 Bo Li , Peng Qi , Bo Liu , Shuai Di , Jingen Liu , Jiquan Pei , Jinfeng Yi , Bowen Zhou

Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon massively amplify…

Artificial Intelligence (AI) has made impressive progress in recent years and represents a key technology that has a crucial impact on the economy and society. However, it is clear that AI and business models based on it can only reach…

We introduce a conceptual framework and provide considerations for the institutional design of AI incident reporting systems, i.e., processes for collecting information about safety- and rights-related events caused by general-purpose AI.…

计算机与社会 · 计算机科学 2026-04-15 Kevin Wei , Lennart Heim

Artificial Intelligence (AI) is a widely developed and adopted technology across entire industry sectors. Integrating environmental, social, and governance (ESG) considerations with AI investments is crucial for ensuring ethical and…

人工智能 · 计算机科学 2024-08-07 Sung Une Lee , Harsha Perera , Yue Liu , Boming Xia , Qinghua Lu , Liming Zhu , Jessica Cairns , Moana Nottage

The Robust Artificial Intelligence System Assurance (RAISA) workshop will focus on research, development and application of robust artificial intelligence (AI) and machine learning (ML) systems. Rather than studying robustness with respect…

人工智能 · 计算机科学 2022-02-11 Olivia Brown , Brad Dillman

Prominent AI experts have suggested that companies developing high-risk AI systems should be required to show that such systems are safe before they can be developed or deployed. The goal of this paper is to expand on this idea and explore…

计算机与社会 · 计算机科学 2024-06-25 Akash R. Wasil , Joshua Clymer , David Krueger , Emily Dardaman , Simeon Campos , Evan R. Murphy

The prevailing methodologies for visualizing AI risks have focused on technical issues such as data biases and model inaccuracies, often overlooking broader societal risks like job loss and surveillance. Moreover, these visualizations are…

人机交互 · 计算机科学 2025-02-11 Edyta Bogucka , Sanja Šćepanović , Daniele Quercia

This paper proposes a novel framework for developing safe Artificial General Intelligence (AGI) by combining Active Inference principles with Large Language Models (LLMs). We argue that traditional approaches to AI safety, focused on…

人工智能 · 计算机科学 2025-08-11 Bo Wen

Modern security environments generate fragmented signals across cloud resources, identities, configurations, and third-party security tools. Although AI-native security assistants improve access to this data, they remain largely reactive:…

密码学与安全 · 计算机科学 2026-05-13 Gal Engelberg , Leon Goldberg , Konstantin Koutsyi , Boris Plotnikov , Tiltan Gilat , Ben Benhemo

Frontier AI both amplifies existing risks and introduces qualitatively novel challenges. Not only is there a notable lack of stable scientific consensus resulting from the rapid pace of technological change, but emerging frontier AI safety…

Artificial intelligence (AI) has been advancing at a fast pace and it is now poised for deployment in a wide range of applications, such as autonomous systems, medical diagnosis and natural language processing. Early adoption of AI…

机器学习 · 计算机科学 2023-09-21 Marta Kwiatkowska , Xiyue Zhang

Artificial intelligence (AI) is rapidly transforming healthcare, enabling fast development of tools like stress monitors, wellness trackers, and mental health chatbots. However, rapid and low-barrier development can introduce risks of bias,…

计算与语言 · 计算机科学 2026-04-09 Xingmeng Zhao , Tongnian Wang , Dan Schumacher , Veronica Rammouz , Anthony Rios