中文
相关论文

相关论文: X-Risk Analysis for AI Research

200 篇论文

The advancements in generative AI inevitably raise concerns about their risks and safety implications, which, in return, catalyzes significant progress in AI safety. However, as this field continues to evolve, a critical question arises:…

计算机与社会 · 计算机科学 2025-01-14 Shanshan Han

The capabilities of artificial intelligence systems have been advancing to a great extent, but these systems still struggle with failure modes, vulnerabilities, and biases. In this paper, we study the current state of the field, and present…

密码学与安全 · 计算机科学 2025-06-12 Xingli Fang , Jianwei Li , Varun Mulchandani , Jung-Eun Kim

Artificial intelligence (AI) is a technology which is increasingly being utilised in society and the economy worldwide, and its implementation is planned to become more prevalent in coming years. AI is increasingly being embedded in our…

计算机与社会 · 计算机科学 2019-07-10 Angela Daly , Thilo Hagendorff , Li Hui , Monique Mann , Vidushi Marda , Ben Wagner , Wei Wang , Saskia Witteborn

Current efforts in AI safety prioritize filtering harmful content, preventing manipulation of human behavior, and eliminating existential risks in cybersecurity or biosecurity. While pressing, this narrow focus overlooks critical…

计算机与社会 · 计算机科学 2025-07-14 Sanchaita Hazra , Bodhisattwa Prasad Majumder , Tuhin Chakrabarty

This position paper explores the broad landscape of AI potentiality in the context of cybersecurity, with a particular emphasis on its possible risk factors with awareness, which can be managed by incorporating human experts in the loop,…

密码学与安全 · 计算机科学 2023-10-20 Iqbal H. Sarker , Helge Janicke , Nazeeruddin Mohammad , Paul Watters , Surya Nepal

As AI systems proliferate in society, the AI community is increasingly preoccupied with the concept of AI Safety, namely the prevention of failures due to accidents that arise from an unanticipated departure of a system's behavior from…

计算机与社会 · 计算机科学 2024-01-23 Inioluwa Deborah Raji , Roel Dobbe

The rapid growth of Artificial Intelligence (AI) has underscored the urgent need for responsible AI practices. Despite increasing interest, a comprehensive AI risk assessment toolkit remains lacking. This study introduces our Responsible AI…

计算机与社会 · 计算机科学 2025-01-23 Sung Une Lee , Harsha Perera , Yue Liu , Boming Xia , Qinghua Lu , Liming Zhu , Olivier Salvado , Jon Whittle

Powerful artificial intelligence poses an existential threat if the AI decides to drastically change the world in pursuit of its goals. The hope of low-impact artificial intelligence is to incentivize AI to not do that just because this…

人工智能 · 计算机科学 2023-03-07 Danilo Naiff , Shashwat Goel

Artificial Intelligence (AI) is a double-edged sword: on one hand, AI promises to provide great advances that could benefit humanity, but on the other hand, AI poses substantial (even existential) risks. With advancements happening daily,…

计算机与社会 · 计算机科学 2024-02-05 Willem van der Maden , Derek Lomas , Malak Sadek , Paul Hekkert

As AI systems become increasingly powerful, the need for safe AI has become more pressing. Humans are an attractive model for AI safety: as the only known agents capable of general intelligence, they perform robustly even under conditions…

Societal cognitive overload, driven by the deluge of information and complexity in the AI age, poses a critical challenge to human well-being and societal resilience. This paper argues that mitigating cognitive overload is not only…

计算机与社会 · 计算机科学 2025-04-29 Salem Lahlou

The integration of Generative Artificial Intelligence (AI) into autonomous machines represents a major paradigm shift in how these systems operate and unlocks new solutions to problems once deemed intractable. Although generative AI agents…

机器人学 · 计算机科学 2024-10-22 Jason Jabbour , Vijay Janapa Reddi

With the advent of the digital era, every day-to-day task is automated due to technological advances. However, technology has yet to provide people with enough tools and safeguards. As the internet connects more-and-more devices around the…

密码学与安全 · 计算机科学 2022-09-28 Abhilash Chakraborty , Anupam Biswas , Ajoy Kumar Khan

This position paper argues that AI agents should be regulated by the extent to which they operate autonomously. AI agents with long-term planning and strategic capabilities can pose significant risks of human extinction and irreversible…

计算机与社会 · 计算机科学 2025-05-27 Takayuki Osogami

Risk thresholds provide a measure of the level of risk exposure that a society or individual is willing to withstand, ultimately shaping how we determine the safety of technological systems. Against the backdrop of the Cold War, the first…

计算机与社会 · 计算机科学 2025-04-22 Heidy Khlaaf , Sarah Myers West

AI Safety is an emerging area of critical importance to the safe adoption and deployment of AI systems. With the rapid proliferation of AI and especially with the recent advancement of Generative AI (or GAI), the technology ecosystem behind…

人工智能 · 计算机科学 2026-05-14 Chen Chen , Xueluan Gong , Ziyao Liu , Weifeng Jiang , Si Qi Goh , Kwok-Yan Lam

Artificial Intelligence (AI) has made impressive progress in recent years and represents a key technology that has a crucial impact on the economy and society. However, it is clear that AI and business models based on it can only reach…

This paper provides policy recommendations to reduce extinction risks from advanced artificial intelligence (AI). First, we briefly provide background information about extinction risks from AI. Second, we argue that voluntary commitments…

人工智能 · 计算机科学 2025-09-30 Andrea Miotti

Conversational AI systems can engage in unsafe behaviour when handling users' medical queries that can have severe consequences and could even lead to deaths. Systems therefore need to be capable of both recognising the seriousness of…

计算与语言 · 计算机科学 2022-10-04 Gavin Abercrombie , Verena Rieser

Artificial Intelligence (AI) is one of the most discussed technologies today. There are many innovative applications such as the diagnosis and treatment of cancer, customer experience, new business, education, contagious diseases…

计算机与社会 · 计算机科学 2020-01-28 Richard Benjamins , Idoia Salazar