English
Related papers

Related papers: Evaluating Language Models for Harmful Manipulatio…

200 papers

A core challenge in the development of increasingly capable AI systems is to make them safe and reliable by ensuring their behaviour is consistent with human values. This challenge, known as the alignment problem, does not merely apply to…

Machine Learning · Computer Science 2023-11-07 Raphaël Millière

AI systems are increasingly used in high-stakes domains such as credit rating, where fairness concerns are critical. Existing fairness assessments are typically conducted by AI experts or regulators using predefined protected attributes and…

Computers and Society · Computer Science 2026-02-10 Lin Luo , Satwik Ghanta , Yuri Nakao , Mathieu Chollet , Simone Stumpf

The complexity of psychological principles underscore a significant societal challenge, given the vast social implications of psychological problems. Bridging the gap between understanding these principles and their actual clinical and…

Artificial Intelligence · Computer Science 2023-12-11 Tianyu He , Guanghui Fu , Yijing Yu , Fan Wang , Jianqiang Li , Qing Zhao , Changwei Song , Hongzhi Qi , Dan Luo , Huijing Zou , Bing Xiang Yang

As large language models (LLMs) achieve advanced persuasive capabilities, concerns about their potential risks have grown. The EU AI Act prohibits AI systems that use manipulative or deceptive techniques to undermine informed…

Computers and Society · Computer Science 2025-05-20 Haein Kong

We develop new experimental paradigms for measuring welfare in language models. We compare verbal reports of models about their preferences with preferences expressed through behavior when navigating a virtual environment and selecting…

Artificial Intelligence · Computer Science 2026-05-26 Valen Tagliabue , Leonard Dung

The influence of Artificial Intelligence (AI), and specifically Large Language Models (LLM), on education is continuously increasing. These models are frequently used by students, giving rise to the question whether current forms of…

Human-Computer Interaction · Computer Science 2025-07-02 Patrick Stokkink

As conversational AI systems increasingly permeate the socio-emotional realms of human life, they bring both benefits and risks to individuals and society. Despite extensive research on detecting and categorizing harms in AI systems, less…

Human-Computer Interaction · Computer Science 2026-05-06 Renwen Zhang , Han Li , Han Meng , Jinyuan Zhan , Hongyuan Gan , Yi-Chieh Lee

We introduce a multi-turn benchmark for evaluating personalised alignment in LLM-based AI assistants, focusing on their ability to handle user-provided safety-critical contexts. Our assessment of ten leading models across five scenarios…

Human-Computer Interaction · Computer Science 2025-01-31 Lize Alberts , Benjamin Ellis , Andrei Lupu , Jakob Foerster

Lived experiences fundamentally shape how individuals interact with AI systems, influencing perceptions of safety, trust, and usability. While prior research has focused on developing techniques to emulate human preferences, and proposed…

Computers and Society · Computer Science 2025-08-12 Sanjana Gautam , Mohit Chandra , Ankolika De , Tatiana Chakravorti , Girik Malik , Munmun De Choudhury

Automated decision systems (ADS) are broadly deployed to inform and support human decision-making across a wide range of consequential settings. However, various context-specific details complicate the goal of establishing meaningful…

Computers and Society · Computer Science 2026-02-05 Inioluwa Deborah Raji , Lydia Liu

This paper examines the responsible integration of artificial intelligence (AI) in human services organizations (HSOs), proposing a nuanced framework for evaluating AI applications across multiple dimensions of risk. The authors argue that…

Computers and Society · Computer Science 2025-01-22 Brian E. Perron , Lauri Goldkind , Zia Qi , Bryan G. Victor

Large language models now possess human-level linguistic abilities in many contexts. This raises the concern that they can be used to deceive and manipulate on unprecedented scales, for instance spreading political misinformation on social…

Computers and Society · Computer Science 2026-01-21 Christian Tarsney

Artificial Intelligence (AI) systems are increasingly used in high-stakes domains of our life, increasing the need to explain these decisions and to make sure that they are aligned with how we want the decision to be made. The field of…

Artificial Intelligence · Computer Science 2023-06-28 Sofie Goethals , David Martens , Theodoros Evgeniou

As stories of human-AI interactions continue to be highlighted in the news and research platforms, the challenges are becoming more pronounced, including potential risks of overreliance, cognitive offloading, social and emotional…

Human-Computer Interaction · Computer Science 2025-10-22 Celeste Riley , Omar Al-Refai , Yadira Colunga Reyes , Eman Hammad

The rapid advancement of artificial intelligence (AI) technologies presents profound challenges to societal safety. As AI systems become more capable, accessible, and integrated into critical services, the dual nature of their potential is…

Artificial Intelligence · Computer Science 2024-12-06 Giulio Corsi , Kyle Kilian , Richard Mallah

This report surveys the landscape of potential security threats from malicious uses of AI, and proposes ways to better forecast, prevent, and mitigate these threats. After analyzing the ways in which AI may influence the threat landscape in…

Over a billion users globally interact with AI systems engineered to mimic human traits. This development raises concerns that anthropomorphism, the attribution of human characteristics to AI, may foster over-reliance and misplaced trust.…

Artificial Intelligence · Computer Science 2026-02-24 Robin Schimmelpfennig , Mark Díaz , Vinodkumar Prabhakaran , Aida Davani

This work explores the impact of moderation on users' enjoyment of conversational AI systems. While recent advancements in Large Language Models (LLMs) have led to highly capable conversational AIs that are increasingly deployed in…

Human-Computer Interaction · Computer Science 2023-04-21 Xiaoding Lu , Aleksey Korshuk , Zongyi Liu , William Beauchamp , Chai Research

A key task in AI practice is to assess potential impacts to prevent harm. Current AI tools assisting AI impact assessment have not been designed or evaluated for collaborative team brainstorming, and they do not capture the range of views…

Human-Computer Interaction · Computer Science 2026-05-01 Jarod Govers , Sanja Šćepanović , Daniele Quercia

To evaluate the societal impacts of GenAI requires a model of how social harms emerge from interactions between GenAI, people, and societal structures. Yet a model is rarely explicitly defined in societal impact evaluations, or in the…

Human-Computer Interaction · Computer Science 2024-10-31 Glen Berman , Ned Cooper , Wesley Hanwen Deng , Ben Hutchinson