中文
相关论文

相关论文: From Hard Refusals to Safe-Completions: Toward Out…

200 篇论文

Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representations. Existing methods do so by ablating a refusal direction from model activations, aiming…

人工智能 · 计算机科学 2026-05-22 Giorgio Piras , Raffaele Mura , Fabio Brau , Maura Pintor , Luca Oneto , Fabio Roli , Battista Biggio

SafePredict is a novel meta-algorithm that works with any base prediction algorithm for online data to guarantee an arbitrarily chosen correctness rate, $1-\epsilon$, by allowing refusals. Allowing refusals means that the meta-algorithm may…

机器学习 · 计算机科学 2017-11-10 Mustafa A. Kocak , David Ramirez , Elza Erkip , Dennis E. Shasha

While ChatGPT may help students to learn to program, it can be misused to do plagiarism, a breach of academic integrity. Students can ask ChatGPT to complete a programming task, generating a solution from other people's work without proper…

Large Language Models (LLMs) require careful safety alignment to prevent malicious outputs. While significant research focuses on mitigating harmful content generation, the enhanced safety often come with the side effect of over-refusal,…

计算与语言 · 计算机科学 2025-06-17 Justin Cui , Wei-Lin Chiang , Ion Stoica , Cho-Jui Hsieh

One of the major drawbacks of modularized task-completion dialogue systems is that each module is trained individually, which presents several challenges. For example, downstream modules are affected by earlier modules, and the performance…

计算与语言 · 计算机科学 2018-02-13 Xiujun Li , Yun-Nung Chen , Lihong Li , Jianfeng Gao , Asli Celikyilmaz

The strive to make AI applications "safe" has led to the development of safety-measures as the main or even sole normative requirement of their permissible use. Similar can be attested to the latest version of chatbots, such as chatGPT. In…

人工智能 · 计算机科学 2023-05-01 Hendrik Kempt , Alon Lavie , Saskia K. Nagel

We identify a structural weakness in current large language model (LLM) alignment: modern refusal mechanisms are fail-open. While existing approaches encode refusal behaviors across multiple latent features, suppressing a single dominant…

机器学习 · 计算机科学 2026-02-20 Zachary Coalson , Beth Sohler , Aiden Gabriel , Sanghyun Hong

In conversational search, agents can interact with users by asking clarifying questions to increase their chance to find better results. Many recent works and shared tasks in both NLP and IR communities have focused on identifying the need…

信息检索 · 计算机科学 2022-01-04 Zhenduo Wang , Qingyao Ai

While large neural-based conversational models have become increasingly proficient dialogue agents, recent work has highlighted safety issues with these systems. For example, these systems can be goaded into generating toxic content, which…

计算与语言 · 计算机科学 2023-10-24 Nicholas Meade , Spandana Gella , Devamanyu Hazarika , Prakhar Gupta , Di Jin , Siva Reddy , Yang Liu , Dilek Hakkani-Tür

Recent advancements in large language models, such as ChatGPT, have demonstrated significant potential to impact various aspects of human life. However, ChatGPT still faces challenges in providing reliable and accurate answers to user…

计算与语言 · 计算机科学 2023-12-05 Shen Zheng , Jie Huang , Kevin Chen-Chuan Chang

Patients with schizophrenia often present with cognitive impairments that may hinder their ability to learn about their condition. These individuals could benefit greatly from education platforms that leverage the adaptability of Large…

计算与语言 · 计算机科学 2024-10-18 Per Niklas Waaler , Musarrat Hussain , Igor Molchanov , Lars Ailo Bongo , Brita Elvevåg

Safety post-training has been studied extensively in single-step "chat" settings where safety typically refers to refusing harmful requests. We study an "agentic" (i.e., multi-step, tool-use) setting where safety refers to harmful actions…

机器学习 · 计算机科学 2026-03-04 Benjamin Plaut

Correctly classifying adversarial examples is an essential but challenging requirement for safely deploying machine learning models. As reported in RobustBench, even the state-of-the-art adversarially trained models struggle to exceed 67%…

机器学习 · 计算机科学 2022-04-01 Tianyu Pang , Huishuai Zhang , Di He , Yinpeng Dong , Hang Su , Wei Chen , Jun Zhu , Tie-Yan Liu

Instruction-following language models are trained to be helpful and safe, yet their safety behavior can deteriorate under benign fine-tuning and worsen under adversarial updates. Existing defenses often offer limited protection or force a…

计算与语言 · 计算机科学 2026-05-12 Jyotin Goel , Souvik Maji , Pratik Mazumder

As the influence of large language models (LLMs) spans across global communities, their safety challenges in multilingual settings become paramount for alignment research. This paper examines the variations in safety challenges faced by…

计算与语言 · 计算机科学 2024-01-25 Lingfeng Shen , Weiting Tan , Sihao Chen , Yunmo Chen , Jingyu Zhang , Haoran Xu , Boyuan Zheng , Philipp Koehn , Daniel Khashabi

Refusal training is widely used to prevent LLMs from generating harmful, undesirable, or illegal outputs. We reveal a curious generalization gap in the current refusal training approaches: simply reformulating a harmful request in the past…

计算与语言 · 计算机科学 2025-04-21 Maksym Andriushchenko , Nicolas Flammarion

AI leaders and safety reports increasingly warn that advances in model reasoning may enable biological misuse, including by low-expertise users, while major labs describe safeguards as expanding but still evolving rather than settled. This…

计算机与社会 · 计算机科学 2026-04-24 Michael Richter

This paper studies recent developments in large language models' (LLM) abilities to pass assessments in introductory and intermediate Python programming courses at the postsecondary level. The emergence of ChatGPT resulted in heated debates…

计算机与社会 · 计算机科学 2023-10-05 Jaromir Savelka , Arav Agarwal , Marshall An , Chris Bogart , Majd Sakr

Safety is one of the main challenges in applying reinforcement learning to realistic environmental tasks. To ensure safety during and after training process, existing methods tend to adopt overly conservative policy to avoid unsafe…

机器学习 · 计算机科学 2023-06-27 Xiao Zhang , Hai Zhang , Hongtu Zhou , Chang Huang , Di Zhang , Chen Ye , Junqiao Zhao

AI assistants are being increasingly used by students enrolled in higher education institutions. While these tools provide opportunities for improved teaching and education, they also pose significant challenges for assessment and learning…

计算机与社会 · 计算机科学 2024-11-28 Beatriz Borges , Negar Foroutan , Deniz Bayazit , Anna Sotnikova , Syrielle Montariol , Tanya Nazaretzky , Mohammadreza Banaei , Alireza Sakhaeirad , Philippe Servant , Seyed Parsa Neshaei , Jibril Frej , Angelika Romanou , Gail Weiss , Sepideh Mamooler , Zeming Chen , Simin Fan , Silin Gao , Mete Ismayilzada , Debjit Paul , Alexandre Schöpfer , Andrej Janchevski , Anja Tiede , Clarence Linden , Emanuele Troiani , Francesco Salvi , Freya Behrens , Giacomo Orsi , Giovanni Piccioli , Hadrien Sevel , Louis Coulon , Manuela Pineros-Rodriguez , Marin Bonnassies , Pierre Hellich , Puck van Gerwen , Sankalp Gambhir , Solal Pirelli , Thomas Blanchard , Timothée Callens , Toni Abi Aoun , Yannick Calvino Alonso , Yuri Cho , Alberto Chiappa , Antonio Sclocchi , Étienne Bruno , Florian Hofhammer , Gabriel Pescia , Geovani Rizk , Leello Dadi , Lucas Stoffl , Manoel Horta Ribeiro , Matthieu Bovel , Yueyang Pan , Aleksandra Radenovic , Alexandre Alahi , Alexander Mathis , Anne-Florence Bitbol , Boi Faltings , Cécile Hébert , Devis Tuia , François Maréchal , George Candea , Giuseppe Carleo , Jean-Cédric Chappelier , Nicolas Flammarion , Jean-Marie Fürbringer , Jean-Philippe Pellet , Karl Aberer , Lenka Zdeborová , Marcel Salathé , Martin Jaggi , Martin Rajman , Mathias Payer , Matthieu Wyart , Michael Gastpar , Michele Ceriotti , Ola Svensson , Olivier Lévêque , Paolo Ienne , Rachid Guerraoui , Robert West , Sanidhya Kashyap , Valerio Piazza , Viesturs Simanis , Viktor Kuncak , Volkan Cevher , Philippe Schwaller , Sacha Friedli , Patrick Jermann , Tanja Käser , Antoine Bosselut