English
Related papers

Related papers: R1-ACT: Efficient Reasoning Model Safety Alignment…

200 papers

The safety of large language models (LLMs) has increasingly emerged as a fundamental aspect of their development. Existing safety alignment for LLMs is predominantly achieved through post-training methods, which are computationally…

Artificial Intelligence · Computer Science 2026-02-03 Sicheng Shen , Mingyang Lv , Han Shen , Jialin Wu , Binghao Wang , Zhou Yang , Guobin Shen , Dongcheng Zhao , Feifei Zhao , Yi Zeng

Reasoning-based language models have demonstrated strong performance across various domains, with the most notable gains seen in mathematical and coding tasks. Recent research has shown that reasoning also offers significant benefits for…

Artificial Intelligence · Computer Science 2025-05-27 Makesh Narsimhan Sreedhar , Traian Rebedea , Christopher Parisien

Although Large Reasoning Models (LRMs) have progressed in solving complex problems, their chain-of-thought (CoT) reasoning often contains harmful content that can persist even when the final responses appear safe. We show that this issue…

Artificial Intelligence · Computer Science 2026-03-03 Yichi Zhang , Yue Ding , Jingwen Yang , Tianwei Luo , Dongbai Li , Ranjie Duan , Qiang Liu , Hang Su , Yinpeng Dong , Jun Zhu

Large language models (LLMs) are promising for autonomous driving, but semantics-only decision policies can yield physically unsafe behavior in dynamic traffic. Existing methods either perform online language reasoning without explicit…

Artificial Intelligence · Computer Science 2026-05-26 Zhengqi Sun , Yiwen Sun , Boxuan Liu , Tailai Chen , Tianxu Guo , Jiabin Liu

In recent years, general-purpose large language models (LLMs) such as GPT, Gemini, Claude, and DeepSeek have advanced at an unprecedented pace. Despite these achievements, their application to finance remains challenging, due to fragmented…

Reasoning large language models (RLLMs) have demonstrated outstanding performance across a variety of tasks, yet they also expose numerous security vulnerabilities. Most of these vulnerabilities have centered on the generation of unsafe…

Cryptography and Security · Computer Science 2025-05-13 Yu Cui , Cong Zuo

Recently, advanced large language models (LLMs) have emerged at an increasingly rapid pace. However, when faced with complex problems, most users are often unable to provide accurate and effective prompts to interact with LLMs, thus…

Computation and Language · Computer Science 2026-04-17 Wenjin Liu , Haoran Luo , Xueyuan Lin , Haoming Liu , Tiesunlong Shen , Jiapu Wang , Rui Mao , Erik Cambria

Recent advances in fine-tuning large language models (LLMs) with reinforcement learning (RL) have shown promising improvements in complex reasoning tasks, particularly when paired with chain-of-thought (CoT) prompting. However, these…

Machine Learning · Computer Science 2025-04-04 Hung Le , Dai Do , Dung Nguyen , Svetha Venkatesh

Large Language Models have shown impressive generative capabilities across diverse tasks, but their safety remains a critical concern. Existing post-training alignment methods, such as SFT and RLHF, reduce harmful outputs yet leave LLMs…

Cryptography and Security · Computer Science 2025-10-21 Zhengyue Zhao , Yingzi Ma , Somesh Jha , Marco Pavone , Patrick McDaniel , Chaowei Xiao

Chain-of-thought (CoT) reasoning has emerged as a powerful tool for multimodal large language models on video understanding tasks. However, its necessity and advantages over direct answering remain underexplored. In this paper, we first…

The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provide new avenues for improving the safety and robustness of our models. In particular, our…

Artificial Intelligence · Computer Science 2026-05-01 OpenAI , : , Aaron Jaech , Adam Kalai , Adam Lerer , Adam Richardson , Ahmed El-Kishky , Aiden Low , Alec Helyar , Aleksander Madry , Alex Beutel , Alex Carney , Alex Iftimie , Alex Karpenko , Alex Tachard Passos , Alexander Neitz , Alexander Prokofiev , Alexander Wei , Allison Tam , Ally Bennett , Ananya Kumar , Andre Saraiva , Andrea Vallone , Andrew Duberstein , Andrew Kondrich , Andrey Mishchenko , Andy Applebaum , Angela Jiang , Ashvin Nair , Barret Zoph , Behrooz Ghorbani , Bohan Zhang , Ben Rossen , Benjamin Sokolowsky , Boaz Barak , Bob McGrew , Borys Minaiev , Botao Hao , Bowen Baker , Brandon Houghton , Brandon McKinzie , Brydon Eastman , Camillo Lugaresi , Cary Bassin , Cary Hudson , Chak Ming Li , Charles de Bourcy , Chelsea Voss , Chen Shen , Chong Zhang , Chris Koch , Chris Orsinger , Christopher Hesse , Claudia Fischer , Clive Chan , Dan Roberts , Daniel Kappler , Daniel Levy , Daniel Selsam , David Dohan , David Farhi , David Mely , David Robinson , Dimitris Tsipras , Doug Li , Dragos Oprica , Eben Freeman , Eddie Zhang , Edmund Wong , Elizabeth Proehl , Enoch Cheung , Eric Mitchell , Eric Wallace , Erik Ritter , Evan Mays , Fan Wang , Felipe Petroski Such , Filippo Raso , Florencia Leoni , Foivos Tsimpourlas , Francis Song , Fred von Lohmann , Freddie Sulit , Geoff Salmon , Giambattista Parascandolo , Gildas Chabot , Grace Zhao , Greg Brockman , Guillaume Leclerc , Hadi Salman , Haiming Bao , Hao Sheng , Hart Andrin , Hessam Bagherinezhad , Hongyu Ren , Hunter Lightman , Hyung Won Chung , Ian Kivlichan , Ian O'Connell , Ian Osband , Ignasi Clavera Gilaberte , Ilge Akkaya , Ilya Kostrikov , Ilya Sutskever , Irina Kofman , Jakub Pachocki , James Lennon , Jason Wei , Jean Harb , Jerry Twore , Jiacheng Feng , Jiahui Yu , Jiayi Weng , Jie Tang , Jieqi Yu , Joaquin Quiñonero Candela , Joe Palermo , Joel Parish , Johannes Heidecke , John Hallman , John Rizzo , Jonathan Gordon , Jonathan Uesato , Jonathan Ward , Joost Huizinga , Julie Wang , Kai Chen , Kai Xiao , Karan Singhal , Karina Nguyen , Karl Cobbe , Katy Shi , Kayla Wood , Kendra Rimbach , Keren Gu-Lemberg , Kevin Liu , Kevin Lu , Kevin Stone , Kevin Yu , Lama Ahmad , Lauren Yang , Leo Liu , Leon Maksin , Leyton Ho , Liam Fedus , Lilian Weng , Linden Li , Lindsay McCallum , Lindsey Held , Lorenz Kuhn , Lukas Kondraciuk , Lukasz Kaiser , Luke Metz , Madelaine Boyd , Maja Trebacz , Manas Joglekar , Mark Chen , Marko Tintor , Mason Meyer , Matt Jones , Matt Kaufer , Max Schwarzer , Meghan Shah , Mehmet Yatbaz , Melody Y. Guan , Mengyuan Xu , Mengyuan Yan , Mia Glaese , Mianna Chen , Michael Lampe , Michael Malek , Michele Wang , Michelle Fradin , Mike McClay , Mikhail Pavlov , Miles Wang , Mingxuan Wang , Mira Murati , Mo Bavarian , Mostafa Rohaninejad , Nat McAleese , Neil Chowdhury , Neil Chowdhury , Nick Ryder , Nikolas Tezak , Noam Brown , Ofir Nachum , Oleg Boiko , Oleg Murk , Olivia Watkins , Patrick Chao , Paul Ashbourne , Pavel Izmailov , Peter Zhokhov , Rachel Dias , Rahul Arora , Randall Lin , Rapha Gontijo Lopes , Raz Gaon , Reah Miyara , Reimar Leike , Renny Hwang , Rhythm Garg , Robin Brown , Roshan James , Rui Shu , Ryan Cheu , Ryan Greene , Saachi Jain , Sam Altman , Sam Toizer , Sam Toyer , Samuel Miserendino , Sandhini Agarwal , Santiago Hernandez , Sasha Baker , Scott McKinney , Scottie Yan , Shengjia Zhao , Shengli Hu , Shibani Santurkar , Shraman Ray Chaudhuri , Shuyuan Zhang , Siyuan Fu , Spencer Papay , Steph Lin , Suchir Balaji , Suvansh Sanjeev , Szymon Sidor , Tal Broda , Aidan Clark , Tao Wang , Taylor Gordon , Ted Sanders , Tejal Patwardhan , Thibault Sottiaux , Thomas Degry , Thomas Dimson , Tianhao Zheng , Timur Garipov , Tom Stasi , Trapit Bansal , Trevor Creech , Troy Peterson , Tyna Eloundou , Valerie Qi , Vineet Kosaraju , Vinnie Monaco , Vitchyr Pong , Vlad Fomenko , Weiyi Zheng , Wenda Zhou , Wenting Zhan , Wes McCabe , Wojciech Zaremba , Yann Dubois , Yinghai Lu , Yining Chen , Young Cha , Yu Bai , Yuchen He , Yuchen Zhang , Yunyun Wang , Zheng Shao , Zhuohan Li

This paper presents AutoRAN, the first framework to automate the hijacking of internal safety reasoning in large reasoning models (LRMs). At its core, AutoRAN pioneers an execution simulation paradigm that leverages a weaker but…

Machine Learning · Computer Science 2026-04-17 Jiacheng Liang , Tanqiu Jiang , Yuhui Wang , Rongyi Zhu , Fenglong Ma , Ting Wang

As LLMs become increasingly prevalent across various applications, it is critical to establish safety guardrails to moderate input/output content of LLMs. Existing guardrail models treat various safety categories independently and fail to…

Artificial Intelligence · Computer Science 2024-07-09 Mintong Kang , Bo Li

Large Reasoning Models (LRMs) have emerged as a powerful advancement in multi-step reasoning tasks, offering enhanced transparency and logical consistency through explicit chains of thought (CoT). However, these models introduce novel…

Cryptography and Security · Computer Science 2026-04-15 Jiawei Chen , Yang Yang , Chao Yu , Yu Tian , Zhi Cao , Xue Yang , Linghao Li , Hang Su , Zhaoxia Yin

Large Reasoning Models (LRMs) achieve remarkable success through explicit thinking steps, yet the thinking steps introduce a novel risk by potentially amplifying unsafe behaviors. Despite this vulnerability, conventional defense mechanisms…

Artificial Intelligence · Computer Science 2026-01-08 Su-Hyeon Kim , Hyundong Jin , Yejin Lee , Yo-Sub Han

Large language models (LLMs) inevitably make mistakes when performing step-by-step mathematical reasoning. Process Reward Models (PRMs) have emerged as a promising solution by evaluating each reasoning step. However, existing PRMs typically…

Computation and Language · Computer Science 2025-03-28 Shuaijie She , Junxiao Liu , Yifeng Liu , Jiajun Chen , Xin Huang , Shujian Huang

Reasoning-enhanced large language models (RLLMs), whether explicitly trained for reasoning or prompted via chain-of-thought (CoT), have achieved state-of-the-art performance on many complex reasoning tasks. However, we uncover a surprising…

Computation and Language · Computer Science 2025-09-03 Xiaomin Li , Zhou Yu , Zhiwei Zhang , Xupeng Chen , Ziji Zhang , Yingying Zhuang , Narayanan Sadagopan , Anurag Beniwal

Large reasoning models (LRMs) increasingly expose chain-of-thought-like reasoning for transparency, verification, and deliberate problem solving. This creates a safety blind spot: harmful or policy-violating content may appear in reasoning…

Artificial Intelligence · Computer Science 2026-05-08 Xiaomin Li , Jianheng Hou , Zheyuan Deng , Zhiwei Zhang , Taoran Li , Binghang Lu , Bing Hu , Yunhan Zhao , Yuexing Hao

Large Reasoning Models like DeepSeek-R1 mark a fundamental shift in how LLMs approach complex problems. Instead of directly producing an answer for a given input, DeepSeek-R1 creates detailed multi-step reasoning chains, seemingly…

As large language models (LLMs) continue to advance in capabilities, ensuring their safety against jailbreak attacks remains a critical challenge. In this paper, we introduce a novel safety alignment approach called Answer-Then-Check, which…

Machine Learning · Computer Science 2026-03-09 Chentao Cao , Xiaojun Xu , Bo Han , Hang Li