English
Related papers

Related papers: Don't Command, Cultivate: An Exploratory Study of …

200 papers

This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers most questions, a deeper reasoning model for harder problems, and a real-time router that…

Computation and Language · Computer Science 2026-05-05 Aaditya Singh , Adam Fry , Adam Perelman , Adam Tart , Adi Ganesh , Ahmed El-Kishky , Aidan McLaughlin , Aiden Low , AJ Ostrow , Akhila Ananthram , Akshay Nathan , Alan Luo , Alec Helyar , Aleksander Madry , Aleksandr Efremov , Aleksandra Spyra , Alex Baker-Whitcomb , Alex Beutel , Alex Karpenko , Alex Makelov , Alex Neitz , Alex Wei , Alexandra Barr , Alexandre Kirchmeyer , Alexey Ivanov , Alexi Christakis , Alistair Gillespie , Allison Tam , Ally Bennett , Alvin Wan , Alyssa Huang , Amy McDonald Sandjideh , Amy Yang , Ananya Kumar , Andre Saraiva , Andrea Vallone , Andrei Gheorghe , Andres Garcia Garcia , Andrew Braunstein , Andrew Liu , Andrew Schmidt , Andrey Mereskin , Andrey Mishchenko , Andy Applebaum , Andy Rogerson , Ann Rajan , Annie Wei , Anoop Kotha , Anubha Srivastava , Anushree Agrawal , Arun Vijayvergiya , Ashley Tyra , Ashvin Nair , Avi Nayak , Ben Eggers , Bessie Ji , Beth Hoover , Bill Chen , Blair Chen , Boaz Barak , Borys Minaiev , Botao Hao , Bowen Baker , Brad Lightcap , Brandon McKinzie , Brandon Wang , Brendan Quinn , Brian Fioca , Brian Hsu , Brian Yang , Brian Yu , Brian Zhang , Brittany Brenner , Callie Riggins Zetino , Cameron Raymond , Camillo Lugaresi , Carolina Paz , Cary Hudson , Cedric Whitney , Chak Li , Charles Chen , Charlotte Cole , Chelsea Voss , Chen Ding , Chen Shen , Chengdu Huang , Chris Colby , Chris Hallacy , Chris Koch , Chris Lu , Christina Kaplan , Christina Kim , CJ Minott-Henriques , Cliff Frey , Cody Yu , Coley Czarnecki , Colin Reid , Colin Wei , Cory Decareaux , Cristina Scheau , Cyril Zhang , Cyrus Forbes , Da Tang , Dakota Goldberg , Dan Roberts , Dana Palmie , Daniel Kappler , Daniel Levine , Daniel Wright , Dave Leo , David Lin , David Robinson , Declan Grabb , Derek Chen , Derek Lim , Derek Salama , Dibya Bhattacharjee , Dimitris Tsipras , Dinghua Li , Dingli Yu , DJ Strouse , Drew Williams , Dylan Hunn , Ed Bayes , Edwin Arbus , Ekin Akyurek , Elaine Ya Le , Elana Widmann , Eli Yani , Elizabeth Proehl , Enis Sert , Enoch Cheung , Eri Schwartz , Eric Han , Eric Jiang , Eric Mitchell , Eric Sigler , Eric Wallace , Erik Ritter , Erin Kavanaugh , Evan Mays , Evgenii Nikishin , Fangyuan Li , Felipe Petroski Such , Filipe de Avila Belbute Peres , Filippo Raso , Florent Bekerman , Foivos Tsimpourlas , Fotis Chantzis , Francis Song , Francis Zhang , Gaby Raila , Garrett McGrath , Gary Briggs , Gary Yang , Giambattista Parascandolo , Gildas Chabot , Grace Kim , Grace Zhao , Gregory Valiant , Guillaume Leclerc , Hadi Salman , Hanson Wang , Hao Sheng , Haoming Jiang , Haoyu Wang , Haozhun Jin , Harshit Sikchi , Heather Schmidt , Henry Aspegren , Honglin Chen , Huida Qiu , Hunter Lightman , Ian Covert , Ian Kivlichan , Ian Silber , Ian Sohl , Ibrahim Hammoud , Ignasi Clavera , Ikai Lan , Ilge Akkaya , Ilya Kostrikov , Irina Kofman , Isak Etinger , Ishaan Singal , Jackie Hehir , Jacob Huh , Jacqueline Pan , Jake Wilczynski , Jakub Pachocki , James Lee , James Quinn , Jamie Kiros , Janvi Kalra , Jasmyn Samaroo , Jason Wang , Jason Wolfe , Jay Chen , Jay Wang , Jean Harb , Jeffrey Han , Jeffrey Wang , Jennifer Zhao , Jeremy Chen , Jerene Yang , Jerry Tworek , Jesse Chand , Jessica Landon , Jessica Liang , Ji Lin , Jiancheng Liu , Jianfeng Wang , Jie Tang , Jihan Yin , Joanne Jang , Joel Morris , Joey Flynn , Johannes Ferstad , Johannes Heidecke , John Fishbein , John Hallman , Jonah Grant , Jonathan Chien , Jonathan Gordon , Jongsoo Park , Jordan Liss , Jos Kraaijeveld , Joseph Guay , Joseph Mo , Josh Lawson , Josh McGrath , Joshua Vendrow , Joy Jiao , Julian Lee , Julie Steele , Julie Wang , Junhua Mao , Kai Chen , Kai Hayashi , Kai Xiao , Kamyar Salahi , Kan Wu , Karan Sekhri , Karan Sharma , Karan Singhal , Karen Li , Kenny Nguyen , Keren Gu-Lemberg , Kevin King , Kevin Liu , Kevin Stone , Kevin Yu , Kristen Ying , Kristian Georgiev , Kristie Lim , Kushal Tirumala , Kyle Miller , Lama Ahmad , Larry Lv , Laura Clare , Laurance Fauconnet , Lauren Itow , Lauren Yang , Laurentia Romaniuk , Leah Anise , Lee Byron , Leher Pathak , Leon Maksin , Leyan Lo , Leyton Ho , Li Jing , Liang Wu , Liang Xiong , Lien Mamitsuka , Lin Yang , Lindsay McCallum , Lindsey Held , Liz Bourgeois , Logan Engstrom , Lorenz Kuhn , Louis Feuvrier , Lu Zhang , Lucas Switzer , Lukas Kondraciuk , Lukasz Kaiser , Manas Joglekar , Mandeep Singh , Mandip Shah , Manuka Stratta , Marcus Williams , Mark Chen , Mark Sun , Marselus Cayton , Martin Li , Marvin Zhang , Marwan Aljubeh , Matt Nichols , Matthew Haines , Max Schwarzer , Mayank Gupta , Meghan Shah , Melody Y. Guan , Melody Huang , Meng Dong , Mengqing Wang , Mia Glaese , Micah Carroll , Michael Lampe , Michael Malek , Michael Sharman , Michael Zhang , Michele Wang , Michelle Pokrass , Mihai Florian , Mikhail Pavlov , Miles Wang , Ming Chen , Mingxuan Wang , Minnia Feng , Mo Bavarian , Molly Lin , Moose Abdool , Mostafa Rohaninejad , Nacho Soto , Natalie Staudacher , Natan LaFontaine , Nathan Marwell , Nelson Liu , Nick Preston , Nick Turley , Nicklas Ansman , Nicole Blades , Nikil Pancha , Nikita Mikhaylin , Niko Felix , Nikunj Handa , Nishant Rai , Nitish Keskar , Noam Brown , Ofir Nachum , Oleg Boiko , Oleg Murk , Olivia Watkins , Oona Gleeson , Pamela Mishkin , Patryk Lesiewicz , Paul Baltescu , Pavel Belov , Peter Zhokhov , Philip Pronin , Phillip Guo , Phoebe Thacker , Qi Liu , Qiming Yuan , Qinghua Liu , Rachel Dias , Rachel Puckett , Rahul Arora , Ravi Teja Mullapudi , Raz Gaon , Reah Miyara , Rennie Song , Rishabh Aggarwal , RJ Marsan , Robel Yemiru , Robert Xiong , Rohan Kshirsagar , Rohan Nuttall , Roman Tsiupa , Ronen Eldan , Rose Wang , Roshan James , Roy Ziv , Rui Shu , Ruslan Nigmatullin , Saachi Jain , Saam Talaie , Sam Altman , Sam Arnesen , Sam Toizer , Sam Toyer , Samuel Miserendino , Sandhini Agarwal , Sarah Yoo , Savannah Heon , Scott Ethersmith , Sean Grove , Sean Taylor , Sebastien Bubeck , Sever Banesiu , Shaokyi Amdo , Shengjia Zhao , Sherwin Wu , Shibani Santurkar , Shiyu Zhao , Shraman Ray Chaudhuri , Shreyas Krishnaswamy , Shuaiqi , Xia , Shuyang Cheng , Shyamal Anadkat , Simón Posada Fishman , Simon Tobin , Siyuan Fu , Somay Jain , Song Mei , Sonya Egoian , Spencer Kim , Spug Golden , SQ Mah , Steph Lin , Stephen Imm , Steve Sharpe , Steve Yadlowsky , Sulman Choudhry , Sungwon Eum , Suvansh Sanjeev , Tabarak Khan , Tal Stramer , Tao Wang , Tao Xin , Tarun Gogineni , Taya Christianson , Ted Sanders , Tejal Patwardhan , Thomas Degry , Thomas Shadwell , Tianfu Fu , Tianshi Gao , Timur Garipov , Tina Sriskandarajah , Toki Sherbakov , Tomek Korbak , Tomer Kaftan , Tomo Hiratsuka , Tongzhou Wang , Tony Song , Tony Zhao , Troy Peterson , Val Kharitonov , Victoria Chernova , Vineet Kosaraju , Vishal Kuo , Vitchyr Pong , Vivek Verma , Vlad Petrov , Wanning Jiang , Weixing Zhang , Wenda Zhou , Wenlei Xie , Wenting Zhan , Wes McCabe , Will DePue , Will Ellsworth , Wulfie Bain , Wyatt Thompson , Xiangning Chen , Xiangyu Qi , Xin Xiang , Xinwei Shi , Yann Dubois , Yaodong Yu , Yara Khakbaz , Yifan Wu , Yilei Qian , Yin Tat Lee , Yinbo Chen , Yizhen Zhang , Yizhong Xiong , Yonglong Tian , Young Cha , Yu Bai , Yu Yang , Yuan Yuan , Yuanzhi Li , Yufeng Zhang , Yuguang Yang , Yujia Jin , Yun Jiang , Yunyun Wang , Yushi Wang , Yutian Liu , Zach Stubenvoll , Zehao Dou , Zheng Wu , Zhigang Wang

Recent advances in Large Language Models (LLMs) have led to impressive alignment where models learn to distinguish harmful from harmless queries through supervised finetuning (SFT) and reinforcement learning from human feedback (RLHF). In…

Artificial Intelligence · Computer Science 2025-06-18 Jiahao Yu , Haozheng Luo , Jerry Yao-Chieh Hu , Wenbo Guo , Han Liu , Xinyu Xing

The governance of open-weight artificial intelligence (AI) models has been framed as a binary choice: openness as risk, restriction as safety. This paper challenges that framing, arguing that access restrictions, without governed…

Computers and Society · Computer Science 2026-04-21 Vinicius Santana Gomes

Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in reasoning and generation tasks and are increasingly deployed in real-world applications. However, their explicit chain-of-thought (CoT) mechanism introduces new…

Artificial Intelligence · Computer Science 2026-05-26 Jianan Li , Simeng Qin , Xiaojun Jia , Lionel Z. Wang , Tianhang Zheng , Xiaoshuang Jia , Yang Liu , Xiaochun Cao

Language models are trained to follow instructions, but they are also powerful pattern completers. What happens when these two objectives conflict? We construct conversations in which a user instruction to behave in a target way T (e.g.,…

Computation and Language · Computer Science 2026-05-21 Carolina Camassa , Derek Shiller

Chain-of-Thought (CoT) monitoring, in which automated systems monitor the CoT of an LLM, is a promising approach for effectively overseeing AI systems. However, the extent to which a model's CoT helps us oversee the model - the…

Machine Learning · Computer Science 2026-04-01 Max Kaufmann , David Lindner , Roland S. Zimmermann , and Rohin Shah

Leading language model (LM) providers like OpenAI and Anthropic allow customers to fine-tune frontier LMs for specific use cases. To prevent abuse, these providers apply filters to block fine-tuning on overtly harmful data. In this setting,…

Cryptography and Security · Computer Science 2025-07-15 Joshua Kazdan , Abhay Puri , Rylan Schaeffer , Lisa Yu , Chris Cundy , Jason Stanley , Sanmi Koyejo , Krishnamurthy Dvijotham

Recent work has developed optimization procedures to find token sequences, called adversarial triggers, which can elicit unsafe responses from aligned language models. These triggers are believed to be highly transferable, i.e., a trigger…

Computation and Language · Computer Science 2025-04-10 Nicholas Meade , Arkil Patel , Siva Reddy

We propose CRAFT, a red-teaming alignment framework that leverages model reasoning capabilities and hidden representations to improve robustness against jailbreak attacks. Unlike prior defenses that operate primarily at the output level,…

Artificial Intelligence · Computer Science 2026-05-20 Haozheng Luo , Yimin Wang , Jiahao Yu , Binghui Wang , Yan Chen

The challenge of ensuring Large Language Models (LLMs) align with societal standards is of increasing interest, as these models are still prone to adversarial jailbreaks that bypass their safety mechanisms. Identifying these vulnerabilities…

Computation and Language · Computer Science 2025-04-29 Mohammad Akbar-Tajari , Mohammad Taher Pilehvar , Mohammad Mahmoody

Large Language Models (LLMs) are known to be susceptible to crafted adversarial attacks or jailbreaks that lead to the generation of objectionable content despite being aligned to human preferences using safety fine-tuning methods. While…

Computation and Language · Computer Science 2025-03-26 Sravanti Addepalli , Yerram Varun , Arun Suggala , Karthikeyan Shanmugam , Prateek Jain

Text-to-image (T2I) models commonly incorporate defense mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreak attacks have shown that adversarial prompts can effectively bypass these mechanisms and induce…

Cryptography and Security · Computer Science 2026-03-25 Chenyu Zhang , Lanjun Wang , Yiwen Ma , Wenhui Li , Yi Tu , An-An Liu

Should LLM reasoning live in a separate module, or within a single model's forward pass and representational space? We study dual-architecture latent reasoning, where a fluent Base exchanges latent messages with a Coprocessor, and test two…

As artificial intelligence (AI) continues to advance, it demonstrates capabilities comparable to human intelligence, with significant potential to transform education and workforce development. This study evaluates OpenAI o1-preview's…

AI agents powered by reasoning models require access to sensitive user data. However, their reasoning traces are difficult to control, which can result in the unintended leakage of private information to external parties. We propose…

Computation and Language · Computer Science 2026-03-02 Haritz Puerto , Haonan Li , Xudong Han , Timothy Baldwin , Iryna Gurevych

System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensitive operational context in agentic AI applications. These instructions may contain sensitive…

Cryptography and Security · Computer Science 2026-04-02 Anubhab Sahu , Diptisha Samanta , Reza Soosahabi

Safety cases, structured arguments that a system is acceptably safe, are becoming central to the governance of AI systems. Yet, traditional safety-case practices from aviation or nuclear engineering rely on well-specified system boundaries,…

Software Engineering · Computer Science 2026-03-09 Sung Une Lee , Liming Zhu , Md Shamsujjoha , Liming Dong , Qinghua Lu , Jieshan Chen , Lionel Briand

Reasoning Language Models (RLMs) have gained traction for their ability to perform complex, multi-step reasoning tasks through mechanisms such as Chain-of-Thought (CoT) prompting or fine-tuned reasoning traces. While these capabilities…

Computation and Language · Computer Science 2025-07-04 Riccardo Cantini , Nicola Gabriele , Alessio Orsino , Domenico Talia

Chain-of-Thought (CoT) prompting has been used to enhance the reasoning capability of LLMs. However, its reliability in security-sensitive analytical tasks remains insufficiently examined, particularly under structured human evaluation.…

Cryptography and Security · Computer Science 2026-04-07 Jiling Zhou , Aisvarya Adeseye , Seppo Virtanen , Antti Hakkala , Jouni Isoaho

Large language models have demonstrated remarkable progress in mathematical reasoning, leveraging chain-of-thought and test-time compute scaling. However, many open questions remain regarding the interplay between reasoning token usage and…

Machine Learning · Computer Science 2025-02-24 Marthe Ballon , Andres Algaba , Vincent Ginis