English
Related papers

Related papers: Stress Testing Deliberative Alignment for Anti-Sch…

200 papers

While the wide adoption of refusal training in large language models (LLMs) has showcased improvements in model safety, recent works have highlighted shortcomings due to the shallow nature of these alignment methods. To this end, the work…

Machine Learning · Computer Science 2026-04-17 Pankayaraj Pathmanathan , Furong Huang

As artificial intelligence (AI) improves, traditional alignment strategies may falter in the face of unpredictable self-improvement, hidden subgoals, and the sheer complexity of intelligent systems. Inspired by contemplative wisdom…

Artificial Intelligence · Computer Science 2025-08-19 Ruben Laukkonen , Fionn Inglis , Shamil Chandaria , Lars Sandved-Smith , Edmundo Lopez-Sola , Jakob Hohwy , Jonathan Gold , Adam Elwood

Large language models (LLMs) have demonstrated remarkable capabilities in tasks requiring reasoning and multi-step problem-solving through the use of chain-of-thought (CoT) prompting. However, generating the full CoT process results in…

Computation and Language · Computer Science 2024-09-16 Tianqiao Liu , Zui Chen , Zitao Liu , Mi Tian , Weiqi Luo

Despite recent advancements, NLP models continue to be vulnerable to bias. This bias often originates from the uneven distribution of real-world data and can propagate through the annotation process. Escalated integration of these models in…

Computation and Language · Computer Science 2023-05-29 Sabit Hassan , Malihe Alikhani

Zero-shot Chain-of-Thought (CoT) prompting emerges as a simple and effective strategy for enhancing the performance of large language models (LLMs) in real-world reasoning tasks. Nonetheless, the efficacy of a singular, task-level prompt…

Computation and Language · Computer Science 2024-11-01 Xiaosong Yuan , Chen Shen , Shaotian Yan , Xiaofeng Zhang , Liang Xie , Wenxiao Wang , Renchu Guan , Ying Wang , Jieping Ye

High false-positive rate is a long-standing challenge for anomaly detection algorithms, especially in high-stake applications. To identify the true anomalies, in practice, analysts or domain experts will be employed to investigate the top…

Machine Learning · Computer Science 2020-09-17 Daochen Zha , Kwei-Herng Lai , Mingyang Wan , Xia Hu

Recent advances in deep learning have brought attention to the possibility of creating advanced, general AI systems that outperform humans across many tasks. However, if these systems pursue unintended goals, there could be catastrophic…

Machine Learning · Computer Science 2024-11-25 Dylan Xu , Juan-Pablo Rivera

Aiming at efficient and dense chain-of-thought (CoT) reasoning, latent reasoning methods fine-tune Large Language Models (LLMs) to substitute discrete language tokens with continuous latent tokens. These methods consume fewer tokens…

Artificial Intelligence · Computer Science 2026-01-30 Zhi Zheng , Wee Sun Lee

Structured deliberation has been found to improve the performance of human forecasters. This study investigates whether a similar intervention, i.e. allowing LLMs to review each other's forecasts before updating, can improve accuracy in…

Artificial Intelligence · Computer Science 2025-12-30 Paul Schneider , Amalie Schramm

In complex industrial and chemical process control rooms, effective decision-making is crucial for safety and efficiency. The experiments in this paper evaluate the impact and applications of an AI-based decision support system integrated…

Chain-of-thought (CoT) reasoning improves large language models (LLMs) on difficult tasks, but it also makes inference expensive because every intermediate step must be generated as a discrete token. Latent reasoning reduces visible token…

Computation and Language · Computer Science 2026-05-11 Xuan Li , Yining Wang , Yuchen Liu , Guanjun Liu , Delai Qiu , Shengping Liu , Jiaen Liang , Wei Huang , Jun Yu , Junnan Zhu

As large-scale language models increasingly impact safety-critical domains, ensuring their reliable adherence to well-defined principles remains a fundamental challenge. We introduce Deliberative Alignment, a new paradigm that directly…

We develop and study new adversarial perturbations that enable an attacker to gain control over decisions in generic Artificial Intelligence (AI) systems including deep learning neural networks. In contrast to adversarial data modification,…

Cryptography and Security · Computer Science 2023-12-07 Ivan Y. Tyukin , Desmond J. Higham , Alexander Bastounis , Eliyas Woldegeorgis , Alexander N. Gorban

The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provide new avenues for improving the safety and robustness of our models. In particular, our…

Artificial Intelligence · Computer Science 2026-05-01 OpenAI , : , Aaron Jaech , Adam Kalai , Adam Lerer , Adam Richardson , Ahmed El-Kishky , Aiden Low , Alec Helyar , Aleksander Madry , Alex Beutel , Alex Carney , Alex Iftimie , Alex Karpenko , Alex Tachard Passos , Alexander Neitz , Alexander Prokofiev , Alexander Wei , Allison Tam , Ally Bennett , Ananya Kumar , Andre Saraiva , Andrea Vallone , Andrew Duberstein , Andrew Kondrich , Andrey Mishchenko , Andy Applebaum , Angela Jiang , Ashvin Nair , Barret Zoph , Behrooz Ghorbani , Bohan Zhang , Ben Rossen , Benjamin Sokolowsky , Boaz Barak , Bob McGrew , Borys Minaiev , Botao Hao , Bowen Baker , Brandon Houghton , Brandon McKinzie , Brydon Eastman , Camillo Lugaresi , Cary Bassin , Cary Hudson , Chak Ming Li , Charles de Bourcy , Chelsea Voss , Chen Shen , Chong Zhang , Chris Koch , Chris Orsinger , Christopher Hesse , Claudia Fischer , Clive Chan , Dan Roberts , Daniel Kappler , Daniel Levy , Daniel Selsam , David Dohan , David Farhi , David Mely , David Robinson , Dimitris Tsipras , Doug Li , Dragos Oprica , Eben Freeman , Eddie Zhang , Edmund Wong , Elizabeth Proehl , Enoch Cheung , Eric Mitchell , Eric Wallace , Erik Ritter , Evan Mays , Fan Wang , Felipe Petroski Such , Filippo Raso , Florencia Leoni , Foivos Tsimpourlas , Francis Song , Fred von Lohmann , Freddie Sulit , Geoff Salmon , Giambattista Parascandolo , Gildas Chabot , Grace Zhao , Greg Brockman , Guillaume Leclerc , Hadi Salman , Haiming Bao , Hao Sheng , Hart Andrin , Hessam Bagherinezhad , Hongyu Ren , Hunter Lightman , Hyung Won Chung , Ian Kivlichan , Ian O'Connell , Ian Osband , Ignasi Clavera Gilaberte , Ilge Akkaya , Ilya Kostrikov , Ilya Sutskever , Irina Kofman , Jakub Pachocki , James Lennon , Jason Wei , Jean Harb , Jerry Twore , Jiacheng Feng , Jiahui Yu , Jiayi Weng , Jie Tang , Jieqi Yu , Joaquin Quiñonero Candela , Joe Palermo , Joel Parish , Johannes Heidecke , John Hallman , John Rizzo , Jonathan Gordon , Jonathan Uesato , Jonathan Ward , Joost Huizinga , Julie Wang , Kai Chen , Kai Xiao , Karan Singhal , Karina Nguyen , Karl Cobbe , Katy Shi , Kayla Wood , Kendra Rimbach , Keren Gu-Lemberg , Kevin Liu , Kevin Lu , Kevin Stone , Kevin Yu , Lama Ahmad , Lauren Yang , Leo Liu , Leon Maksin , Leyton Ho , Liam Fedus , Lilian Weng , Linden Li , Lindsay McCallum , Lindsey Held , Lorenz Kuhn , Lukas Kondraciuk , Lukasz Kaiser , Luke Metz , Madelaine Boyd , Maja Trebacz , Manas Joglekar , Mark Chen , Marko Tintor , Mason Meyer , Matt Jones , Matt Kaufer , Max Schwarzer , Meghan Shah , Mehmet Yatbaz , Melody Y. Guan , Mengyuan Xu , Mengyuan Yan , Mia Glaese , Mianna Chen , Michael Lampe , Michael Malek , Michele Wang , Michelle Fradin , Mike McClay , Mikhail Pavlov , Miles Wang , Mingxuan Wang , Mira Murati , Mo Bavarian , Mostafa Rohaninejad , Nat McAleese , Neil Chowdhury , Neil Chowdhury , Nick Ryder , Nikolas Tezak , Noam Brown , Ofir Nachum , Oleg Boiko , Oleg Murk , Olivia Watkins , Patrick Chao , Paul Ashbourne , Pavel Izmailov , Peter Zhokhov , Rachel Dias , Rahul Arora , Randall Lin , Rapha Gontijo Lopes , Raz Gaon , Reah Miyara , Reimar Leike , Renny Hwang , Rhythm Garg , Robin Brown , Roshan James , Rui Shu , Ryan Cheu , Ryan Greene , Saachi Jain , Sam Altman , Sam Toizer , Sam Toyer , Samuel Miserendino , Sandhini Agarwal , Santiago Hernandez , Sasha Baker , Scott McKinney , Scottie Yan , Shengjia Zhao , Shengli Hu , Shibani Santurkar , Shraman Ray Chaudhuri , Shuyuan Zhang , Siyuan Fu , Spencer Papay , Steph Lin , Suchir Balaji , Suvansh Sanjeev , Szymon Sidor , Tal Broda , Aidan Clark , Tao Wang , Taylor Gordon , Ted Sanders , Tejal Patwardhan , Thibault Sottiaux , Thomas Degry , Thomas Dimson , Tianhao Zheng , Timur Garipov , Tom Stasi , Trapit Bansal , Trevor Creech , Troy Peterson , Tyna Eloundou , Valerie Qi , Vineet Kosaraju , Vinnie Monaco , Vitchyr Pong , Vlad Fomenko , Weiyi Zheng , Wenda Zhou , Wenting Zhan , Wes McCabe , Wojciech Zaremba , Yann Dubois , Yinghai Lu , Yining Chen , Young Cha , Yu Bai , Yuchen He , Yuchen Zhang , Yunyun Wang , Zheng Shao , Zhuohan Li

Pretraining corpora contain extensive discourse about AI systems, yet the causal influence of this discourse on downstream alignment remains poorly understood. If prevailing descriptions of AI behaviour are predominantly negative, LLMs may…

Computation and Language · Computer Science 2026-02-23 Cameron Tice , Puria Radmard , Samuel Ratnam , Andy Kim , David Africa , Kyle O'Brien

Reasoning is a cognitive process of using evidence to reach a sound conclusion. The reasoning capability is essential for large language models (LLMs) to serve as the brain of the artificial general intelligence agent. Recent studies reveal…

Computation and Language · Computer Science 2023-09-06 Peiyi Wang , Lei Li , Liang Chen , Feifan Song , Binghuai Lin , Yunbo Cao , Tianyu Liu , Zhifang Sui

The AI-alignment problem arises when there is a discrepancy between the goals that a human designer specifies to an AI learner and a potential catastrophic outcome that does not reflect what the human designer really wants. We argue that a…

Machine Learning · Computer Science 2020-04-10 Shai Shalev-Shwartz , Shaked Shammah , Amnon Shashua

The AI alignment problem, which focusses on ensuring that artificial intelligence (AI), including AGI and ASI, systems act according to human values, presents profound challenges. With the progression from narrow AI to Artificial General…

Artificial Intelligence · Computer Science 2025-07-25 Alberto Hernández-Espinosa , Felipe S. Abrahão , Olaf Witkowski , Hector Zenil

The cost of error in many high-stakes settings is asymmetric: misdiagnosing pneumonia when absent is an inconvenience, but failing to detect it when present can be life-threatening. Because of this, artificial intelligence (AI) models used…

General Economics · Economics 2025-11-12 David Autor , Andrew Caplin , Daniel Martin , Philip Marx

As large language models (LLMs) are increasingly deployed as interactive agents, open-ended human-AI interactions can involve deceptive behaviors with serious real-world consequences, yet existing evaluations remain largely…

Artificial Intelligence · Computer Science 2026-02-09 Yichen Wu , Qianqian Gao , Xudong Pan , Geng Hong , Min Yang
‹ Prev 1 3 4 5 6 7 10 Next ›