English
Related papers

Related papers: System 2 thinking in OpenAI's o1-preview model: Ne…

200 papers

As artificial intelligence (AI) continues to advance, it demonstrates capabilities comparable to human intelligence, with significant potential to transform education and workforce development. This study evaluates OpenAI o1-preview's…

This study evaluates the performance of OpenAI's o1-preview model in higher-order cognitive domains, including critical thinking, systematic thinking, computational thinking, data literacy, creative thinking, logical reasoning, and…

Computers and Society · Computer Science 2024-12-10 Ehsan Latif , Yifan Zhou , Shuchen Guo , Lehong Shi , Yizhu Gao , Matthew Nyaaba , Arne Bewerdorff , Xiantong Yang , Xiaoming Zhai

Enabling Large Language Models (LLMs) to handle a wider range of complex tasks (e.g., coding, math) has drawn great attention from many researchers. As LLMs continue to evolve, merely increasing the number of model parameters yields…

The remarkable performance of models like the OpenAI o1 can be attributed to their ability to emulate human-like long-time thinking during inference. These models employ extended chain-of-thought (CoT) processes, exploring multiple…

Computation and Language · Computer Science 2025-02-04 Xingyu Chen , Jiahao Xu , Tian Liang , Zhiwei He , Jianhui Pang , Dian Yu , Linfeng Song , Qiuzhi Liu , Mengfei Zhou , Zhuosheng Zhang , Rui Wang , Zhaopeng Tu , Haitao Mi , Dong Yu

Recently, slow-thinking reasoning systems, such as o1, have demonstrated remarkable capabilities in solving complex reasoning tasks. These systems typically engage in an extended thinking process before responding to a query, allowing them…

Test-time scaling is a promising new approach to language modeling that uses extra test-time compute to improve performance. Recently, OpenAI's o1 model showed this capability but did not publicly share its methodology, leading to many…

Large language models (LLMs) such as OpenAI's o1 have demonstrated remarkable abilities in complex reasoning tasks by scaling test-time compute and exhibiting human-like deep thinking. However, we identify a phenomenon we term…

Computation and Language · Computer Science 2025-02-19 Yue Wang , Qiuzhi Liu , Jiahao Xu , Tian Liang , Xingyu Chen , Zhiwei He , Linfeng Song , Dian Yu , Juntao Li , Zhuosheng Zhang , Rui Wang , Zhaopeng Tu , Haitao Mi , Dong Yu

The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provide new avenues for improving the safety and robustness of our models. In particular, our…

Artificial Intelligence · Computer Science 2026-05-01 OpenAI , : , Aaron Jaech , Adam Kalai , Adam Lerer , Adam Richardson , Ahmed El-Kishky , Aiden Low , Alec Helyar , Aleksander Madry , Alex Beutel , Alex Carney , Alex Iftimie , Alex Karpenko , Alex Tachard Passos , Alexander Neitz , Alexander Prokofiev , Alexander Wei , Allison Tam , Ally Bennett , Ananya Kumar , Andre Saraiva , Andrea Vallone , Andrew Duberstein , Andrew Kondrich , Andrey Mishchenko , Andy Applebaum , Angela Jiang , Ashvin Nair , Barret Zoph , Behrooz Ghorbani , Bohan Zhang , Ben Rossen , Benjamin Sokolowsky , Boaz Barak , Bob McGrew , Borys Minaiev , Botao Hao , Bowen Baker , Brandon Houghton , Brandon McKinzie , Brydon Eastman , Camillo Lugaresi , Cary Bassin , Cary Hudson , Chak Ming Li , Charles de Bourcy , Chelsea Voss , Chen Shen , Chong Zhang , Chris Koch , Chris Orsinger , Christopher Hesse , Claudia Fischer , Clive Chan , Dan Roberts , Daniel Kappler , Daniel Levy , Daniel Selsam , David Dohan , David Farhi , David Mely , David Robinson , Dimitris Tsipras , Doug Li , Dragos Oprica , Eben Freeman , Eddie Zhang , Edmund Wong , Elizabeth Proehl , Enoch Cheung , Eric Mitchell , Eric Wallace , Erik Ritter , Evan Mays , Fan Wang , Felipe Petroski Such , Filippo Raso , Florencia Leoni , Foivos Tsimpourlas , Francis Song , Fred von Lohmann , Freddie Sulit , Geoff Salmon , Giambattista Parascandolo , Gildas Chabot , Grace Zhao , Greg Brockman , Guillaume Leclerc , Hadi Salman , Haiming Bao , Hao Sheng , Hart Andrin , Hessam Bagherinezhad , Hongyu Ren , Hunter Lightman , Hyung Won Chung , Ian Kivlichan , Ian O'Connell , Ian Osband , Ignasi Clavera Gilaberte , Ilge Akkaya , Ilya Kostrikov , Ilya Sutskever , Irina Kofman , Jakub Pachocki , James Lennon , Jason Wei , Jean Harb , Jerry Twore , Jiacheng Feng , Jiahui Yu , Jiayi Weng , Jie Tang , Jieqi Yu , Joaquin Quiñonero Candela , Joe Palermo , Joel Parish , Johannes Heidecke , John Hallman , John Rizzo , Jonathan Gordon , Jonathan Uesato , Jonathan Ward , Joost Huizinga , Julie Wang , Kai Chen , Kai Xiao , Karan Singhal , Karina Nguyen , Karl Cobbe , Katy Shi , Kayla Wood , Kendra Rimbach , Keren Gu-Lemberg , Kevin Liu , Kevin Lu , Kevin Stone , Kevin Yu , Lama Ahmad , Lauren Yang , Leo Liu , Leon Maksin , Leyton Ho , Liam Fedus , Lilian Weng , Linden Li , Lindsay McCallum , Lindsey Held , Lorenz Kuhn , Lukas Kondraciuk , Lukasz Kaiser , Luke Metz , Madelaine Boyd , Maja Trebacz , Manas Joglekar , Mark Chen , Marko Tintor , Mason Meyer , Matt Jones , Matt Kaufer , Max Schwarzer , Meghan Shah , Mehmet Yatbaz , Melody Y. Guan , Mengyuan Xu , Mengyuan Yan , Mia Glaese , Mianna Chen , Michael Lampe , Michael Malek , Michele Wang , Michelle Fradin , Mike McClay , Mikhail Pavlov , Miles Wang , Mingxuan Wang , Mira Murati , Mo Bavarian , Mostafa Rohaninejad , Nat McAleese , Neil Chowdhury , Neil Chowdhury , Nick Ryder , Nikolas Tezak , Noam Brown , Ofir Nachum , Oleg Boiko , Oleg Murk , Olivia Watkins , Patrick Chao , Paul Ashbourne , Pavel Izmailov , Peter Zhokhov , Rachel Dias , Rahul Arora , Randall Lin , Rapha Gontijo Lopes , Raz Gaon , Reah Miyara , Reimar Leike , Renny Hwang , Rhythm Garg , Robin Brown , Roshan James , Rui Shu , Ryan Cheu , Ryan Greene , Saachi Jain , Sam Altman , Sam Toizer , Sam Toyer , Samuel Miserendino , Sandhini Agarwal , Santiago Hernandez , Sasha Baker , Scott McKinney , Scottie Yan , Shengjia Zhao , Shengli Hu , Shibani Santurkar , Shraman Ray Chaudhuri , Shuyuan Zhang , Siyuan Fu , Spencer Papay , Steph Lin , Suchir Balaji , Suvansh Sanjeev , Szymon Sidor , Tal Broda , Aidan Clark , Tao Wang , Taylor Gordon , Ted Sanders , Tejal Patwardhan , Thibault Sottiaux , Thomas Degry , Thomas Dimson , Tianhao Zheng , Timur Garipov , Tom Stasi , Trapit Bansal , Trevor Creech , Troy Peterson , Tyna Eloundou , Valerie Qi , Vineet Kosaraju , Vinnie Monaco , Vitchyr Pong , Vlad Fomenko , Weiyi Zheng , Wenda Zhou , Wenting Zhan , Wes McCabe , Wojciech Zaremba , Yann Dubois , Yinghai Lu , Yining Chen , Young Cha , Yu Bai , Yuchen He , Yuchen Zhang , Yunyun Wang , Zheng Shao , Zhuohan Li

The Orion-1 model by OpenAI is claimed to have more robust logical reasoning capabilities than previous large language models. However, some suggest the excellence might be partially due to the model "memorizing" solutions, resulting in…

Artificial Intelligence · Computer Science 2024-11-12 Leo Li , Ye Luo , Tingyou Pan

Achieving human-level intelligence requires refining the transition from the fast, intuitive System 1 to the slower, more deliberate System 2 reasoning. While System 1 excels in quick, heuristic decisions, System 2 relies on logical…

This paper presents a critical examination of current approaches to replicating OpenAI's O1 model capabilities, with particular focus on the widespread but often undisclosed use of knowledge distillation techniques. While our previous work…

Computation and Language · Computer Science 2024-11-26 Zhen Huang , Haoyang Zou , Xuefeng Li , Yixiu Liu , Yuxiang Zheng , Ethan Chern , Shijie Xia , Yiwei Qin , Weizhe Yuan , Pengfei Liu

The o1 system card identifies the o1 models as the most robust within OpenAI, with their defining characteristic being the progression from rapid, intuitive thinking to slower, more deliberate reasoning. This observation motivated us to…

Computation and Language · Computer Science 2025-01-15 Yuhang Wang , Yuxiang Zhang , Yanxu Zhu , Xinyan Wen , Jitao Sang

This paper presents a case study of coding tasks by the latest reasoning models of OpenAI, i.e. o1-preview and o1-mini, in comparison with other frontier models. The o1 models deliver SOTA results for WebApp1K, a single-task benchmark. To…

Software Engineering · Computer Science 2024-09-24 Yi Cui

In "Embers of Autoregression" (McCoy et al., 2023), we showed that several large language models (LLMs) have some important limitations that are attributable to their origins in next-word prediction. Here we investigate whether these issues…

Computation and Language · Computer Science 2024-10-07 R. Thomas McCoy , Shunyu Yao , Dan Friedman , Mathew D. Hardy , Thomas L. Griffiths

Recent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1…

Large language models (LLMs) have exhibited remarkable capabilities across various domains and tasks, pushing the boundaries of our knowledge in learning and cognition. The latest model, OpenAI's o1, stands out as the first LLM with an…

Computation and Language · Computer Science 2024-09-24 Yunfei Xie , Juncheng Wu , Haoqin Tu , Siwei Yang , Bingchen Zhao , Yongshuo Zong , Qiao Jin , Cihang Xie , Yuyin Zhou

Large Language Models (LLMs) are increasingly utilized in AI-driven educational instruction and assessment, particularly within mathematics education. The capability of LLMs to generate accurate answers and detailed solutions for math…

Artificial Intelligence · Computer Science 2025-08-15 Liang Zhang , Edith Aurora Graf

OpenAI o1 represents a significant milestone in Artificial Inteiligence, which achieves expert-level performances on many challanging tasks that require strong reasoning ability.OpenAI has claimed that the main techinique behinds o1 is the…

Artificial Intelligence · Computer Science 2024-12-19 Zhiyuan Zeng , Qinyuan Cheng , Zhangyue Yin , Bo Wang , Shimin Li , Yunhua Zhou , Qipeng Guo , Xuanjing Huang , Xipeng Qiu

The releases of OpenAI's o-[n] series, such as o1, o3, and o4-mini, mark a significant paradigm shift in Large Language Models towards advanced reasoning capabilities. Notably, models like o3 have demonstrated strong performance on…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Vernon Y. H. Toh , Yew Ken Chia , Deepanway Ghosal , Soujanya Poria
‹ Prev 1 2 3 10 Next ›