English
Related papers

Related papers: Evaluation of OpenAI o1: Opportunities and Challen…

200 papers

The advent of large reasoning models, such as OpenAI o1 and DeepSeek R1, has significantly advanced complex reasoning tasks. However, their capabilities in multilingual complex reasoning remain underexplored, with existing efforts largely…

Computation and Language · Computer Science 2025-05-27 Wenyang Luo , Wayne Xin Zhao , Jing Sha , Shijin Wang , Ji-Rong Wen

Large Language Models (LLMs) are increasingly utilized in AI-driven educational instruction and assessment, particularly within mathematics education. The capability of LLMs to generate accurate answers and detailed solutions for math…

Artificial Intelligence · Computer Science 2025-08-15 Liang Zhang , Edith Aurora Graf

Recent advancements in reasoning-enhanced large language models (LLMs), such as DeepSeek-R1 and OpenAI-o3, have demonstrated significant progress. However, their application in professional medical contexts remains underexplored,…

Computation and Language · Computer Science 2025-03-11 Pengcheng Qiu , Chaoyi Wu , Shuyu Liu , Weike Zhao , Zhuoxia Chen , Hongfei Gu , Chuanjin Peng , Ya Zhang , Yanfeng Wang , Weidi Xie

Currently OpenAI o1 sparks a surge of interest in the study of large reasoning models (LRM). Building on this momentum, Marco-o1 not only focuses on disciplines with standard answers, such as mathematics, physics, and coding -- which are…

Computation and Language · Computer Science 2024-11-26 Yu Zhao , Huifeng Yin , Bo Zeng , Hao Wang , Tianqi Shi , Chenyang Lyu , Longyue Wang , Weihua Luo , Kaifu Zhang

This paper introduces a pioneering approach to artificial intelligence research, embodied in our O1 Replication Journey. In response to the announcement of OpenAI's groundbreaking O1 model, we embark on a transparent, real-time exploration…

Artificial Intelligence · Computer Science 2024-10-28 Yiwei Qin , Xuefeng Li , Haoyang Zou , Yixiu Liu , Shijie Xia , Zhen Huang , Yixin Ye , Weizhe Yuan , Hector Liu , Yuanzhi Li , Pengfei Liu

The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provide new avenues for improving the safety and robustness of our models. In particular, our…

Artificial Intelligence · Computer Science 2026-05-01 OpenAI , : , Aaron Jaech , Adam Kalai , Adam Lerer , Adam Richardson , Ahmed El-Kishky , Aiden Low , Alec Helyar , Aleksander Madry , Alex Beutel , Alex Carney , Alex Iftimie , Alex Karpenko , Alex Tachard Passos , Alexander Neitz , Alexander Prokofiev , Alexander Wei , Allison Tam , Ally Bennett , Ananya Kumar , Andre Saraiva , Andrea Vallone , Andrew Duberstein , Andrew Kondrich , Andrey Mishchenko , Andy Applebaum , Angela Jiang , Ashvin Nair , Barret Zoph , Behrooz Ghorbani , Bohan Zhang , Ben Rossen , Benjamin Sokolowsky , Boaz Barak , Bob McGrew , Borys Minaiev , Botao Hao , Bowen Baker , Brandon Houghton , Brandon McKinzie , Brydon Eastman , Camillo Lugaresi , Cary Bassin , Cary Hudson , Chak Ming Li , Charles de Bourcy , Chelsea Voss , Chen Shen , Chong Zhang , Chris Koch , Chris Orsinger , Christopher Hesse , Claudia Fischer , Clive Chan , Dan Roberts , Daniel Kappler , Daniel Levy , Daniel Selsam , David Dohan , David Farhi , David Mely , David Robinson , Dimitris Tsipras , Doug Li , Dragos Oprica , Eben Freeman , Eddie Zhang , Edmund Wong , Elizabeth Proehl , Enoch Cheung , Eric Mitchell , Eric Wallace , Erik Ritter , Evan Mays , Fan Wang , Felipe Petroski Such , Filippo Raso , Florencia Leoni , Foivos Tsimpourlas , Francis Song , Fred von Lohmann , Freddie Sulit , Geoff Salmon , Giambattista Parascandolo , Gildas Chabot , Grace Zhao , Greg Brockman , Guillaume Leclerc , Hadi Salman , Haiming Bao , Hao Sheng , Hart Andrin , Hessam Bagherinezhad , Hongyu Ren , Hunter Lightman , Hyung Won Chung , Ian Kivlichan , Ian O'Connell , Ian Osband , Ignasi Clavera Gilaberte , Ilge Akkaya , Ilya Kostrikov , Ilya Sutskever , Irina Kofman , Jakub Pachocki , James Lennon , Jason Wei , Jean Harb , Jerry Twore , Jiacheng Feng , Jiahui Yu , Jiayi Weng , Jie Tang , Jieqi Yu , Joaquin Quiñonero Candela , Joe Palermo , Joel Parish , Johannes Heidecke , John Hallman , John Rizzo , Jonathan Gordon , Jonathan Uesato , Jonathan Ward , Joost Huizinga , Julie Wang , Kai Chen , Kai Xiao , Karan Singhal , Karina Nguyen , Karl Cobbe , Katy Shi , Kayla Wood , Kendra Rimbach , Keren Gu-Lemberg , Kevin Liu , Kevin Lu , Kevin Stone , Kevin Yu , Lama Ahmad , Lauren Yang , Leo Liu , Leon Maksin , Leyton Ho , Liam Fedus , Lilian Weng , Linden Li , Lindsay McCallum , Lindsey Held , Lorenz Kuhn , Lukas Kondraciuk , Lukasz Kaiser , Luke Metz , Madelaine Boyd , Maja Trebacz , Manas Joglekar , Mark Chen , Marko Tintor , Mason Meyer , Matt Jones , Matt Kaufer , Max Schwarzer , Meghan Shah , Mehmet Yatbaz , Melody Y. Guan , Mengyuan Xu , Mengyuan Yan , Mia Glaese , Mianna Chen , Michael Lampe , Michael Malek , Michele Wang , Michelle Fradin , Mike McClay , Mikhail Pavlov , Miles Wang , Mingxuan Wang , Mira Murati , Mo Bavarian , Mostafa Rohaninejad , Nat McAleese , Neil Chowdhury , Neil Chowdhury , Nick Ryder , Nikolas Tezak , Noam Brown , Ofir Nachum , Oleg Boiko , Oleg Murk , Olivia Watkins , Patrick Chao , Paul Ashbourne , Pavel Izmailov , Peter Zhokhov , Rachel Dias , Rahul Arora , Randall Lin , Rapha Gontijo Lopes , Raz Gaon , Reah Miyara , Reimar Leike , Renny Hwang , Rhythm Garg , Robin Brown , Roshan James , Rui Shu , Ryan Cheu , Ryan Greene , Saachi Jain , Sam Altman , Sam Toizer , Sam Toyer , Samuel Miserendino , Sandhini Agarwal , Santiago Hernandez , Sasha Baker , Scott McKinney , Scottie Yan , Shengjia Zhao , Shengli Hu , Shibani Santurkar , Shraman Ray Chaudhuri , Shuyuan Zhang , Siyuan Fu , Spencer Papay , Steph Lin , Suchir Balaji , Suvansh Sanjeev , Szymon Sidor , Tal Broda , Aidan Clark , Tao Wang , Taylor Gordon , Ted Sanders , Tejal Patwardhan , Thibault Sottiaux , Thomas Degry , Thomas Dimson , Tianhao Zheng , Timur Garipov , Tom Stasi , Trapit Bansal , Trevor Creech , Troy Peterson , Tyna Eloundou , Valerie Qi , Vineet Kosaraju , Vinnie Monaco , Vitchyr Pong , Vlad Fomenko , Weiyi Zheng , Wenda Zhou , Wenting Zhan , Wes McCabe , Wojciech Zaremba , Yann Dubois , Yinghai Lu , Yining Chen , Young Cha , Yu Bai , Yuchen He , Yuchen Zhang , Yunyun Wang , Zheng Shao , Zhuohan Li

The goal of achieving Artificial General Intelligence (AGI) is to imitate humans and surpass them. Models such as OpenAI's o1, o3, and DeepSeek's R1 have demonstrated that large language models (LLMs) with human-like reasoning capabilities…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Yansheng Qiu , Li Xiao , Zhaopan Xu , Pengfei Zhou , Zheng Wang , Kaipeng Zhang

Large reasoning models (LRMs) like OpenAI-o1 have demonstrated impressive long stepwise reasoning capabilities through large-scale reinforcement learning. However, their extended reasoning processes often suffer from knowledge…

Artificial Intelligence · Computer Science 2025-01-10 Xiaoxi Li , Guanting Dong , Jiajie Jin , Yuyao Zhang , Yujia Zhou , Yutao Zhu , Peitian Zhang , Zhicheng Dou

Frontier AI models demonstrate formidable breadth of knowledge. But how close are they to true human -- or superhuman -- expertise? Genuine experts can tackle the hardest problems and push the boundaries of scientific understanding. To…

Existing benchmarks for frontier models often test specialized, "PhD-level" knowledge that is difficult for non-experts to grasp. In contrast, we present a benchmark with 613 problems based on the NPR Sunday Puzzle Challenge that requires…

Recent advances in test-time scaling of large language models (LLMs), exemplified by DeepSeek-R1 and OpenAI's o1, show that extending the chain of thought during inference can significantly improve general reasoning performance. However,…

Computation and Language · Computer Science 2025-11-11 Yinghao Hu , Yaoyao Yu , Leilei Gan , Bin Wei , Kun Kuang , Fei Wu

Language has long been conceived as an essential tool for human reasoning. The breakthrough of Large Language Models (LLMs) has sparked significant research interest in leveraging these models to tackle complex reasoning tasks. Researchers…

Large Language Models (LLMs) have demonstrated impressive real-world utility, exemplifying artificial useful intelligence (AUI). However, their ability to reason adaptively and robustly -- the hallmarks of artificial general intelligence…

Machine Learning · Computer Science 2025-08-27 Seungwook Han , Jyothish Pari , Samuel J. Gershman , Pulkit Agrawal

Purpose: To evaluate the accuracy and reasoning ability of DeepSeek-R1 and three other recently released large language models (LLMs) in bilingual complex ophthalmology cases. Methods: A total of 130 multiple-choice questions (MCQs) related…

Computation and Language · Computer Science 2025-02-26 Pusheng Xu , Yue Wu , Kai Jin , Xiaolan Chen , Mingguang He , Danli Shi

Recent developments, particularly OpenAI's O1 model, have demonstrated the remarkable potential of Large Language Models (LLMs) for complex reasoning tasks. Through analysis of O1's outputs and provided sample Chain-of-Thought (CoT)…

Artificial Intelligence · Computer Science 2024-12-09 Toby Simonds , Jey Han Lau , Chaithanya Bandi

Large reasoning models (LRMs) like OpenAI o1 and DeepSeek R1 have demonstrated impressive performance on complex reasoning tasks like mathematics and programming with long Chain-of-Thought (CoT) reasoning sequences (slow-thinking), compared…

Artificial Intelligence · Computer Science 2025-07-15 Jason Zhu , Hongyu Li

Recent advancements in Artificial Intelligence (AI), particularly with Large Language Models (LLMs), have led to significant progress in narrow tasks such as image classification, language translation, coding, and writing. However, these…

Artificial Intelligence · Computer Science 2024-12-02 Daniel A. Dollinger , Michael Singleton

Reasoning models are the new generation of Large Language Models (LLMs) capable of complex problem solving. Their reliability in solving introductory physics problems was tested by evaluating a sample of n = 5 solutions generated by one…

Physics Education · Physics 2025-08-29 Amir Bralin , N. Sanjay Rebello

In recent years, the development of Large Language Models (LLMs) has made significant breakthroughs in the field of natural language processing and has gradually been applied to the field of humanities and social sciences research. LLMs…

Computers and Society · Computer Science 2025-04-16 Peiran Gu , Fuhao Duan , Wenhao Li , Bochen Xu , Ying Cai , Teng Yao , Chenxun Zhuo , Tianming Liu , Bao Ge

In this short note, we report and analyze a striking event: OpenAI's large language model o3 has outwitted all students in a university exam on thermodynamics. The thermodynamics exam is a difficult hurdle for most students, where they must…

Computational Engineering, Finance, and Science · Computer Science 2025-06-12 Rebecca Loubet , Pascal Zittlau , Marco Hoffmann , Luisa Vollmer , Sophie Fellenz , Heike Leitte , Fabian Jirasek , Johannes Lenhard , Hans Hasse