English
Related papers

Related papers: OpenDeepThink: Parallel Reasoning via Bradley-Terr…

200 papers

Large Language Models (LLMs) have shown impressive performance on a range of educational tasks, but are still understudied for their potential to solve mathematical problems. In this study, we compare three prominent LLMs, including GPT-4o,…

Artificial Intelligence · Computer Science 2025-07-01 Ruonan Wang , Runxi Wang , Yunwen Shen , Chengfeng Wu , Qinglin Zhou , Rohitash Chandra

Achieving human-level intelligence requires refining the transition from the fast, intuitive System 1 to the slower, more deliberate System 2 reasoning. While System 1 excels in quick, heuristic decisions, System 2 relies on logical…

DeepSeek R1 has significantly advanced complex reasoning for large language models (LLMs). While recent methods have attempted to replicate R1's reasoning capabilities in multimodal settings, they face limitations, including inconsistencies…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Zhehan Kan , Yanlin Liu , Kun Yin , Xinghua Jiang , Xin Li , Haoyu Cao , Yinsong Liu , Deqiang Jiang , Xing Sun , Qingmin Liao , Wenming Yang

Test-time compute scaling, the practice of spending extra computation during inference via repeated sampling, search, or extended reasoning, has become a powerful lever for improving large language model performance. Yet deploying these…

Machine Learning · Computer Science 2026-04-17 Zhiyuan Zhai , Bingcong Li , Bingnan Xiao , Ming Li , Xin Wang

When not using reasoning, repeating the input prompt improves performance for popular models (Gemini, GPT, Claude, and Deepseek) without increasing the number of generated tokens or latency.

Machine Learning · Computer Science 2025-12-18 Yaniv Leviathan , Matan Kalman , Yossi Matias

Current evaluations of mathematical reasoning in large language models (LLMs) are dominated by static benchmarks, either derived from competition-style problems or curated through costly expert effort, resulting in limited coverage of…

Computation and Language · Computer Science 2026-05-08 Jicheng Ma , Guohua Wang , Xinhua Feng , Yiming Liu , Zhichao Hu , Yuhong Liu

Several closed-source LLMs have consistently outperformed open-source alternatives in program repair tasks, primarily due to their superior reasoning capabilities and extensive pre-training. This paper introduces Repairity, a novel…

Software Engineering · Computer Science 2025-06-05 Xunzhu Tang , Jacques Klein , Tegawendé F. Bissyandé

Recent advancements in large reasoning models (LRMs) have demonstrated the effectiveness of scaling test-time computation to enhance reasoning capabilities on various tasks. However, LRMs often suffer from an ``overthinking'' problem, where…

Computation and Language · Computer Science 2025-08-05 Yule Liu , Jingyi Zheng , Zhen Sun , Zifan Peng , Wenhan Dong , Zeyang Sha , Shiwen Cui , Weiqiang Wang , Xinlei He

K2-Think is a reasoning system that achieves state-of-the-art performance with a 32B parameter model, matching or surpassing much larger models like GPT-OSS 120B and DeepSeek v3.1. Built on the Qwen2.5 base model, our system shows that…

Using large language models (LLMs) to solve complex robotics problems requires understanding their planning capabilities. Yet while we know that LLMs can plan on some problems, the extent to which these planning capabilities cover the space…

Robotics · Computer Science 2025-10-02 Jorge Mendez-Mendez

Recent advances of reasoning models, exemplified by OpenAI's o1 and DeepSeek's R1, highlight the significant potential of Reinforcement Learning (RL) to enhance the reasoning capabilities of Large Language Models (LLMs). However,…

Training Large Language Models (LLMs) for chain-of-thought reasoning presents a significant challenge: supervised fine-tuning on a single "golden" rationale hurts generalization as it penalizes equally valid alternatives, whereas…

Computation and Language · Computer Science 2025-11-14 Mingye Zhu , Yi Liu , Zheren Fu , Quan Wang , Yongdong Zhang

We introduce Seed1.5-Thinking, capable of reasoning through thinking before responding, resulting in improved performance on a wide range of benchmarks. Seed1.5-Thinking achieves 86.7 on AIME 2024, 55.0 on Codeforces and 77.3 on GPQA,…

Computation and Language · Computer Science 2025-04-30 ByteDance Seed , : , Jiaze Chen , Tiantian Fan , Xin Liu , Lingjun Liu , Zhiqi Lin , Mingxuan Wang , Chengyi Wang , Xiangpeng Wei , Wenyuan Xu , Yufeng Yuan , Yu Yue , Lin Yan , Qiying Yu , Xiaochen Zuo , Chi Zhang , Ruofei Zhu , Zhecheng An , Zhihao Bai , Yu Bao , Xingyan Bin , Jiangjie Chen , Feng Chen , Hongmin Chen , Riwei Chen , Liangqiang Chen , Zixin Chen , Jinsong Chen , Siyan Chen , Kaiyuan Chen , Zhi Chen , Jin Chen , Jiecao Chen , Jinxin Chi , Weinan Dai , Ning Dai , Jiahui Dai , Shihan Dou , Yantao Du , Zhengyin Du , Jianhui Duan , Chen Dun , Ting-Han Fan , Jiazhan Feng , Junda Feng , Ziyuan Feng , Yuwei Fu , Wenqi Fu , Hanjie Fu , Hao Ge , Hongyi Guo , Mingji Han , Li Han , Wenhao Hao , Xintong Hao , Qianyu He , Jerry He , Feng He , Wen Heng , Zehua Hong , Qi Hou , Liang Hu , Shengding Hu , Nan Hu , Kai Hua , Qi Huang , Ziyue Huang , Hongzhi Huang , Zihao Huang , Ting Huang , Wenhao Huang , Wei Jia , Bin Jia , Xiaoying Jia , Yuhua Jiang , Haobin Jiang , Ziheng Jiang , Kaihua Jiang , Chengquan Jiang , Jianpeng Jiao , Xiaoran Jin , Xing Jin , Xunhao Lai , Zheng Li , Xiang Li , Liyi Li , Hongkai Li , Zheng Li , Shengxian Wan , Ya Wang , Yunshui Li , Chenggang Li , Niuniu Li , Siyu Li , Xi Li , Xiao Li , Aoyan Li , Yuntao Li , Nianning Liang , Xinnian Liang , Haibin Lin , Weijian Lin , Ye Lin , Zhicheng Liu , Guanlin Liu , Guanlin Liu , Chenxiao Liu , Yan Liu , Gaohong Liu , Juncai Liu , Chundian Liu , Deyi Liu , Kaibo Liu , Siyao Liu , Qi Liu , Yongfei Liu , Kang Liu , Gan Liu , Boyi Liu , Rui Long , Weiqiang Lou , Chenwei Lou , Xiang Luo , Yao Luo , Caiping Lv , Heyang Lv , Bole Ma , Qianli Ma , Hongzhi Ma , Yiyuan Ma , Jin Ma , Wenchang Ma , Tingting Ma , Chen Mao , Qiyang Min , Zhe Nan , Guanghan Ning , Jinxiang Ou , Haojie Pan , Renming Pang , Yanghua Peng , Tao Peng , Lihua Qian , Lihua Qian , Mu Qiao , Meng Qu , Cheng Ren , Hongbin Ren , Yong Shan , Wei Shen , Ke Shen , Kai Shen , Guangming Sheng , Jinlong Shi , Wenlei Shi , Guang Shi , Shuai Shuai Cao , Yuxin Song , Zuquan Song , Jing Su , Yifan Sun , Tao Sun , Zewei Sun , Borui Wan , Zihan Wang , Xiaohui Wang , Xi Wang , Shuguang Wang , Jun Wang , Qinlong Wang , Chenyuan Wang , Shuai Wang , Zihan Wang , Changbao Wang , Jiaqiang Wang , Shihang Wang , Xuwu Wang , Zaiyuan Wang , Yuxuan Wang , Wenqi Wang , Taiqing Wang , Chengzhi Wei , Houmin Wei , Ziyun Wei , Shufa Wei , Zheng Wu , Yonghui Wu , Yangjun Wu , Bohong Wu , Shuang Wu , Jingqiao Wu , Ning Wu , Shuangzhi Wu , Jianmin Wu , Chenguang Xi , Fan Xia , Yuqiao Xian , Liang Xiang , Boren Xiang , Bowen Xiao , Zhen Xiao , Xia Xiao , Yongsheng Xiao , Chao Xin , Shulin Xin , Yuwen Xiong , Jingjing Xu , Ziwen Xu , Chenyin Xu , Jiayi Xu , Yifan Xu , Wei Xu , Yufei Xu , Shikun Xu , Shipeng Yan , Shen Yan , Qingping Yang , Xi Yang , Tianhao Yang , Yuehang Yang , Yuan Yang , Ximing Yang , Zeyu Yang , Guang Yang , Yifan Yang , Xuesong Yao , Bairen Yi , Fan Yin , Jianian Yin , Ziqiang Ying , Xiangyu Yu , Hongli Yu , Song Yu , Menghan Yu , Huan Yu , Siyu Yuan , Jun Yuan , Yutao Zeng , Tianyang Zhan , Zheng Zhang , Yun Zhang , Mofan Zhang , Wang Zhang , Ru Zhang , Zhi Zhang , Tianqi Zhang , Xinyi Zhang , Zhexi Zhang , Sijun Zhang , Wenqiang Zhang , Xiangxiang Zhang , Yongtao Zhang , Yuyu Zhang , Ge Zhang , He Zhang , Yue Zhang , Renjie Zheng , Ningxin Zheng , Zhuolin Zheng , Yaowei Zheng , Chen Zheng , Xiaoyun Zhi , Wanjun Zhong , Cheng Zhong , Zheng Zhong , Baoquan Zhong , Xun Zhou , Na Zhou , Huan Zhou , Hang Zhu , Defa Zhu , Wenjia Zhu , Lei Zuo

Despite recent efforts to develop large language models with robust long-context capabilities, the lack of long-context benchmarks means that relatively little is known about their performance. To alleviate this gap, in this paper, we…

Computation and Language · Computer Science 2024-12-25 Mingyang Song , Mao Zheng , Xuan Luo

Large reasoning models (LRMs) are proficient at generating explicit, step-by-step reasoning sequences before producing final answers. However, such detailed reasoning can introduce substantial computational overhead and latency,…

Computation and Language · Computer Science 2025-10-10 Songjun Tu , Jiahao Lin , Qichao Zhang , Xiangyu Tian , Linjing Li , Xiangyuan Lan , Dongbin Zhao

Despite significant advancements in the general capability of large language models (LLMs), they continue to struggle with consistent and accurate reasoning, especially in complex tasks such as mathematical and code reasoning. One key…

Machine Learning · Computer Science 2024-10-10 Zhenwen Liang , Ye Liu , Tong Niu , Xiangliang Zhang , Yingbo Zhou , Semih Yavuz

We introduce ThinkTwice, a simple two-phase framework that jointly optimizes LLMs to solve reasoning problems and refine the answers, based on Group Relative Policy Optimization (GRPO). In each pair of training steps, ThinkTwice first…

Artificial Intelligence · Computer Science 2026-04-08 Difan Jiao , Qianfeng Wen , Blair Yang , Zhenwei Tang , Ashton Anderson

Large Language Models (LLMs), despite their remarkable capabilities, are prone to generating hallucinated or outdated content due to their static internal knowledge. While Retrieval-Augmented Generation (RAG) integrated with Reinforcement…

Computation and Language · Computer Science 2026-01-14 Zhiwen Tan , Jiaming Huang , Qintong Wu , Hongxuan Zhang , Chenyi Zhuang , Jinjie Gu

Large language models (LLMs) can face factual limitations when responding to time-sensitive queries about recent events that arise after their knowledge thresholds in the training corpus. Existing search-augmented approaches fall into two…

Information Retrieval · Computer Science 2025-06-11 Wentao Shi , Yiqing Shen

Test-time scaling has significantly improved large language model performance, enabling deeper reasoning to solve complex problems. However, this increased reasoning capability also leads to excessive token generation and unnecessary…

‹ Prev 1 8 9 10 Next ›