English
Related papers

Related papers: Ring-lite: Scalable Reasoning via C3PO-Stabilized …

200 papers

This technical report presents Ring-Lite-Distill, a lightweight reasoning model derived from our open-source Mixture-of-Experts (MoE) Large Language Models (LLMs) Ling-Lite. This study demonstrates that through meticulous high-quality data…

We present Ring-1T, the first open-source, state-of-the-art thinking model with a trillion-scale parameter. It features 1 trillion total parameters and activates approximately 50 billion per token. Training such models at a…

Recent advances in reinforcement learning (RL) have substantially improved the training of large-scale language models, leading to significant gains in generation quality and reasoning ability. However, most existing research focuses on…

Machine Learning · Computer Science 2026-01-13 Di Zhang , Xun Wu , Shaohan Huang , Lingjie Jiang , Yaru Hao , Li Dong , Zewen Chi , Zhifang Sui , Furu Wei

Recent advancements in the reasoning capabilities of large language models (LLMs) show that employing group relative policy optimization (GRPO) algorithm for reinforcement learning (RL) training allows the models to use more…

Computation and Language · Computer Science 2025-07-04 Purbesh Mitra , Sennur Ulukus

Large language models have recently evolved from fluent text generation to advanced reasoning across diverse domains, giving rise to reasoning language models. Among these domains, mathematical reasoning serves as a representative benchmark…

The effective training of Large Language Models (LLMs) for function calling faces a critical challenge: balancing exploration of complex reasoning paths with stable policy optimization. Standard methods like Supervised Fine-Tuning (SFT)…

We introduce Ling 2.0, a series reasoning-oriented language foundation built upon the principle that every activation boosts reasoning capability. Designed to scale from tens of billions to one trillion parameters under a unified…

Computation and Language · Computer Science 2025-11-10 Ling Team , Ang Li , Ben Liu , Binbin Hu , Bing Li , Bingwei Zeng , Borui Ye , Caizhi Tang , Changxin Tian , Chao Huang , Chao Zhang , Chen Qian , Chenchen Ju , Chenchen Li , Chengfu Tang , Chilin Fu , Chunshao Ren , Chunwei Wu , Cong Zhang , Cunyin Peng , Dafeng Xu , Daixin Wang , Dalong Zhang , Dingnan Jin , Dingyuan Zhu , Dongke Hu , Fangzheng Zhao , Feifan Wu , Feng Zhu , Gangshan Wang , Haitao Zhang , Hailin Zhao , Hanxiao Zhang , Hanzi Wang , Hao Qian , Haoyi Yu , Heng Zhang , Hongliang Zhang , Hongzhi Luan , Huirong Dong , Huizhong Li , Jia Li , Jia Liu , Jialong Zhu , Jian Sha , Jianping Wei , Jiaolong Yang , Jieyue Ma , Jiewei Wu , Jinjing Huang , Jingyun Tian , Jingyuan Zhang , Jinquan Sun , Juanhui Tu , Jun Liu , Jun Xu , Jun Zhou , Junjie Ou , Junpeng Fang , Kaihong Zhang , Kaiqin Hu , Ke Shi , Kun Tang , Kunlong Chen , Lanyin Mei , Lei Liang , Lei Xu , Libo Zhang , Lin Ju , Lin Yuan , Ling Zhong , Lintao Ma , Lu Liu , Lu Yu , Lun Cai , Meiqi Zhu , Mengying Li , Min Chen , Minghao Xue , Minghong Cai , Mingming Yin , Peijie Jiang , Peilong Zhao , Pingping Liu , Qian Zhao , Qing Cui , Qingxiang Huang , Qingyuan Yang , Quankun Yu , Shaowei Wei , Shijie Lian , Shoujian Zheng , Shun Song , Shungen Zhang , Shuo Zhang , Siyuan Li , Song Liu , Ting Guo , Tong Zhao , Wanli Gu , Weichang Wu , Weiguang Han , Wenjing Fang , Wubin Wang , Xiang Shu , Xiao Shi , Xiaoshun Lan , Xiaolu Zhang , Xiaqing Sun , Xin Zhao , Xingyu Lu , Xiong Xu , Xudong Wang , Xudong Wang , Xuemin Yang , Yajie Yang , Yang Xiang , Yanzhe Li , Yi Zhang , Yilong Wang , Yingxue Li , Yongzhen Guo , Yuzhuo Fu , Yuanyuan Wang , Yue Yang , Yue Yu , Yufeng Deng , Yun Zhang , Yunfei Yu , Yuqi Zhang , Yuxiao He , Zengke Gui , Zhaoxin Huan , Zhaoyang Wang , Zhibo Zhu , Zhihao Wang , Zhiqiang Zhang , Zhoufei Wang , Zihang Zeng , Ziqi Liu , Zitao Xuan , Zuoli Tang

Current Large Language Models (LLMs) exhibit significant limitations, notably in structured, interpretable, and verifiable medical reasoning, alongside practical deployment challenges related to computational resources and data privacy.…

Computation and Language · Computer Science 2025-06-06 Boqin Zhuang , Chenxiao Song , Huitong Lu , Jiacheng Qiao , Mingqian Liu , Mingxing Yu , Ping Hong , Rui Li , Xiaoxia Song , Xiangjun Xu , Xu Chen , Yaoyao Ma , Yujie Gao

Recent advances in reinforcement learning (RL) have significantly enhanced the reasoning capabilities of large language models (LLMs). Group Relative Policy Optimization (GRPO), a lightweight variant of Proximal Policy Optimization (PPO),…

Machine Learning · Computer Science 2025-10-13 Chen Wang , Lai Wei , Yanzhi Zhang , Chenyang Shao , Zedong Dan , Weiran Huang , Yuzhi Zhang , Yue Wang

Recent advances in fine-tuning large language models (LLMs) with reinforcement learning (RL) have shown promising improvements in complex reasoning tasks, particularly when paired with chain-of-thought (CoT) prompting. However, these…

Machine Learning · Computer Science 2025-04-04 Hung Le , Dai Do , Dung Nguyen , Svetha Venkatesh

Long chain-of-thought (CoT) significantly enhances the reasoning capabilities of large language models (LLMs). However, extensive reasoning traces lead to inefficiencies and increased time-to-first-token (TTFT). We propose a training…

Computation and Language · Computer Science 2026-01-08 Roy Xie , David Qiu , Deepak Gopinath , Dong Lin , Yanchao Sun , Chong Wang , Saloni Potdar , Bhuwan Dhingra

Diffusion large language models (dLLMs) are promising alternatives to autoregressive large language models (AR-LLMs), as they potentially allow higher inference throughput. Reinforcement learning (RL) is a crucial component for dLLMs to…

Machine Learning · Computer Science 2026-02-24 Yuchen Zhu , Wei Guo , Jaemoo Choi , Petr Molodyk , Bo Yuan , Molei Tao , Yongxin Chen

Reinforcement learning (RL) has become the dominant paradigm for improving the performance of language models on complex reasoning tasks. Despite the substantial empirical gains demonstrated by RL-based training methods like GRPO, a…

Artificial Intelligence · Computer Science 2025-10-27 Jiayu Wang , Yifei Ming , Zixuan Ke , Caiming Xiong , Shafiq Joty , Aws Albarghouthi , Frederic Sala

Enhancing the reasoning capabilities of large language models (LLMs) typically relies on massive computational resources and extensive datasets, limiting accessibility for resource-constrained settings. Our study investigates the potential…

Machine Learning · Computer Science 2026-01-21 Quy-Anh Dang , Chris Ngo

Multimodal large language models (MLLMs) have shown promising capabilities in reasoning tasks, yet still struggle with complex problems requiring explicit self-reflection and self-correction, especially compared to their unimodal text-based…

Computation and Language · Computer Science 2025-10-07 Zhongwei Wan , Zhihao Dou , Che Liu , Yu Zhang , Dongfei Cui , Qinjian Zhao , Hui Shen , Jing Xiong , Yi Xin , Yifan Jiang , Chaofan Tao , Yangfan He , Mi Zhang , Shen Yan

Recent advances of reasoning models, exemplified by OpenAI's o1 and DeepSeek's R1, highlight the significant potential of Reinforcement Learning (RL) to enhance the reasoning capabilities of Large Language Models (LLMs). However,…

Despite recent advancements in language models (LMs), their application to dialogue management (DM) problems and ability to carry on rich conversations remain a challenge. We use reinforcement learning (RL) to develop a dialogue agent that…

Computation and Language · Computer Science 2022-06-02 Yinlam Chow , Aza Tulepbergenov , Ofir Nachum , MoonKyung Ryu , Mohammad Ghavamzadeh , Craig Boutilier

While Reinforcement Learning (RL) shows promise in training tool-use Large Language Models (LLMs) using verifiable outcome rewards, existing methods largely overlook the potential of reasoning rewards based on chain-of-thought quality for…

Computation and Language · Computer Science 2026-01-16 Zihan Lin , Xiaohan Wang , Hexiong Yang , Jiajun Chai , Jie Cao , Guojun Yin , Wei Lin , Ran He

We present MoE-MLA-RoPE, a novel architecture combination that combines Mixture of Experts (MoE) with Multi-head Latent Attention (MLA) and Rotary Position Embeddings (RoPE) for efficient language modeling. Our approach addresses the…

Artificial Intelligence · Computer Science 2025-08-05 Sushant Mehta , Raj Dandekar , Rajat Dandekar , Sreedath Panat
‹ Prev 1 2 3 10 Next ›