English
Related papers

Related papers: Nanbeige4-3B Technical Report: Exploring the Front…

200 papers

In this report, we introduce Qwen2.5, a comprehensive series of large language models (LLMs) designed to meet diverse needs. Compared to previous iterations, Qwen 2.5 has been significantly improved during both the pre-training and…

We introduce FuseChat-3.0, a suite of large language models (LLMs) developed by integrating the strengths of heterogeneous source LLMs into more compact target LLMs. Our source models include the powerful Gemma-2-27B-it,…

Computation and Language · Computer Science 2025-03-07 Ziyi Yang , Fanqi Wan , Longguang Zhong , Canbin Huang , Guosheng Liang , Xiaojun Quan

Large language models (LLMs) have demonstrated remarkable capabilities in problem-solving. However, their proficiency in solving mathematical problems remains inadequate. We propose MathScale, a simple and scalable method to create…

Computation and Language · Computer Science 2024-03-06 Zhengyang Tang , Xingxing Zhang , Benyou Wang , Furu Wei

The introduction of the transformer architecture and the self-attention mechanism has led to an explosive production of language models trained on specific downstream tasks and data domains. With over 200, 000 models in the Hugging Face…

Machine Learning · Computer Science 2023-08-24 Surya Narayanan Hari , Matt Thomson

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 are as follows: (1) DeepSeek Sparse Attention (DSA): We…

Computation and Language · Computer Science 2025-12-03 DeepSeek-AI , Aixin Liu , Aoxue Mei , Bangcai Lin , Bing Xue , Bingxuan Wang , Bingzheng Xu , Bochao Wu , Bowei Zhang , Chaofan Lin , Chen Dong , Chengda Lu , Chenggang Zhao , Chengqi Deng , Chenhao Xu , Chong Ruan , Damai Dai , Daya Guo , Dejian Yang , Deli Chen , Erhang Li , Fangqi Zhou , Fangyun Lin , Fucong Dai , Guangbo Hao , Guanting Chen , Guowei Li , H. Zhang , Hanwei Xu , Hao Li , Haofen Liang , Haoran Wei , Haowei Zhang , Haowen Luo , Haozhe Ji , Honghui Ding , Hongxuan Tang , Huanqi Cao , Huazuo Gao , Hui Qu , Hui Zeng , Jialiang Huang , Jiashi Li , Jiaxin Xu , Jiewen Hu , Jingchang Chen , Jingting Xiang , Jingyang Yuan , Jingyuan Cheng , Jinhua Zhu , Jun Ran , Junguang Jiang , Junjie Qiu , Junlong Li , Junxiao Song , Kai Dong , Kaige Gao , Kang Guan , Kexin Huang , Kexing Zhou , Kezhao Huang , Kuai Yu , Lean Wang , Lecong Zhang , Lei Wang , Liang Zhao , Liangsheng Yin , Lihua Guo , Lingxiao Luo , Linwang Ma , Litong Wang , Liyue Zhang , M. S. Di , M. Y Xu , Mingchuan Zhang , Minghua Zhang , Minghui Tang , Mingxu Zhou , Panpan Huang , Peixin Cong , Peiyi Wang , Qiancheng Wang , Qihao Zhu , Qingyang Li , Qinyu Chen , Qiushi Du , Ruiling Xu , Ruiqi Ge , Ruisong Zhang , Ruizhe Pan , Runji Wang , Runqiu Yin , Runxin Xu , Ruomeng Shen , Ruoyu Zhang , S. H. Liu , Shanghao Lu , Shangyan Zhou , Shanhuang Chen , Shaofei Cai , Shaoyuan Chen , Shengding Hu , Shengyu Liu , Shiqiang Hu , Shirong Ma , Shiyu Wang , Shuiping Yu , Shunfeng Zhou , Shuting Pan , Songyang Zhou , Tao Ni , Tao Yun , Tian Pei , Tian Ye , Tianyuan Yue , Wangding Zeng , Wen Liu , Wenfeng Liang , Wenjie Pang , Wenjing Luo , Wenjun Gao , Wentao Zhang , Xi Gao , Xiangwen Wang , Xiao Bi , Xiaodong Liu , Xiaohan Wang , Xiaokang Chen , Xiaokang Zhang , Xiaotao Nie , Xin Cheng , Xin Liu , Xin Xie , Xingchao Liu , Xingkai Yu , Xingyou Li , Xinyu Yang , Xinyuan Li , Xu Chen , Xuecheng Su , Xuehai Pan , Xuheng Lin , Xuwei Fu , Y. Q. Wang , Yang Zhang , Yanhong Xu , Yanru Ma , Yao Li , Yao Li , Yao Zhao , Yaofeng Sun , Yaohui Wang , Yi Qian , Yi Yu , Yichao Zhang , Yifan Ding , Yifan Shi , Yiliang Xiong , Ying He , Ying Zhou , Yinmin Zhong , Yishi Piao , Yisong Wang , Yixiao Chen , Yixuan Tan , Yixuan Wei , Yiyang Ma , Yiyuan Liu , Yonglun Yang , Yongqiang Guo , Yongtong Wu , Yu Wu , Yuan Cheng , Yuan Ou , Yuanfan Xu , Yuduan Wang , Yue Gong , Yuhan Wu , Yuheng Zou , Yukun Li , Yunfan Xiong , Yuxiang Luo , Yuxiang You , Yuxuan Liu , Yuyang Zhou , Z. F. Wu , Z. Z. Ren , Zehua Zhao , Zehui Ren , Zhangli Sha , Zhe Fu , Zhean Xu , Zhenda Xie , Zhengyan Zhang , Zhewen Hao , Zhibin Gou , Zhicheng Ma , Zhigang Yan , Zhihong Shao , Zhixian Huang , Zhiyu Wu , Zhuoshu Li , Zhuping Zhang , Zian Xu , Zihao Wang , Zihui Gu , Zijia Zhu , Zilin Li , Zipeng Zhang , Ziwei Xie , Ziyi Gao , Zizheng Pan , Zongqing Yao , Bei Feng , Hui Li , J. L. Cai , Jiaqi Ni , Lei Xu , Meng Li , Ning Tian , R. J. Chen , R. L. Jin , S. S. Li , Shuang Zhou , Tianyu Sun , X. Q. Li , Xiangyue Jin , Xiaojin Shen , Xiaosha Chen , Xinnan Song , Xinyi Zhou , Y. X. Zhu , Yanping Huang , Yaohui Li , Yi Zheng , Yuchen Zhu , Yunxian Ma , Zhen Huang , Zhipeng Xu , Zhongyu Zhang , Dongjie Ji , Jian Liang , Jianzhong Guo , Jin Chen , Leyi Xia , Miaojun Wang , Mingming Li , Peng Zhang , Ruyi Chen , Shangmian Sun , Shaoqing Wu , Shengfeng Ye , T. Wang , W. L. Xiao , Wei An , Xianzu Wang , Xiaowen Sun , Xiaoxiang Wang , Ying Tang , Yukun Zha , Zekai Zhang , Zhe Ju , Zhen Zhang , Zihua Qu

Recent years have witnessed a clear trend towards language models with an ever-increasing number of parameters, as well as the growing training overhead and memory usage. Distributed training, particularly through Sharded Data Parallelism…

Machine Learning · Computer Science 2024-11-26 Jinda Jia , Cong Xie , Hanlin Lu , Daoce Wang , Hao Feng , Chengming Zhang , Baixi Sun , Haibin Lin , Zhi Zhang , Xin Liu , Dingwen Tao

Discrete diffusion models have emerged as strong alternatives to autoregressive language models, with recent work initializing and fine-tuning a base unimodal model for bimodal generation. Diverging from previous approaches, we introduce…

Diffusion models have demonstrated strong potential in language modeling, offering various advantages over traditional autoregressive approaches. Their ability to generate and revise entire responses in parallel enables faster generation…

Machine Learning · Computer Science 2026-03-03 Michael Hersche , Samuel Moor-Smith , Thomas Hofmann , Abbas Rahimi

This work introduces Falcon-H1R, a 7B-parameter reasoning-optimized model that establishes the feasibility of achieving competitive reasoning performance with small language models (SLMs). Falcon-H1R stands out for its parameter efficiency,…

Large language models have recently enabled a generative paradigm for query expansion, but their high inference cost makes direct deployment difficult in practical retrieval systems. To address this issue, a retrieval-feedback-driven…

Information Retrieval · Computer Science 2026-03-17 Minghan Li , Guodong Zhou

In this work, we introduce Speech-Copilot, a modular framework for instruction-oriented speech-processing tasks that minimizes human effort in toolset construction. Unlike end-to-end methods using large audio-language models, Speech-Copilot…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 Chun-Yi Kuan , Chih-Kai Yang , Wei-Ping Huang , Ke-Han Lu , Hung-yi Lee

Large Language Models exhibit impressive reasoning capabilities across diverse tasks, motivating efforts to distill these capabilities into smaller models through generated reasoning data. However, direct training on such synthesized…

Computation and Language · Computer Science 2025-02-05 Shengmin Piao , Sanghyun Park

In this work, we investigate the synergy between supervised fine-tuning (SFT) and reinforcement learning (RL) in developing strong reasoning models. We begin by curating the SFT training data through two scaling strategies: increasing the…

Computation and Language · Computer Science 2025-06-17 Zihan Liu , Zhuolin Yang , Yang Chen , Chankyu Lee , Mohammad Shoeybi , Bryan Catanzaro , Wei Ping

Large language models (LLMs) achieve strong performance across many natural language processing tasks, yet their decision processes remain difficult to interpret. This lack of transparency creates challenges for trust, debugging, and…

Computation and Language · Computer Science 2026-04-20 Venkata Abhinandan Kancharla

Multimodal Large Language Models (MLLMs) are undergoing rapid progress and represent the frontier of AI development. However, their training and inference efficiency have emerged as a core bottleneck in making MLLMs more accessible and…

Small-scale models offer various computational advantages, and yet to which extent size is critical for problem-solving abilities remains an open question. Specifically for solving grade school math, the smallest model size so far required…

Machine Learning · Computer Science 2023-12-15 Bingbin Liu , Sebastien Bubeck , Ronen Eldan , Janardhan Kulkarni , Yuanzhi Li , Anh Nguyen , Rachel Ward , Yi Zhang

The rapid development of open-source large language models (LLMs) has been truly remarkable. However, the scaling law described in previous literature presents varying conclusions, which casts a dark cloud over scaling LLMs. We delve into…

Word Sense Disambiguation (WSD) remains a key challenge in Natural Language Processing (NLP), especially when dealing with rare or domain-specific senses that are often misinterpreted. While modern high-parameter Large Language Models…

Computation and Language · Computer Science 2026-03-06 Deshan Sumanathilaka , Nicholas Micallef , Julian Hough

Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still requires…