English
Related papers

Related papers: Holistic Capability Preservation: Towards Compact …

200 papers

Large Language Models (LLMs) demonstrate exceptional reasoning capabilities, often achieving state-of-the-art performance in various tasks. However, their substantial computational and memory demands, due to billions of parameters, hinder…

Computation and Language · Computer Science 2024-11-25 Xunyu Zhu , Jian Li , Can Ma , Weiping Wang

Large Language Models (LLMs) have shown promising performance in knowledge-intensive reasoning tasks that require a compound understanding of knowledge. However, deployment of the LLMs in real-world applications can be challenging due to…

Computation and Language · Computer Science 2023-10-31 Minki Kang , Seanie Lee , Jinheon Baek , Kenji Kawaguchi , Sung Ju Hwang

Improving performance on complex tasks and enabling interpretable decision making in large language models (LLMs), especially for clinical applications, requires effective reasoning. Yet this remains challenging without supervised…

Computation and Language · Computer Science 2025-05-26 Che Liu , Haozhe Wang , Jiazhen Pan , Zhongwei Wan , Yong Dai , Fangzhen Lin , Wenjia Bai , Daniel Rueckert , Rossella Arcucci

Current long chain-of-thought (long-CoT) models excel at mathematical reasoning but rely on slow and error-prone natural language traces. Tool-augmented agents address arithmetic via code execution, but often falter on complex logical…

Computation and Language · Computer Science 2025-09-03 Weihua Du , Pranjal Aggarwal , Sean Welleck , Yiming Yang

Recent advancements in Large Language Models (LLMs) have revealed a significant performance gap between closed-source and open-source models, particularly in tasks requiring complex reasoning and precise instruction following. This paper…

Artificial Intelligence · Computer Science 2025-07-01 Ziqi Zhong , Xunzhu Tang

The paradigm shift in large language models (LLMs) from instinctive responses to chain-of-thought (CoT) reasoning has fueled two prevailing assumptions: (1) reasoning capabilities only emerge in sufficiently large models, and (2) such…

Pre-trained Language Models (LMs) have become an integral part of Natural Language Processing (NLP) in recent years, due to their superior performance in downstream applications. In spite of this resounding success, the usability of LMs is…

Computation and Language · Computer Science 2023-05-02 Mohammadmahdi Nouriborji , Omid Rohanian , Samaneh Kouchaki , David A. Clifton

We introduce Ling 2.0, a series reasoning-oriented language foundation built upon the principle that every activation boosts reasoning capability. Designed to scale from tens of billions to one trillion parameters under a unified…

Computation and Language · Computer Science 2025-11-10 Ling Team , Ang Li , Ben Liu , Binbin Hu , Bing Li , Bingwei Zeng , Borui Ye , Caizhi Tang , Changxin Tian , Chao Huang , Chao Zhang , Chen Qian , Chenchen Ju , Chenchen Li , Chengfu Tang , Chilin Fu , Chunshao Ren , Chunwei Wu , Cong Zhang , Cunyin Peng , Dafeng Xu , Daixin Wang , Dalong Zhang , Dingnan Jin , Dingyuan Zhu , Dongke Hu , Fangzheng Zhao , Feifan Wu , Feng Zhu , Gangshan Wang , Haitao Zhang , Hailin Zhao , Hanxiao Zhang , Hanzi Wang , Hao Qian , Haoyi Yu , Heng Zhang , Hongliang Zhang , Hongzhi Luan , Huirong Dong , Huizhong Li , Jia Li , Jia Liu , Jialong Zhu , Jian Sha , Jianping Wei , Jiaolong Yang , Jieyue Ma , Jiewei Wu , Jinjing Huang , Jingyun Tian , Jingyuan Zhang , Jinquan Sun , Juanhui Tu , Jun Liu , Jun Xu , Jun Zhou , Junjie Ou , Junpeng Fang , Kaihong Zhang , Kaiqin Hu , Ke Shi , Kun Tang , Kunlong Chen , Lanyin Mei , Lei Liang , Lei Xu , Libo Zhang , Lin Ju , Lin Yuan , Ling Zhong , Lintao Ma , Lu Liu , Lu Yu , Lun Cai , Meiqi Zhu , Mengying Li , Min Chen , Minghao Xue , Minghong Cai , Mingming Yin , Peijie Jiang , Peilong Zhao , Pingping Liu , Qian Zhao , Qing Cui , Qingxiang Huang , Qingyuan Yang , Quankun Yu , Shaowei Wei , Shijie Lian , Shoujian Zheng , Shun Song , Shungen Zhang , Shuo Zhang , Siyuan Li , Song Liu , Ting Guo , Tong Zhao , Wanli Gu , Weichang Wu , Weiguang Han , Wenjing Fang , Wubin Wang , Xiang Shu , Xiao Shi , Xiaoshun Lan , Xiaolu Zhang , Xiaqing Sun , Xin Zhao , Xingyu Lu , Xiong Xu , Xudong Wang , Xudong Wang , Xuemin Yang , Yajie Yang , Yang Xiang , Yanzhe Li , Yi Zhang , Yilong Wang , Yingxue Li , Yongzhen Guo , Yuzhuo Fu , Yuanyuan Wang , Yue Yang , Yue Yu , Yufeng Deng , Yun Zhang , Yunfei Yu , Yuqi Zhang , Yuxiao He , Zengke Gui , Zhaoxin Huan , Zhaoyang Wang , Zhibo Zhu , Zhihao Wang , Zhiqiang Zhang , Zhoufei Wang , Zihang Zeng , Ziqi Liu , Zitao Xuan , Zuoli Tang

In this paper, we present EasyDistill, a comprehensive toolkit designed for effective black-box and white-box knowledge distillation (KD) of large language models (LLMs). Our framework offers versatile functionalities, including data…

Computation and Language · Computer Science 2025-06-30 Chengyu Wang , Junbing Yan , Wenrui Cai , Yuanhao Yue , Jun Huang

Reasoning-centric large language models (LLMs) achieve strong performance by generating intermediate reasoning trajectories, but often incur excessive token usage and high inference-time decoding cost. We observe that, when solving the same…

Artificial Intelligence · Computer Science 2026-05-12 Han Yang , Mingyan Wu , Bailan He , Zeyu Cao , Sikuan Yan , Kevin Qinghong Lin , Zifeng Ding

Large language models (LLMs) have transformed sentiment analysis, yet balancing accuracy, efficiency, and explainability remains a critical challenge. This study presents the first comprehensive evaluation of DeepSeek-R1--an open-source…

Computation and Language · Computer Science 2026-02-05 Donghao Huang , Zhaoxia Wang

The emergence of Large Audio-Language Models (LALMs) has advanced Speech Emotion Recognition (SER), but their size limits deployment in resource-constrained environments. While Knowledge Distillation is effective for LALM compression,…

Sparse Mixture-of-Experts (MoE) has been a successful approach for scaling multilingual translation models to billions of parameters without a proportional increase in training computation. However, MoE models are prohibitively large and…

Computation and Language · Computer Science 2021-10-11 Sneha Kudugunta , Yanping Huang , Ankur Bapna , Maxim Krikun , Dmitry Lepikhin , Minh-Thang Luong , Orhan Firat

Large language models (LLMs), especially Explicit Long Chain-of-Thought (CoT) reasoning models like DeepSeek-R1 and QWQ, have demonstrated powerful reasoning capabilities, achieving impressive performance in commonsense reasoning and…

Computation and Language · Computer Science 2025-08-13 Jiatong Li , Weida Wang , Qinggang Zhang , Junxian Li , Di Zhang , Changmeng Zheng , Shufei Zhang , Xiaoyong Wei , Qing Li

Reasoning segmentation enables open-set object segmentation via implicit text queries, therefore serving as a foundation for embodied agents that should operate autonomously in real-world environments. However, existing methods for…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Yiqing Shen , Mathias Unberath

In this paper, we present BitNet Distillation (BitDistill), a lightweight pipeline that fine-tunes off-the-shelf full-precision LLMs (e.g., Qwen) into 1.58-bit precision (i.e., ternary weights {-1, 0, 1}) for specific downstream tasks,…

Machine Learning · Computer Science 2025-10-17 Xun Wu , Shaohan Huang , Wenhui Wang , Ting Song , Li Dong , Yan Xia , Furu Wei

Lipreading has witnessed a lot of progress due to the resurgence of neural networks. Recent works have placed emphasis on aspects such as improving performance by finding the optimal architecture or improving generalization. However, there…

Computer Vision and Pattern Recognition · Computer Science 2021-06-03 Pingchuan Ma , Brais Martinez , Stavros Petridis , Maja Pantic

Existing chain-of-thought (CoT) distillation methods can effectively transfer reasoning abilities to base models but suffer from two major limitations: excessive verbosity of reasoning traces and inadequate adaptability to problem…

Artificial Intelligence · Computer Science 2025-05-27 Yifan Wu , Jingze Shi , Bingheng Wu , Jiayi Zhang , Xiaotian Lin , Nan Tang , Yuyu Luo

Language models often lack grounded reasoning capabilities in specialized domains where training data is scarce but bespoke systems excel. We introduce a general framework for distilling expert system reasoning into natural language…

Artificial Intelligence · Computer Science 2026-03-24 Zhenwei Tang , Qianfeng Wen , Seth Grief-Albert , Yahya Elgabra , Blair Yang , Honghua Dong , Ashton Anderson

Recent think-answer approaches in VLMs, such as Qwen3-VL-Thinking, boost reasoning performance by leveraging intermediate thinking steps before the final answer, but their computational cost becomes substantial, especially for larger VLMs.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Seonghoon Yu , Dongjun Nam , Byung-Kwan Lee , Jeany Son