Computation and Language · Computer Science
Distilling Token-Trained Models into Byte-Level Models
Zishuo Bao, Jiaqi Leng, Junxiong Wang, Bowen Peng +1
2026-02-03
Computation and Language · Computer Science
Alexa Teacher Model: Pretraining and Distilling Multi-Billion-Parameter Encoders for Natural Language Understanding Systems
Jack FitzGerald, Shankar Ananthakrishnan, Konstantine Arkoudas, Davide Bernardi +37
2022-06-17
Computation and Language · Computer Science
DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
Jongwoo Ko, Tianyi Chen, Sungnyun Kim, Tianyu Ding +3
2025-06-02
Machine Learning · Computer Science
Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling
Sachin Goyal, David Lopez-Paz, Kartik Ahuja
2025-09-03
Computation and Language · Computer Science
Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost +5
2023-07-06
Computation and Language · Computer Science
DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training
Xiaoyu Tian, Sitong Zhao, Haotian Wang, Shuaiting Chen +4
2025-05-14
Computation and Language · Computer Science
Effective Distillation of Table-based Reasoning Ability from LLMs
Bohao Yang, Chen Tang, Kun Zhao, Chenghao Xiao +1
2024-03-26
Computation and Language · Computer Science
Benchmarking Distilled Language Models: Performance and Efficiency in Resource-Constrained Settings
Sachin Gopal Wani, Eric Page, Ajay Dholakia, David Ellison
2026-02-25
Computation and Language · Computer Science
DiLM: Distilling Dataset into Language Model for Text-level Dataset Distillation
Aru Maekawa, Satoshi Kosugi, Kotaro Funakoshi, Manabu Okumura
2024-04-02
Machine Learning · Computer Science
Knowledge Distillation Using Frontier Open-source LLMs: Generalizability and the Role of Synthetic Data
Anup Shirgaonkar, Nikhil Pandey, Nazmiye Ceren Abay, Tolga Aktas +1
2024-10-25
Computation and Language · Computer Science
Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
Michael Y. Hu, Aaron Mueller, Candace Ross, Adina Williams +6
2024-12-09
Computation and Language · Computer Science
Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?
Yudi Zhang, Lu Wang, Meng Fang, Yali Du +7
2025-02-28
Computation and Language · Computer Science
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
Jaehun Jung, Peter West, Liwei Jiang, Faeze Brahman +4
2024-08-21
Computation and Language · Computer Science
Are BabyLMs Second Language Learners?
Lukas Edman, Lisa Bylinina, Faeze Ghorbanpour, Alexander Fraser
2024-10-29
Artificial Intelligence · Computer Science
Pedagogically-Inspired Data Synthesis for Language Model Knowledge Distillation
Bowei He, Yankai Chen, Xiaokun Zhang, Linghe Kong +3
2026-02-13
Computation and Language · Computer Science
Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge Distillation
Yijun Tian, Yikun Han, Xiusi Chen, Wei Wang +1
2024-11-26