English
Related papers

Related papers: LongCat-Flash Technical Report

200 papers

Industrial recommender systems critically depend on high-quality ranking models. However, traditional pipelines still rely on manual feature engineering and scenario-specific architectures, which hinder cross-scenario transfer and…

Information Retrieval · Computer Science 2025-10-20 Xianyang Qi , Yuan Tian , Zhaoyu Hu , Zhirui Kuai , Chang Liu , Hongxiang Lin , Lei Wang

Scaling large language models has driven remarkable advancements across various domains, yet the continual increase in model size presents significant challenges for real-world deployment. The Mixture of Experts (MoE) architecture offers a…

Machine Learning · Computer Science 2025-03-18 Shwai He , Daize Dong , Liang Ding , Ang Li

The proliferation of large language models (LLMs) has driven the adoption of Mixture-of-Experts (MoE) architectures as a promising solution to scale model capacity while controlling computational costs. However, deploying MoE models in…

Networking and Internet Architecture · Computer Science 2025-08-14 Muqing Li , Ning Li , Xin Yuan , Wenchao Xu , Quan Chen , Song Guo , Haijun Zhang

Mixture-of-Experts large language models (MoE-LLMs) marks a significant step forward of language models, however, they encounter two critical challenges in practice: 1) expert parameters lead to considerable memory consumption and loading…

Machine Learning · Computer Science 2025-02-25 Wei Huang , Yue Liao , Jianhui Liu , Ruifei He , Haoru Tan , Shiming Zhang , Hongsheng Li , Si Liu , Xiaojuan Qi

Large language models allocate uniform computation across all tokens, ignoring that some sequences are trivially predictable while others require deep reasoning. We introduce ConceptMoE, which dynamically merges semantically similar tokens…

Machine Learning · Computer Science 2026-01-30 Zihao Huang , Jundong Zhou , Xingwei Qu , Qiyang Min , Ge Zhang

Mixture of Experts (MoE) models have emerged as a promising paradigm for scaling language models efficiently by activating only a subset of parameters for each input token. In this report, we present dots.llm1, a large-scale MoE model that…

We present IntelliCAT, an interactive translation interface with neural models that streamline the post-editing process on machine translation output. We leverage two quality estimation (QE) models at different granularities: sentence-level…

Computation and Language · Computer Science 2021-05-27 Dongjun Lee , Junhyeong Ahn , Heesoo Park , Jaemin Jo

Scaling large language models (LLMs) significantly improves performance but comes with prohibitive computational costs. Mixture-of-Experts (MoE) models offer an efficient alternative, increasing capacity without a proportional rise in…

Machine Learning · Computer Science 2024-12-16 Aditya Vavre , Ethan He , Dennis Liu , Zijie Yan , June Yang , Nima Tajbakhsh , Ashwath Aithal

We introduce Kimi K2, a Mixture-of-Experts (MoE) large language model with 32 billion activated parameters and 1 trillion total parameters. We propose the MuonClip optimizer, which improves upon Muon with a novel QK-clip technique to…

Machine Learning · Computer Science 2026-02-04 Kimi Team , Yifan Bai , Yiping Bao , Y. Charles , Cheng Chen , Guanduo Chen , Haiting Chen , Huarong Chen , Jiahao Chen , Ningxin Chen , Ruijue Chen , Yanru Chen , Yuankun Chen , Yutian Chen , Zhuofu Chen , Jialei Cui , Hao Ding , Mengnan Dong , Angang Du , Chenzhuang Du , Dikang Du , Yulun Du , Yu Fan , Yichen Feng , Kelin Fu , Bofei Gao , Chenxiao Gao , Hongcheng Gao , Peizhong Gao , Tong Gao , Yuyao Ge , Shangyi Geng , Qizheng Gu , Xinran Gu , Longyu Guan , Haiqing Guo , Jianhang Guo , Xiaoru Hao , Tianhong He , Weiran He , Wenyang He , Yunjia He , Chao Hong , Hao Hu , Yangyang Hu , Zhenxing Hu , Weixiao Huang , Zhiqi Huang , Zihao Huang , Tao Jiang , Zhejun Jiang , Xinyi Jin , Yongsheng Kang , Guokun Lai , Cheng Li , Fang Li , Haoyang Li , Ming Li , Wentao Li , Yang Li , Yanhao Li , Yiwei Li , Zhaowei Li , Zheming Li , Hongzhan Lin , Xiaohan Lin , Zongyu Lin , Chengyin Liu , Chenyu Liu , Hongzhang Liu , Jingyuan Liu , Junqi Liu , Liang Liu , Shaowei Liu , T. Y. Liu , Tianwei Liu , Weizhou Liu , Yangyang Liu , Yibo Liu , Yiping Liu , Yue Liu , Zhengying Liu , Enzhe Lu , Haoyu Lu , Lijun Lu , Yashuo Luo , Shengling Ma , Xinyu Ma , Yingwei Ma , Shaoguang Mao , Jie Mei , Xin Men , Yibo Miao , Siyuan Pan , Yebo Peng , Ruoyu Qin , Zeyu Qin , Bowen Qu , Zeyu Shang , Lidong Shi , Shengyuan Shi , Feifan Song , Jianlin Su , Zhengyuan Su , Lin Sui , Xinjie Sun , Flood Sung , Yunpeng Tai , Heyi Tang , Jiawen Tao , Qifeng Teng , Chaoran Tian , Chensi Wang , Dinglu Wang , Feng Wang , Hailong Wang , Haiming Wang , Jianzhou Wang , Jiaxing Wang , Jinhong Wang , Shengjie Wang , Shuyi Wang , Si Wang , Xinyuan Wang , Yao Wang , Yejie Wang , Yiqin Wang , Yuxin Wang , Yuzhi Wang , Zhaoji Wang , Zhengtao Wang , Zhengtao Wang , Zhexu Wang , Chu Wei , Qianqian Wei , Haoning Wu , Wenhao Wu , Xingzhe Wu , Yuxin Wu , Chenjun Xiao , Jin Xie , Xiaotong Xie , Weimin Xiong , Boyu Xu , Jinjing Xu , L. H. Xu , Lin Xu , Suting Xu , Weixin Xu , Xinran Xu , Yangchuan Xu , Ziyao Xu , Jing Xu , Jing Xu , Junjie Yan , Yuzi Yan , Hao Yang , Xiaofei Yang , Yi Yang , Ying Yang , Zhen Yang , Zhilin Yang , Zonghan Yang , Haotian Yao , Xingcheng Yao , Wenjie Ye , Zhuorui Ye , Bohong Yin , Longhui Yu , Enming Yuan , Hongbang Yuan , Mengjie Yuan , Siyu Yuan , Haobing Zhan , Dehao Zhang , Hao Zhang , Wanlu Zhang , Xiaobin Zhang , Yadong Zhang , Yangkun Zhang , Yichi Zhang , Yizhi Zhang , Yongting Zhang , Yu Zhang , Yutao Zhang , Yutong Zhang , Zheng Zhang , Haotian Zhao , Yikai Zhao , Zijia Zhao , Huabin Zheng , Shaojie Zheng , Longguang Zhong , Jianren Zhou , Xinyu Zhou , Zaida Zhou , Jinguo Zhu , Zhen Zhu , Weiyu Zhuang , Xinxing Zu

Mixture-of-Experts (MoE) models, the state-of-the-art in large-scale AI, achieve high quality by sparsely activating parameters. However, their reliance on routing between a few monolithic experts via a top-k mechanism creates a "quality…

Computation and Language · Computer Science 2025-10-23 Xinfeng Xia , Jiacheng Liu , Xiaofeng Hou , Peng Tang , Mingxuan Zhang , Wenfeng Wang , Chao Li

The Mixture of Experts (MoE) selects a few feed-forward networks (FFNs) per token, achieving an effective trade-off between computational cost and performance. In conventional MoE, each expert is treated as entirely independent, and experts…

Machine Learning · Computer Science 2026-01-27 Shota Takashiro , Takeshi Kojima , Shohei Taniguchi , Yusuke Iwasawa , Yutaka Matsuo

We propose Tensor-Trained Low-Rank Adaptation Mixture of Experts (TT-LoRA MoE), a novel computational framework integrating Parameter-Efficient Fine-Tuning (PEFT) with sparse MoE routing to address scalability challenges in large model…

Machine Learning · Computer Science 2026-01-27 Pradip Kunwar , Minh N. Vu , Maanak Gupta , Mahmoud Abdelsalam , Manish Bhattarai

In this technical report, we introduce the training methodologies implemented in the development of Skywork-MoE, a high-performance mixture-of-experts (MoE) large language model (LLM) with 146 billion parameters and 16 experts. It is…

Attention mechanisms, primarily designed to capture pairwise correlations between words, have become the backbone of machine learning, expanding beyond natural language processing into other domains. This growth in adaptation comes at the…

Machine Learning · Computer Science 2022-09-27 Sheng-Chun Kao , Suvinay Subramanian , Gaurav Agrawal , Amir Yazdanbakhsh , Tushar Krishna

In this work, we aim to simultaneously enhance the effectiveness and efficiency of Mixture-of-Experts (MoE) methods. To achieve this, we propose MoE++, a general and heterogeneous MoE framework that integrates both Feed-Forward…

Machine Learning · Computer Science 2024-10-11 Peng Jin , Bo Zhu , Li Yuan , Shuicheng Yan

Mixture-of-Experts (MoE) models have gained popularity as a means of scaling the capacity of large language models (LLMs) while maintaining sparse activations and reduced per-token compute. However, in memory-constrained inference settings,…

Machine Learning · Computer Science 2026-03-23 Vivan Madan , Prajwal Singhania , Abhinav Bhatele , Tom Goldstein , Ashwinee Panda

Mixture-of-Expert (MoE) based large language models (LLMs), such as the recent Mixtral and DeepSeek-MoE, have shown great promise in scaling model size without suffering from the quadratic growth of training cost of dense transformers. Like…

Machine Learning · Computer Science 2024-04-04 Longfei Yun , Yonghao Zhuang , Yao Fu , Eric P Xing , Hao Zhang

Research on LLM technologies is rapidly emerging, with most of them employ a 'fast thinking' approach to inference. Most LLMs generate the final result based solely on a single query and LLM's reasoning capabilities. However, with the…

Computation and Language · Computer Science 2025-11-14 Jianfeng Pan , Senyou Deng , Shaomang Huang

Larger transformer models always perform better on various tasks but require more costs to scale up the model size. To efficiently enlarge models, the mixture-of-experts (MoE) architecture is widely adopted, which consists of a gate network…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-11-14 Xiaonan Nie , Qibin Liu , Fangcheng Fu , Shenhan Zhu , Xupeng Miao , Xiaoyang Li , Yang Zhang , Shouda Liu , Bin Cui

Most recent state-of-the-art (SOTA) large language models (LLMs) use Mixture-of-Experts (MoE) architectures to scale model capacity without proportional per-token compute, enabling higher-quality outputs at manageable serving costs.…

‹ Prev 1 4 5 6 7 8 10 Next ›