English
Related papers

Related papers: INTELLECT-1 Technical Report

200 papers

As large language models (LLMs) are increasingly integrated into emotionally sensitive domains, the structural integrity of their emotional intelligence (EI) becomes a critical frontier for safety and alignment. Current benchmarks often…

Artificial Intelligence · Computer Science 2026-05-26 Minghao Lv , Lu Chen , Enchang Zhang , Anji Zhou , Xiaoran Xue , Hanyi Zhang , Fenghua Tang , Zhuo Rachel Han , Mengyue Wu

Training machine learning models in parallel is an increasingly important workload. We accelerate distributed parallel training by designing a communication primitive that uses a programmable switch dataplane to execute a key step of the…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-10-01 Amedeo Sapio , Marco Canini , Chen-Yu Ho , Jacob Nelson , Panos Kalnis , Changhoon Kim , Arvind Krishnamurthy , Masoud Moshref , Dan R. K. Ports , Peter Richtárik

We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond…

Machine Learning · Computer Science 2026-04-03 Yicheng Zou , Dongsheng Zhu , Lin Zhu , Tong Zhu , Yunhua Zhou , Peiheng Zhou , Xinyu Zhou , Dongzhan Zhou , Zhiwang Zhou , Yuhao Zhou , Bowen Zhou , Zhanping Zhong , Zhijie Zhong , Haiteng Zhao , Penghao Zhao , Xiaomeng Zhao , Zhiyuan Zhao , Yechen Zhang , Jin Zhang , Wenwei Zhang , Hongjie Zhang , Zhuo Zhang , Wenlong Zhang , Bo Zhang , Chao Zhang , Chen Zhang , Yuhang Zang , Fei Yuan , Jiakang Yuan , Jiashuo Yu , Jinhui Yin , Haochen Ye , Qian Yao , Bowen Yang , Danni Yang , Kaichen Yang , Ziang Yan , Jun Xu , Yicheng Xu , Wanghan Xu , Xuenan Xu , Chao Xu , Ruiliang Xu , Shuhao Xing , Long Xing , Xinchen Xie , Ling-I Wu , Zijian Wu , Zhenyu Wu , Lijun Wu , Yue Wu , Jianyu Wu , Wen Wu , Fan Wu , Xilin Wei , Qi Wei , Bingli Wang , Rui Wang , Ziyi Wang , Zun Wang , Yi Wang , Haomin Wang , Yizhou Wang , Lintao Wang , Yiheng Wang , Longjiang Wang , Bin Wang , Jian Tong , Zhongbo Tian , Huanze Tang , Chen Tang , Shixiang Tang , Yu Sun , Qiushi Sun , Xuerui Su , Qisheng Su , Chenlin Su , Demin Song , Jin Shi , Fukai Shang , Yuchen Ren , Pengli Ren , Xiaoye Qu , Yuan Qu , Jiantao Qiu , Yu Qiao , Biqing Qi , Runyu Peng , Tianshuo Peng , Jiahui Peng , Qizhi Pei , Zhuoshi Pan , Linke Ouyang , Wenchang Ning , Yichuan Ma , Zerun Ma , Ningsheng Ma , Runyuan Ma , Chengqi Lyu , Haijun Lv , Han Lv , Lindong Lu , Kuikun Liu , Jiangning Liu , Yuhong Liu , Kai Liu , Hongwei Liu , Zhoumianze Liu , Mengjie Liu , Ziyu Liu , Wenran Liu , Yang Liu , Liwei Liu , Kaiwen Liu , Junyao Lin , Junming Lin , Tianyang Lin , Dahua Lin , Jianze Liang , Linyang Li , Peiji Li , Zonglin Li , Zehao Li , Pengze Li , Guoyan Li , Lingkai Kong , Linglin Jing , Zhenjiang Jin , Feifei Jiang , Qian Jiang , Junhao Huang , Zixian Huang , Haian Huang , Zhouqi Hua , Ermo Hua , Han Hu , Linfeng Hou , Yinan He , Conghui He , Tianyao He , Xu Guo , Qipeng Guo , Aijia Guo , Yuzhe Gu , Lixin Gu , Jingyang Gong , Qiming Ge , Jiaye Ge , Songyang Gao , Jianfei Gao , Xinyu Fang , Caihua fan , Yue Fan , Yanhui Duan , Zichen Ding , Shengyuan Ding , Ning Ding , Xuanlang Dai , Erfei Cui , Ganqu Cui , Pei Chu , Tao Chu , Guangran Cheng , Yu Cheng , Kai Chen , Yongkang Chen , Chiyu Chen , Guanzhou Chen , Qiaosheng Chen , Sitao Chen , Xin Chen , Haojiong Chen , Yicheng Chen , Weihan Cao , Yuhang Cao , Qinglong Cao , Lei Bai

Large Language Models (LLMs) have seen great advance in both academia and industry, and their popularity results in numerous open-source frameworks and techniques in accelerating LLM pre-training, fine-tuning, and inference. Training and…

Performance · Computer Science 2023-12-04 Longteng Zhang , Xiang Liu , Zeyu Li , Xinglin Pan , Peijie Dong , Ruibo Fan , Rui Guo , Xin Wang , Qiong Luo , Shaohuai Shi , Xiaowen Chu

Mixture-of-Experts (MoE) language models can reduce computational costs by 2-4$\times$ compared to dense models without sacrificing performance, making them more efficient in computation-bounded scenarios. However, MoE models generally…

Machine Learning · Computer Science 2024-04-09 Bowen Pan , Yikang Shen , Haokun Liu , Mayank Mishra , Gaoyuan Zhang , Aude Oliva , Colin Raffel , Rameswar Panda

Large language models are typically trained densely: all parameters are updated with respect to all inputs. This requires synchronization of billions of parameters across thousands of GPUs. We introduce a simple but effective method to…

Computation and Language · Computer Science 2023-03-27 Suchin Gururangan , Margaret Li , Mike Lewis , Weijia Shi , Tim Althoff , Noah A. Smith , Luke Zettlemoyer

Federated learning is a distributed machine learning approach where local weight parameters trained by clients locally are aggregated as global parameters by a server. The global parameters can be trained without uploading privacy-sensitive…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-12-19 Naoki Shibahara , Michihiro Koibuchi , Hiroki Matsutani

The workflow of pretraining and fine-tuning has emerged as a popular paradigm for solving various NLP and V&L (Vision-and-Language) downstream tasks. With the capacity of pretrained models growing rapidly, how to perform parameter-efficient…

Computation and Language · Computer Science 2022-03-09 Zhengkun Zhang , Wenya Guo , Xiaojun Meng , Yasheng Wang , Yadao Wang , Xin Jiang , Qun Liu , Zhenglu Yang

Communication-efficient distributed training algorithms have received considerable interest recently due to their benefits for training Large Language Models (LLMs) in bandwidth-constrained settings, such as across datacenters and over the…

Machine Learning · Computer Science 2025-11-07 Amir Sarfi , Benjamin Thérien , Joel Lidin , Eugene Belilovsky

Training Large Language Models(LLMs) is one of the most compute-intensive tasks in high-performance computing. Predicting end-to-end training time for multi-billion parameter models distributed across hundreds of GPUs remains challenging…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-30 Biyao Zhang , Mingkai Zheng , Debargha Ganguly , Xuecen Zhang , Vikash Singh , Vipin Chaudhary , Zhao Zhang

In this work, we study to release the potential of massive heterogeneous weak computing power to collaboratively train large-scale models on dispersed datasets. In order to improve both efficiency and accuracy in resource-adaptive…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-10-24 Yan Li , Xiao Zhang , Mingyi Li , Guangwei Xu , Feng Chen , Yuan Yuan , Yifei Zou , Mengying Zhao , Jianbo Lu , Dongxiao Yu

Federated Learning (FL) enables multiple clients to collaboratively train a shared model while preserving data privacy. However, the high memory demand during model training severely limits the deployment of FL on resource-constrained…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-10-14 Yebo Wu , Jingguang Li , Chunlin Tian , Kahou Tam , Li Li , Chengzhong Xu

Scaling model parameters improves model quality at the price of high computation overhead. Sparsely activated models, usually in the form of Mixture of Experts (MoE) architecture, have sub-linear scaling of computation cost with model size,…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-04-30 Jiamin Li , Yimin Jiang , Yibo Zhu , Cong Wang , Hong Xu

Collaborative learning across heterogeneous model architectures presents significant challenges in ensuring interoperability and preserving privacy. We propose a communication-efficient distributed learning framework that supports model…

Machine Learning · Computer Science 2025-09-30 Mounssif Krouka , Mehdi Bennis

Billions of dollars are lost every year in DeFi platforms by transactions exploiting business logic or accounting vulnerabilities. Existing defenses focus on static code analysis, public mempool screening, attacker contract detection, or…

Cryptography and Security · Computer Science 2025-10-21 Abdulrahman Alhaidari , Balaji Palanisamy , Prashant Krishnamurthy

Efficiently deploying large language models (LLMs) in real-world scenarios remains a critical challenge, primarily due to hardware heterogeneity, inference framework limitations, and workload complexities.Efficiently deploying large…

Artificial Intelligence · Computer Science 2025-01-28 Yanyu Chen , Ganhong Huang

Creating multilingual LLMs poses a significant challenge. Pretraining or fine-tuning LLMs to adopt new languages is evidently very costly. Furthermore, there exist limitations concerning benchmark datasets and the metrics used to measure…

Computation and Language · Computer Science 2024-04-08 Bibek Upadhayay , Vahid Behzadan

Frontier models increasingly adopt Mixture-of-Experts (MoE) architectures to achieve large-model performance at reduced cost. However, training MoE models on HPC platforms is hindered by large memory footprints, frequent large-scale…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-07 Sajal Dash , Feiyi Wang

Mixture-of-Experts (MoE) architectures have become standard in large language models, yet many of their core design choices - expert count, granularity, shared experts, load balancing, token dropping - have only been studied one or two at a…

Machine Learning · Computer Science 2026-05-13 Margaret Li , Sneha Kudugunta , Danielle Rothermel , Luke Zettlemoyer