English
Related papers

Related papers: GPT-NeoX-20B: An Open-Source Autoregressive Langua…

200 papers

We introduce BitNet b1.58 2B4T, the first open-source, native 1-bit Large Language Model (LLM) at the 2-billion parameter scale. Trained on a corpus of 4 trillion tokens, the model has been rigorously evaluated across benchmarks covering…

Computation and Language · Computer Science 2025-04-28 Shuming Ma , Hongyu Wang , Shaohan Huang , Xingxing Zhang , Ying Hu , Ting Song , Yan Xia , Furu Wei

The training of large language models (LLMs) requires substantial computational resources, complex software stacks, and carefully designed workflows to achieve scalability and efficiency. This report presents best practices and insights…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-08 Carolin Penke , Chelsea Maria John , Jan Ebert , Stefan Kesselheim , Andreas Herten

Groundbreaking language-vision architectures like CLIP and DALL-E proved the utility of training on large amounts of noisy image-text data, without relying on expensive accurate labels used in standard vision unimodal supervised learning.…

We introduce Nemotron-4 15B, a 15-billion-parameter large multilingual language model trained on 8 trillion text tokens. Nemotron-4 15B demonstrates strong performance when assessed on English, multilingual, and coding tasks: it outperforms…

Large language models (LLMs) have recently achieved human-level performance on a range of professional and academic benchmarks. The accessibility of these models has lagged behind their performance. State-of-the-art LLMs require costly…

Computation and Language · Computer Science 2023-11-10 Yuvanesh Anand , Zach Nussbaum , Adam Treat , Aaron Miller , Richard Guo , Ben Schmidt , GPT4All Community , Brandon Duderstadt , Andriy Mulyar

As social-media platforms emerge and evolve faster than the regulations meant to oversee them, automated detoxification might serve as a timely tool for moderators to enforce safe discourse at scale. We here describe our submission to the…

Computation and Language · Computer Science 2026-02-03 Trung Duc Anh Dang , Ferdinando Pio D'Elia

Large language models have a range of beneficial uses: they can assist in prose, poetry, and programming; analyze dataset biases; and more. However, their flexibility and generative capabilities also raise misuse concerns. This report…

Large Language Models (LLMs) have made significant strides in natural language generation but often face challenges in tasks requiring precise calculations and structural analysis. This paper investigates the performance of state-of-the-art…

Computation and Language · Computer Science 2025-02-18 Birger Moell , Johan Boye

Recently, significant public efforts have been directed towards developing low-cost models with capabilities akin to ChatGPT, thereby fostering the growth of open-source conversational models. However, there remains a scarcity of…

Computation and Language · Computer Science 2023-04-18 Yunjie Ji , Yan Gong , Yong Deng , Yiping Peng , Qiang Niu , Baochang Ma , Xiangang Li

GPT series models, such as GPT-3, CodeX, InstructGPT, ChatGPT, and so on, have gained considerable attention due to their exceptional natural language processing capabilities. However, despite the abundance of research on the difference in…

Computation and Language · Computer Science 2023-12-27 Junjie Ye , Xuanting Chen , Nuo Xu , Can Zu , Zekai Shao , Shichun Liu , Yuhan Cui , Zeyang Zhou , Chao Gong , Yang Shen , Jie Zhou , Siming Chen , Tao Gui , Qi Zhang , Xuanjing Huang

In recent years, pretrained models have been widely used in various fields, including natural language understanding, computer vision, and natural language generation. However, the performance of these language generation models is highly…

Computation and Language · Computer Science 2023-04-14 Zhengqing Yuan , Huiwen Xue , Chao Zhang , Yongming Liu

Large language models (LLMs) with billions of parameters have demonstrated outstanding performance on various natural language processing tasks. This report presents OpenBA, an open-sourced 15B bilingual asymmetric seq2seq model, to…

Computation and Language · Computer Science 2024-11-26 Juntao Li , Zecheng Tang , Yuyang Ding , Pinzheng Wang , Pei Guo , Wangjie You , Dan Qiao , Wenliang Chen , Guohong Fu , Qiaoming Zhu , Guodong Zhou , Min Zhang

Recently, we have witnessed great progress in image editing with natural language instructions. Several closed-source models like GPT-Image-1, Seedream, and Google-Nano-Banana have shown highly promising progress. However, the open-source…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Keming Wu , Sicong Jiang , Max Ku , Ping Nie , Minghao Liu , Wenhu Chen

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 are as follows: (1) DeepSeek Sparse Attention (DSA): We…

Computation and Language · Computer Science 2025-12-03 DeepSeek-AI , Aixin Liu , Aoxue Mei , Bangcai Lin , Bing Xue , Bingxuan Wang , Bingzheng Xu , Bochao Wu , Bowei Zhang , Chaofan Lin , Chen Dong , Chengda Lu , Chenggang Zhao , Chengqi Deng , Chenhao Xu , Chong Ruan , Damai Dai , Daya Guo , Dejian Yang , Deli Chen , Erhang Li , Fangqi Zhou , Fangyun Lin , Fucong Dai , Guangbo Hao , Guanting Chen , Guowei Li , H. Zhang , Hanwei Xu , Hao Li , Haofen Liang , Haoran Wei , Haowei Zhang , Haowen Luo , Haozhe Ji , Honghui Ding , Hongxuan Tang , Huanqi Cao , Huazuo Gao , Hui Qu , Hui Zeng , Jialiang Huang , Jiashi Li , Jiaxin Xu , Jiewen Hu , Jingchang Chen , Jingting Xiang , Jingyang Yuan , Jingyuan Cheng , Jinhua Zhu , Jun Ran , Junguang Jiang , Junjie Qiu , Junlong Li , Junxiao Song , Kai Dong , Kaige Gao , Kang Guan , Kexin Huang , Kexing Zhou , Kezhao Huang , Kuai Yu , Lean Wang , Lecong Zhang , Lei Wang , Liang Zhao , Liangsheng Yin , Lihua Guo , Lingxiao Luo , Linwang Ma , Litong Wang , Liyue Zhang , M. S. Di , M. Y Xu , Mingchuan Zhang , Minghua Zhang , Minghui Tang , Mingxu Zhou , Panpan Huang , Peixin Cong , Peiyi Wang , Qiancheng Wang , Qihao Zhu , Qingyang Li , Qinyu Chen , Qiushi Du , Ruiling Xu , Ruiqi Ge , Ruisong Zhang , Ruizhe Pan , Runji Wang , Runqiu Yin , Runxin Xu , Ruomeng Shen , Ruoyu Zhang , S. H. Liu , Shanghao Lu , Shangyan Zhou , Shanhuang Chen , Shaofei Cai , Shaoyuan Chen , Shengding Hu , Shengyu Liu , Shiqiang Hu , Shirong Ma , Shiyu Wang , Shuiping Yu , Shunfeng Zhou , Shuting Pan , Songyang Zhou , Tao Ni , Tao Yun , Tian Pei , Tian Ye , Tianyuan Yue , Wangding Zeng , Wen Liu , Wenfeng Liang , Wenjie Pang , Wenjing Luo , Wenjun Gao , Wentao Zhang , Xi Gao , Xiangwen Wang , Xiao Bi , Xiaodong Liu , Xiaohan Wang , Xiaokang Chen , Xiaokang Zhang , Xiaotao Nie , Xin Cheng , Xin Liu , Xin Xie , Xingchao Liu , Xingkai Yu , Xingyou Li , Xinyu Yang , Xinyuan Li , Xu Chen , Xuecheng Su , Xuehai Pan , Xuheng Lin , Xuwei Fu , Y. Q. Wang , Yang Zhang , Yanhong Xu , Yanru Ma , Yao Li , Yao Li , Yao Zhao , Yaofeng Sun , Yaohui Wang , Yi Qian , Yi Yu , Yichao Zhang , Yifan Ding , Yifan Shi , Yiliang Xiong , Ying He , Ying Zhou , Yinmin Zhong , Yishi Piao , Yisong Wang , Yixiao Chen , Yixuan Tan , Yixuan Wei , Yiyang Ma , Yiyuan Liu , Yonglun Yang , Yongqiang Guo , Yongtong Wu , Yu Wu , Yuan Cheng , Yuan Ou , Yuanfan Xu , Yuduan Wang , Yue Gong , Yuhan Wu , Yuheng Zou , Yukun Li , Yunfan Xiong , Yuxiang Luo , Yuxiang You , Yuxuan Liu , Yuyang Zhou , Z. F. Wu , Z. Z. Ren , Zehua Zhao , Zehui Ren , Zhangli Sha , Zhe Fu , Zhean Xu , Zhenda Xie , Zhengyan Zhang , Zhewen Hao , Zhibin Gou , Zhicheng Ma , Zhigang Yan , Zhihong Shao , Zhixian Huang , Zhiyu Wu , Zhuoshu Li , Zhuping Zhang , Zian Xu , Zihao Wang , Zihui Gu , Zijia Zhu , Zilin Li , Zipeng Zhang , Ziwei Xie , Ziyi Gao , Zizheng Pan , Zongqing Yao , Bei Feng , Hui Li , J. L. Cai , Jiaqi Ni , Lei Xu , Meng Li , Ning Tian , R. J. Chen , R. L. Jin , S. S. Li , Shuang Zhou , Tianyu Sun , X. Q. Li , Xiangyue Jin , Xiaojin Shen , Xiaosha Chen , Xinnan Song , Xinyi Zhou , Y. X. Zhu , Yanping Huang , Yaohui Li , Yi Zheng , Yuchen Zhu , Yunxian Ma , Zhen Huang , Zhipeng Xu , Zhongyu Zhang , Dongjie Ji , Jian Liang , Jianzhong Guo , Jin Chen , Leyi Xia , Miaojun Wang , Mingming Li , Peng Zhang , Ruyi Chen , Shangmian Sun , Shaoqing Wu , Shengfeng Ye , T. Wang , W. L. Xiao , Wei An , Xianzu Wang , Xiaowen Sun , Xiaoxiang Wang , Ying Tang , Yukun Zha , Zekai Zhang , Zhe Ju , Zhen Zhang , Zihua Qu

Transfer learning from pretrained language models recently became the dominant approach for solving many NLP tasks. A common approach to transfer learning for multiple tasks that maximize parameter sharing trains one or more task-specific…

Computation and Language · Computer Science 2021-06-03 Karen Hambardzumyan , Hrant Khachatrian , Jonathan May

Large Language Models are traditionally finetuned on large instruction datasets. However recent studies suggest that small, high-quality datasets can suffice for general purpose instruction following. This lack of consensus surrounding…

Machine Learning · Computer Science 2023-12-29 Aditi Jha , Sam Havens , Jeremy Dohmann , Alex Trott , Jacob Portes

Since the release of ChatGPT in November 2022, large language models (LLMs) have seen considerable success, including in the open-source community, with many open-weight models available. However, the requirements to deploy such a service…

Performance · Computer Science 2025-06-13 Yannis Bendi-Ouis , Dan Dutartre , Xavier Hinaut

Parameter-efficient fine-tuning (PEFT) methods have provided an effective way for adapting large vision-language models to specific tasks or scenarios. Typically, they learn a very small scale of parameters for pre-trained models in a…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Zixian Guo , Yuxiang Wei , Ming Liu , Zhilong Ji , Jinfeng Bai , Yiwen Guo , Wangmeng Zuo

We introduce Dream 7B, the most powerful open diffusion large language model to date. Unlike autoregressive (AR) models that generate tokens sequentially, Dream 7B employs discrete diffusion modeling to refine sequences in parallel through…

Computation and Language · Computer Science 2025-08-22 Jiacheng Ye , Zhihui Xie , Lin Zheng , Jiahui Gao , Zirui Wu , Xin Jiang , Zhenguo Li , Lingpeng Kong

Large language models (LLMs) have shown impressive ability for open-domain NLP tasks. However, LLMs are sometimes too footloose for natural language understanding (NLU) tasks which always have restricted output and input format. Their…