English
Related papers

Related papers: Qwen2 Technical Report

200 papers

We introduce InternVL3, a significant advancement in the InternVL series featuring a native multimodal pre-training paradigm. Rather than adapting a text-only large language model (LLM) into a multimodal large language model (MLLM) that…

Large-scale Transformer models have significantly promoted the recent development of natural language processing applications. However, little effort has been made to unify the effective models. In this paper, driven by providing a new set…

Computation and Language · Computer Science 2022-04-12 Dezhou Shen

Large language models (LLMs) are evolving from conversational systems into strong reasoners for tasks such as Olympiad mathematics and competitive programming. While scaling parameters and test-time computation has driven progress, a key…

Machine Learning · Computer Science 2025-09-25 Xueliang Zhao , Wei Wu , Jian Guan , Zhuocheng Gong , Lingpeng Kong

We benchmark different strategies of adding new languages (German and Korean) into the BigScience's pretrained multilingual language model with 1.3 billion parameters that currently supports 13 languages. We investigate the factors that…

Computation and Language · Computer Science 2022-04-12 Zheng-Xin Yong , Vassilina Nikoulina

Large Language Models (LLMs) have showcased remarkable impacts across a wide spectrum of natural language processing tasks. Fine-tuning these pretrained models on downstream datasets provides further significant performance gains; however,…

Computation and Language · Computer Science 2026-03-19 Zhikai Li , Xiaoxuan Liu , Banghua Zhu , Zhen Dong , Qingyi Gu , Kurt Keutzer

Despite the major advances in NLP, significant disparities in NLP system performance across languages still exist. Arguably, these are due to uneven resource allocation and sub-optimal incentives to work on less resourced languages. To…

Effective pre-training of large language models (LLMs) has been challenging due to the immense resource demands and the complexity of the technical processes involved. This paper presents a detailed technical report on YuLan-Mini, a highly…

Computation and Language · Computer Science 2024-12-25 Yiwen Hu , Huatong Song , Jia Deng , Jiapeng Wang , Jie Chen , Kun Zhou , Yutao Zhu , Jinhao Jiang , Zican Dong , Wayne Xin Zhao , Ji-Rong Wen

We introduce Kimi K2, a Mixture-of-Experts (MoE) large language model with 32 billion activated parameters and 1 trillion total parameters. We propose the MuonClip optimizer, which improves upon Muon with a novel QK-clip technique to…

Machine Learning · Computer Science 2026-02-04 Kimi Team , Yifan Bai , Yiping Bao , Y. Charles , Cheng Chen , Guanduo Chen , Haiting Chen , Huarong Chen , Jiahao Chen , Ningxin Chen , Ruijue Chen , Yanru Chen , Yuankun Chen , Yutian Chen , Zhuofu Chen , Jialei Cui , Hao Ding , Mengnan Dong , Angang Du , Chenzhuang Du , Dikang Du , Yulun Du , Yu Fan , Yichen Feng , Kelin Fu , Bofei Gao , Chenxiao Gao , Hongcheng Gao , Peizhong Gao , Tong Gao , Yuyao Ge , Shangyi Geng , Qizheng Gu , Xinran Gu , Longyu Guan , Haiqing Guo , Jianhang Guo , Xiaoru Hao , Tianhong He , Weiran He , Wenyang He , Yunjia He , Chao Hong , Hao Hu , Yangyang Hu , Zhenxing Hu , Weixiao Huang , Zhiqi Huang , Zihao Huang , Tao Jiang , Zhejun Jiang , Xinyi Jin , Yongsheng Kang , Guokun Lai , Cheng Li , Fang Li , Haoyang Li , Ming Li , Wentao Li , Yang Li , Yanhao Li , Yiwei Li , Zhaowei Li , Zheming Li , Hongzhan Lin , Xiaohan Lin , Zongyu Lin , Chengyin Liu , Chenyu Liu , Hongzhang Liu , Jingyuan Liu , Junqi Liu , Liang Liu , Shaowei Liu , T. Y. Liu , Tianwei Liu , Weizhou Liu , Yangyang Liu , Yibo Liu , Yiping Liu , Yue Liu , Zhengying Liu , Enzhe Lu , Haoyu Lu , Lijun Lu , Yashuo Luo , Shengling Ma , Xinyu Ma , Yingwei Ma , Shaoguang Mao , Jie Mei , Xin Men , Yibo Miao , Siyuan Pan , Yebo Peng , Ruoyu Qin , Zeyu Qin , Bowen Qu , Zeyu Shang , Lidong Shi , Shengyuan Shi , Feifan Song , Jianlin Su , Zhengyuan Su , Lin Sui , Xinjie Sun , Flood Sung , Yunpeng Tai , Heyi Tang , Jiawen Tao , Qifeng Teng , Chaoran Tian , Chensi Wang , Dinglu Wang , Feng Wang , Hailong Wang , Haiming Wang , Jianzhou Wang , Jiaxing Wang , Jinhong Wang , Shengjie Wang , Shuyi Wang , Si Wang , Xinyuan Wang , Yao Wang , Yejie Wang , Yiqin Wang , Yuxin Wang , Yuzhi Wang , Zhaoji Wang , Zhengtao Wang , Zhengtao Wang , Zhexu Wang , Chu Wei , Qianqian Wei , Haoning Wu , Wenhao Wu , Xingzhe Wu , Yuxin Wu , Chenjun Xiao , Jin Xie , Xiaotong Xie , Weimin Xiong , Boyu Xu , Jinjing Xu , L. H. Xu , Lin Xu , Suting Xu , Weixin Xu , Xinran Xu , Yangchuan Xu , Ziyao Xu , Jing Xu , Jing Xu , Junjie Yan , Yuzi Yan , Hao Yang , Xiaofei Yang , Yi Yang , Ying Yang , Zhen Yang , Zhilin Yang , Zonghan Yang , Haotian Yao , Xingcheng Yao , Wenjie Ye , Zhuorui Ye , Bohong Yin , Longhui Yu , Enming Yuan , Hongbang Yuan , Mengjie Yuan , Siyu Yuan , Haobing Zhan , Dehao Zhang , Hao Zhang , Wanlu Zhang , Xiaobin Zhang , Yadong Zhang , Yangkun Zhang , Yichi Zhang , Yizhi Zhang , Yongting Zhang , Yu Zhang , Yutao Zhang , Yutong Zhang , Zheng Zhang , Haotian Zhao , Yikai Zhao , Zijia Zhao , Huabin Zheng , Shaojie Zheng , Longguang Zhong , Jianren Zhou , Xinyu Zhou , Zaida Zhou , Jinguo Zhu , Zhen Zhu , Weiyu Zhuang , Xinxing Zu

Multimodal Large Language Models (MLLMs) are gaining increasing popularity in both academia and industry due to their remarkable performance in various applications such as visual question answering, visual perception, understanding, and…

Computation and Language · Computer Science 2024-09-09 Jian Li , Weiheng Lu , Hao Fei , Meng Luo , Ming Dai , Min Xia , Yizhang Jin , Zhenye Gan , Ding Qi , Chaoyou Fu , Ying Tai , Wankou Yang , Yabiao Wang , Chengjie Wang

Open large language models (LLMs) have demonstrated improving multilingual capabilities in recent years. In this paper, we present a study of open LLMs for multilingual machine translation (MT) across a range of languages, and investigate…

Computation and Language · Computer Science 2026-02-26 Yuzhe Shang , Pengzhi Gao , Wei Liu , Jian Luan , Jinsong Su

We present a systematic evaluation of large language models on quantum mechanics problem-solving. Our study evaluates 15 models from five providers (OpenAI, Anthropic, Google, Alibaba, DeepSeek) spanning three capability tiers on 20 tasks…

Artificial Intelligence · Computer Science 2026-02-24 S. K. Rithvik

We evaluate the reasoning abilities of large language models in multilingual settings. We introduce the Multilingual Grade School Math (MGSM) benchmark, by manually translating 250 grade-school math problems from the GSM8K dataset (Cobbe et…

Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-linguistic reasoning abilities. This dual limitation makes it…

Multimodal Large Language Models (MLLMs) are undergoing rapid progress and represent the frontier of AI development. However, their training and inference efficiency have emerged as a core bottleneck in making MLLMs more accessible and…

Open-weight LLMs have been released by frontier labs; however, sovereign Large Language Models (for languages other than English) remain low in supply yet high in demand. Training large language models (LLMs) for low-resource languages such…

Computation and Language · Computer Science 2026-02-03 Shaltiel Shmidman , Avi Shmidman , Amir DN Cohen , Moshe Koppel

Large language models (LLMs) have recently demonstrated strong capabilities in generating machine learning (ML) code, enabling end-to-end pipeline construction from natural language instructions. However, existing benchmarks for ML code…

The development of open-source, multilingual medical language models can benefit a wide, linguistically diverse audience from different regions. To promote this domain, we present contributions from the following: First, we construct a…

Computation and Language · Computer Science 2024-06-04 Pengcheng Qiu , Chaoyi Wu , Xiaoman Zhang , Weixiong Lin , Haicheng Wang , Ya Zhang , Yanfeng Wang , Weidi Xie

We present a comprehensive adoption snapshot of the leading open language models and who is building them, focusing on the ~1.5K mainline open models from the likes of Alibaba's Qwen, DeepSeek, Meta's Llama, that are the foundation of an…

Computers and Society · Computer Science 2026-05-27 Nathan Lambert , Florian Brand

As large language models (LLMs) advance in conversational and reasoning capabilities, their practical application in healthcare has become a critical research focus. However, there is a notable gap between the performance of medical LLMs on…

Large language models (LLMs) have revolutionized natural language processing tasks. However, their practical deployment is hindered by their immense memory and computation requirements. Although recent post-training quantization (PTQ)…

Machine Learning · Computer Science 2024-03-19 Wenqi Shao , Mengzhao Chen , Zhaoyang Zhang , Peng Xu , Lirui Zhao , Zhiqian Li , Kaipeng Zhang , Peng Gao , Yu Qiao , Ping Luo
‹ Prev 1 8 9 10 Next ›