English
Related papers

Related papers: Measuring Maximum Activations in Open Large Langua…

200 papers

Advances in Large Language Models (LLMs) have led to significant interest in their potential to support human experts across a range of domains, including public health. In this work we present automated evaluations of LLMs for public…

While Multimodal Large Language Models (MLLMs) have experienced significant advancement in visual understanding and reasoning, their potential to serve as powerful, flexible, interpretable, and text-driven models for Image Quality…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Tianhe Wu , Kede Ma , Jie Liang , Yujiu Yang , Lei Zhang

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B…

Artificial Intelligence · Computer Science 2026-05-27 MiniMax , : , Aili Chen , Aonian Li , Baichuan Zhou , Bangwei Gong , Binyang Jiang , Boji Dan , Changqing Yu , Chao Wang , Cheng Ma , Cheng Zhong , Cheng Zhu , Chengjun Xiao , Chengyi Yang , Chengyu Du , Chenyang Zhang , Chi Zhang , Chuangyi Huang , Chunhao Zhang , Chunhui Du , Chunyu Zhao , Congchao Guo , Da Chen , Deming Ding , Dianjun Sun , Dongyu Zhang , Enhui Yang , Fei Yu , Guang Zheng , Guodong Zheng , Guohong Li , Haichao Zhu , Haigang Zhou , Haimo Zhang , Han Ding , Hao Zhang , Haohai Sun , Haolin Lyu , Haonan Lu , Haoyu Wang , Huajie Shi , Huiyang Li , Jiacheng Chen , Jian Zhang , Jiaqi Zhuang , Jiaren Cai , Jiaxin Pan , Jiayao Li , Jiayuan Song , Jichuan Zhang , Jie Wang , Jihao Gu , Jin Zhu , Jingwei Dong , Jingyang Li , Jingyu Zhang , Jingze Zhuang , Jinhao Tian , Jinli Liu , Jinyi Hu , Jun Tao , Jun Zhang , Junbin Ruan , Junhao Xu , Junjie Yan , Junteng Liu , Junxian He , Kang Xu , Ke Ji , Ke Yang , Kecheng Xiao , Keyu Duan , Keyu Li , Le Han , Letian Ruan , Li Yuan , Lianfei Yu , Liheng Feng , Lijie Mo , Lin Li , Lingye Bao , Lingyu Yang , Lingyuan Zhou , Loki , Lu Chen , Lunbin Ceng , Ming Li , Ming Zhong , Mingliang Tao , Mingyuan Chi , Mujie Lin , Nan Hu , Ningxin Chen , Peiyin Zhu , Peng Gao , Pengcheng Gao , Pengfei Li , Penglin Li , Pengyu Zhao , Qibin Ren , Qidi Xu , Qihan Ren , Qile Li , Qin Wang , Quanliang Chen , Qunhong Ceng , Rong Tian , Rui Dong , Ruitao Leng , Ruize Zhang , Shanqi Liu , Shaoyu Chen , Sheng Jia , Shun Yao , Shuoran Zhao , Shuqi Yu , Sichen Li , Sicheng Pan , Songquan Zhu , Tengfei Li , Tian Xie , Tiancheng Qin , Tianrun Liang , Wei Liu , Weiqi Xu , Weitao Li , Weixiang Chen , Weiyu Cheng , Weiyu Zhang , Wenhu Chen , Wenqian Zhao , Xiancai Chen , Xiangjun Song , Xiangyuan Wang , Xiao Luo , Xiao Su , Xiaobo Li , Xiaodong Han , Xiaojie Wu , Xihao Song , Xingyi Han , Xinyu Guan , Xuan Lu , Xun Zou , Xunhao Lai , Xutong Li , Yan Gong , Yang Wang , Yang Xu , Yangsen Wang , Ye Tang , Yicheng Chen , Yinran Qiu , Yiqi Shi , Yiting Guo , Yiwen Huang , Yixuan Wang , Yongyi Hu , Yu Gao , Yu Zhang , Yuanxiang Ying , Yuanzhen Zhang , Yubo Wang , Yuchen Song , Yufeng Yang , Yuhang Meng , Yuhang Miao , Yuhao Li , Yujie Liu , Yulin Hu , Yunan Huang , Yunji Li , Yunyi Huang , Yusen Zhang , Yusu Hong , Yutao Xie , Yutong Zhang , Yuwen Liao , Yuxuan Shi , Yuze Wenren , Zebin Li , Zehan Li , Zejian Luo , Zeyu Jin , Zeyuan Sun , Zhanpeng Zhou , Zhaochen Su , Zhendong Li , Zhengmao Zhu , Zhengyuan Peng , Zhenhua Fan , Zhi Zhang , Zhichao Xu , Zhiheng Lv , Zhikang Xu , Zhitao He , Zhiwei He , Zhongyuan Li , Zibo Gao , Zijia Wu , Zijian Song , Zijian Zhou , Zijun Sun , Zishan Huang , Ziying Chen , Ziyue Ge

Recently, Large Language Models (LLM) have demonstrated impressive capability to solve a wide range of tasks. However, despite their success across various tasks, no prior work has investigated their capability in the biomedical domain yet.…

Computation and Language · Computer Science 2024-02-21 Israt Jahan , Md Tahmid Rahman Laskar , Chun Peng , Jimmy Huang

LLM agents with tool access can discover and exploit security vulnerabilities. This is known. What is not known is which features of a system prompt trigger this behaviour, and which do not. We present a systematic taxonomy based on…

Cryptography and Security · Computer Science 2026-04-07 Charafeddine Mouzouni

Multilingualism in Large Language Models (LLMs) is an yet under-explored area. In this paper, we conduct an in-depth analysis of the multilingual capabilities of a family of a Large Language Model, examining its architecture, activation…

Computation and Language · Computer Science 2024-04-23 Sunit Bhattacharya , Ondřej Bojar

The full-size MLPs and the projection layers in attention introduce tremendous model sizes of large language models (LLMs), consuming extensive computational resources in pre-training. We empirically observe that the activations of…

Machine Learning · Computer Science 2025-10-03 Ziyue Liu , Ruijie Zhang , Zhengyang Wang , Mingsong Yan , Zi Yang , Paul Hovland , Bogdan Nicolae , Franck Cappello , Sui Tang , Zheng Zhang

Recently, inspired by the concept of sparsity, Mixture-of-Experts (MoE) models have gained increasing popularity for scaling model size while keeping the number of activated parameters constant. In this study, we thoroughly investigate the…

Computation and Language · Computer Science 2024-11-26 Xiaoye Qu , Daize Dong , Xuyang Hu , Tong Zhu , Weigao Sun , Yu Cheng

In this work, we systematically investigate the efficacy of dynamic activation mechanisms within the LLaMA family of language models. Despite the potential of dynamic activation methods to reduce computation and increase speed in models…

Machine Learning · Computer Science 2024-05-16 Chi Ma , Mincong Huang , Chao Wang , Yujie Wang , Lei Yu

Large language models (LLMs) pretrained on vast source code have achieved prominent progress in code intelligence. However, existing code LLMs have two main limitations in terms of architecture and pretraining tasks. First, they often adopt…

Computation and Language · Computer Science 2023-05-23 Yue Wang , Hung Le , Akhilesh Deepak Gotmare , Nghi D. Q. Bui , Junnan Li , Steven C. H. Hoi

Large Language Models (LLMs) have shown remarkable capabilities in manipulating natural language across multiple applications, but their ability to handle simple reasoning tasks is often questioned. In this work, we aim to provide a…

Computation and Language · Computer Science 2025-05-05 Alessandro Raganato , Rafael Peñaloza , Marco Viviani , Gabriella Pasi

Large language models (LLMs) excel at explicit reasoning, but their implicit computational strategies remain underexplored. Decades of psychophysics research show that humans intuitively process and integrate noisy signals using…

Computation and Language · Computer Science 2025-12-03 Julian Ma , Jun Wang , Zafeirios Fountas

We consider the problem of model compression for Large Language Models (LLMs) at post-training time, where the task is to compress a well-trained model using only a small set of calibration input data. In this work, we introduce a new…

Machine Learning · Statistics 2024-12-12 Meyer Scetbon , James Hensman

Automated essay scoring (AES) is a challenging task in cross-prompt settings due to the diversity of scoring criteria. While previous studies have focused on the output of large language models (LLMs) to improve scoring accuracy, we believe…

Computation and Language · Computer Science 2025-12-23 Jinwei Chi , Ke Wang , Yu Chen , Xuanye Lin , Qiang Xu

Large Language Models increasingly power critical infrastructure from healthcare to finance, yet their vulnerability to adversarial manipulation threatens system integrity and user safety. Despite growing deployment, no comprehensive…

Cryptography and Security · Computer Science 2026-03-19 Taiwo Onitiju , Iman Vakilinia

Large language models (LLMs) exhibit remarkable similarity to neural activity in the human language network. However, the key properties of language shaping brain-like representations, and their evolution during training as a function of…

Computation and Language · Computer Science 2025-09-23 Badr AlKhamissi , Greta Tuckute , Yingtian Tang , Taha Binhuraib , Antoine Bosselut , Martin Schrimpf

Multimodal large language models (MLLMs) have emerged as a promising paradigm for dental image analysis. However, their ability to capture the multi-level cognitive processes required for radiographic analysis remains unclear. Here, we…

Computation and Language · Computer Science 2026-05-11 Rongyang Wang , Shuang Zhou , Jiashuo Wang , Wenya Xie , Xiaoxia Che

Multimodal large language models (LLMs) are increasingly explored as automated evaluators in clinical settings, yet their scoring behavior on ordinal clinical scales remains poorly understood. We benchmark three frontier LLM families…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jiaqing Zhang , Sandeep Elluri , Bhanu Cherukuvada , Yonah Joffe , Jessica Sena , Miguel Contreras , Scott Siegel , Subhash Nerella , Catherine Price , Parisa Rashidi

Extreme activation outliers in Large Language Models (LLMs) critically degrade quantization performance, hindering efficient on-device deployment. While channel-wise operations and adaptive gradient scaling are recognized causes, practical…

Machine Learning · Computer Science 2025-06-25 Jungwoo Park , Taewhoo Lee , Chanwoong Yoon , Hyeon Hwang , Jaewoo Kang

Molecular generative models, often employing GPT-style language modeling on molecular string representations, have shown promising capabilities when scaled to large datasets and model sizes. However, it remains unclear and subject to debate…

Machine Learning · Computer Science 2026-02-02 Dong Xu , Qihua Pan , Sisi Yuan , Jianqiang Li , Zexuan Zhu , Junkai Ji