English
Related papers

Related papers: Fine-Tuning and Serving Gemma 4 31B on Google Clou…

200 papers

Modern GPU software stacks demand developers who can anticipate performance bottlenecks before ever launching a kernel; misjudging floating-point workloads upstream can derail tuning, scheduling, and even hardware procurement. Yet despite…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-05 Gregory Bolet , Giorgis Georgakoudis , Konstantinos Parasyris , Harshitha Menon , Niranjan Hasabnis , Kirk W. Cameron , Gal Oren

This paper presents Llama Guard 3-1B-INT4, a compact and efficient Llama Guard model, which has been open-sourced to the community during Meta Connect 2024. We demonstrate that Llama Guard 3-1B-INT4 can be deployed on resource-constrained…

Cybersecurity education is challenging and it is helpful for educators to understand Large Language Models' (LLMs') capabilities for supporting education. This study evaluates the effectiveness of LLMs in conducting a variety of penetration…

Cryptography and Security · Computer Science 2026-03-30 Martin Nizon-Deladoeuille , Brynjólfur Stefánsson , Helmut Neukirchen , Thomas Welsh

Due to the cost-prohibitive nature of training Large Language Models (LLMs), fine-tuning has emerged as an attractive alternative for specializing LLMs for specific tasks using limited compute resources in a cost-effective manner. In this…

Computation and Language · Computer Science 2024-08-15 Yuchen Xia , Jiho Kim , Yuhan Chen , Haojie Ye , Souvik Kundu , Cong Hao , Nishil Talati

Current FP8 grouped GEMM implementations require padding each group to a fixed alignment (e.g., 128), incurring memory and computational overhead. We propose \textit{TMA-Adaptive FP8 Grouped GEMM}, which eliminates padding by dynamically…

Hardware Architecture · Computer Science 2025-08-26 Zhongling Su , Rong Fu , Weihan Cao , Jianfei Gao , Minxi Jin , Zhilin Pei , Hui Wang

In this work we explore recent advances in instruction-tuning language models on a range of open instruction-following datasets. Despite recent claims that open models can be on par with state-of-the-art proprietary models, these claims are…

This paper introduces and evaluates a freely available cellular nonlinear network simulator optimized for the effective use of GPUs, to achieve fast modelling and simulations. Its relevance is demonstrated for several applications in…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-02-23 Radu Dogaru , Ioana Dogaru

The scaling of Large Language Models (LLMs) for retrieval-based tasks, particularly in Retrieval Augmented Generation (RAG), faces significant memory constraints, especially when fine-tuning extensive prompt sequences. Current open-source…

Machine Learning · Computer Science 2024-03-20 Anique Tahir , Lu Cheng , Huan Liu

This report evaluates the performance impact of enabling Trusted Execution Environments (TEE) on NVIDIA Hopper GPUs for large language model (LLM) inference tasks. We benchmark the overhead introduced by TEE mode across various LLMs and…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-11-06 Jianwei Zhu , Hang Yin , Peng Deng , Aline Almeida , Shunfan Zhou

We adopt CUDA-capable Graphic Processing Units (GPUs) for Coulomb, Landau and maximally Abelian gauge fixing in 3+1 dimensional SU(3) lattice gauge field theories. The local overrelaxation algorithm is perfectly suited for highly parallel…

High Energy Physics - Lattice · Physics 2012-12-07 Mario Schröck , Hannes Vogt

The LHCb experiment at CERN is undergoing an upgrade in preparation for the Run 3 data taking period of the LHC. As part of this upgrade the trigger is moving to a fully software implementation operating at the LHC bunch crossing rate. We…

Instrumentation and Detectors · Physics 2022-01-06 R. Aaij , M. Adinolfi , S. Aiola , S. Akar , J. Albrecht , M. Alexander , S. Amato , Y. Amhis , F. Archilli , M. Bala , G. Bassi , L. Bian , M. P. Blago , T. Boettcher , A. Boldyrev , S. Borghi , A. Brea Rodriguez , L. Calefice , M. Calvo Gomez , D. H. Cámpora Pérez , A. Cardini , M. Cattaneo , V. Chobanova , G. Ciezarek , X. Cid Vidal , J. L. Cobbledick , J. A. B. Coelho , T. Colombo , A. Contu , B. Couturier , D. C. Craik , R. Currie , P. d'Argent , M. De Cian , D. Derkach , F. Dordei , M. Dorigo , L. Dufour , P. Durante , A. Dziurda , A. Dzyuba , S. Easo , S. Esen , P. Fernandez Declara , S. Filippov , C. Fitzpatrick , M. Frank , P. Gandini , V. V. Gligorov , E. Golobardes , G. Graziani , L. Grillo , P. A. Günther , S. Hansmann-Menzemer , A. M. Hennequin , L. Henry , D. Hill , S. E. Hollitt , J. Hu , W. Hulsbergen , R. J. Hunter , M. Hushchyn , B. K. Jashal , C. R. Jones , S. Klaver , K. Klimaszewski , R. Kopecna , W. Krzemien , M. Kucharczyk , R. Lane , F. Lazzari , R. Le Gac , P. Li , J. H. Lopes , M. Lucio Martinez , A. Lupato , O. Lupton , X. Lyu , F. Machefert , O. Madejczyk , S. Malde , J. F. Marchand , S. Mariani , C. Marin Benito , D. Martinez Santos , F. Martinez Vidal , R. Matev , M. Mazurek , B. Mitreska , D. S. Mitzel , M. J. Morello , H. Mu , P. Muzzetto , P. Naik , M. Needham , N. Neri , N. Neufeld , N. S. Nolte , D. O'Hanlon , A. Oyanguren , M. Pepe Altarelli , S. Petrucci , M. Petruzzo , L. Pica , F. Pisani , A. Piucci , F. Polci , A. Poluektov , E. Polycarpo , C. Prouve , G. Punzi , R. Quagliani , R. I. Rabadan Trejo , M. Ramos Pernas , M. S. Rangel , F. Ratnikov , G. Raven , F. Reiss , V. Renaudin , P. Robbe , A. Ryzhikov , M. Santimaria , M. Saur , M. Schiller , R. Schwemmer , B. Sciascia , A. Solomin , F. Suljik , N. Skidmore , M. D. Sokoloff , P. Spradlin , M. Stahl , S. Stahl , H. Stevens , L. Sun , A. Szabelski , T. Szumlak , M. Szymanski , D. Y. Tou , G. Tuci , A. Usachov , N. Valls Canudas , R. Vazquez Gomez , S. Vecchi , M. Vesterinen , X. Vilasis-Cardona , D. Vom Bruch , Z. Wang , T. Wojton , M. Whitehead , M. Williams , M. Witek , Y. Xie , A. Xu , H. Yin , M. Zdybal , O. Zenaiev , D. Zhang , L. Zhang , X. Zhu

Deploying deep neural networks on mobile devices is increasingly important but remains challenging due to limited computing resources. On the other hand, their unified memory architecture and narrower gap between CPU and GPU performance…

Machine Learning · Computer Science 2026-02-20 Zhuojin Li , Marco Paolieri , Leana Golubchik

The rapid adoption of Large Language Models (LLMs) has made GPU inference efficiency an increasingly critical system concern. The runtime of LLM workloads is largely dominated by tile-based kernels, particularly General Matrix…

Performance · Computer Science 2026-04-14 Kaixuan Zhang , Chutong Ding , Shiyou Qian , Luping Wang , Jian Cao , Guangtao Xue , Cheng Huang , Guodong Yang , Liping Zhang

Large Language Models (LLMs) face significant deployment challenges due to their substantial resource requirements. While low-bit quantized weights can reduce memory usage and improve inference efficiency, current hardware lacks native…

Machine Learning · Computer Science 2025-06-10 Pengxiang Zhao , Xiaoming Yuan

Current LLM structured pruning methods typically involve two steps: (1) compression with calibration data and (2) costly continued pretraining on billions of tokens to recover lost performance. This second step is necessary as the first…

Machine Learning · Computer Science 2024-12-31 Yaya Sy , Christophe Cerisara , Irina Illina

In Scientific Computing and modern Machine Learning (ML) workloads, sequences of dependent General Matrix Multiplications (GEMMs) often dominate execution time. While state-of-the-art BLAS libraries aggressively optimize individual GEMM…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-07 César Guedes Carneiro , Lucas Alvarenga , Guido Araujo , Sandro Rigo

The use of large language models (LLMs) is expanding rapidly, and open-source versions are becoming available, offering users safer and more adaptable options. These models enable users to protect data privacy by eliminating the need to…

Machine Learning · Computer Science 2024-08-06 Hui Yin , Amir Aryani , Nakul Nambiar

This technical report briefly describes our JDExplore d-team's Vega v2 submission on the SuperGLUE leaderboard. SuperGLUE is more challenging than the widely used general language understanding evaluation (GLUE) benchmark, containing eight…

Computation and Language · Computer Science 2022-12-06 Qihuang Zhong , Liang Ding , Yibing Zhan , Yu Qiao , Yonggang Wen , Li Shen , Juhua Liu , Baosheng Yu , Bo Du , Yixin Chen , Xinbo Gao , Chunyan Miao , Xiaoou Tang , Dacheng Tao

Large Language Model (LLM) deployment is increasingly shifting to cost-efficient accelerators like Google's Tensor Processing Units (TPUs), prioritizing both performance and total cost of ownership (TCO). However, existing LLM inference…

Performance · Computer Science 2026-04-20 Jevin Jiang , Ying Chen , Blake A. Hechtman , Fenghui Zhang , Yarong Mu

The latency and power consumption of large language models (LLMs) are major constraints when serving them across a wide spectrum of hardware platforms, from mobile edge devices to cloud GPU clusters. Benchmarking is crucial for optimizing…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-12 Hung-Yueh Chiang , Bokun Wang , Diana Marculescu