This paper presents Llama Guard 3-1B-INT4, a compact and efficient Llama Guard model, which has been open-sourced to the community during Meta Connect 2024. We demonstrate that Llama Guard 3-1B-INT4 can be deployed on resource-constrained devices, achieving a throughput of at least 30 tokens per second and a time-to-first-token of 2.5 seconds or less on a commodity Android mobile CPU. Notably, our experiments show that Llama Guard 3-1B-INT4 attains comparable or superior safety moderation scores to its larger counterpart, Llama Guard 3-1B, despite being approximately 7 times smaller in size (440MB).
Cite
@article{arxiv.2411.17713,
title = {Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations},
author = {Igor Fedorov and Kate Plawiak and Lemeng Wu and Tarek Elgamal and Naveen Suda and Eric Smith and Hongyuan Zhan and Jianfeng Chi and Yuriy Hulovatyy and Kimish Patel and Zechun Liu and Changsheng Zhao and Yangyang Shi and Tijmen Blankevoort and Mahesh Pasupuleti and Bilge Soran and Zacharie Delpierre Coudert and Rachad Alao and Raghuraman Krishnamoorthi and Vikas Chandra},
journal= {arXiv preprint arXiv:2411.17713},
year = {2024}
}