English

Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?

Computation and Language 2025-03-13 v3 Artificial Intelligence Machine Learning

Abstract

The rapid rise of Language Models (LMs) has expanded their use in several applications. Yet, due to constraints of model size, associated cost, or proprietary restrictions, utilizing state-of-the-art (SOTA) LLMs is not always feasible. With open, smaller LMs emerging, more applications can leverage their capabilities, but selecting the right LM can be challenging as smaller LMs do not perform well universally. This work tries to bridge this gap by proposing a framework to experimentally evaluate small, open LMs in practical settings through measuring semantic correctness of outputs across three practical aspects: task types, application domains, and reasoning types, using diverse prompt styles. It also conducts an in-depth comparison of 10 small, open LMs to identify the best LM and prompt style depending on specific application requirements using the proposed framework. We also show that if selected appropriately, they can outperform SOTA LLMs like DeepSeek-v2, GPT-4o, GPT-4o-mini, Gemini-1.5-Pro, and even compete with GPT-4o.

Keywords

Cite

@article{arxiv.2406.11402,
  title  = {Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?},
  author = {Neelabh Sinha and Vinija Jain and Aman Chadha},
  journal= {arXiv preprint arXiv:2406.11402},
  year   = {2025}
}

Comments

Accepted at The Fifth Workshop on Trustworthy Natural Language Processing (TrustNLP 2025) in Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL), 2025. 8 pages + references + Appendix

R2 v1 2026-06-28T17:08:26.621Z