Foundations

Small Language Model

SLM

A language model with far fewer parameters than the largest models, built to run cheaply and quickly, often on a single device or on a company's own servers.

Large model (LLM)the highest capabilityRuns in the cloud, through an APICosts more per tokenFor hard, open-ended workSmall model (SLM)enough capability, low costOn a device or your own serverFast and cheapFor narrow, repetitive workThe question to ask: which is the smallest model that is good enough for this job?

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/small-language-model

In plain terms

Not every job needs the most capable model. Sorting support tickets, pulling an invoice number out of an email or answering questions about one product can be done by a much smaller model that responds faster, costs a fraction as much and can run on a laptop or a phone, with no data leaving the building.

Why it matters

There are three reasons to care: cost, speed and control. At high volume the price gap between a large and a small model decides whether a use case pays for itself. Small models answer quickly enough for real-time use. And they can run inside your own environment, which settles many data-protection questions. The trade-off is capability: they are weaker at open-ended reasoning and broad knowledge. The useful question is always which is the smallest model that does this job well enough.

Example

An insurer sorts 200,000 incoming emails a month into 14 categories. A frontier model does it with 97% accuracy. A small model fine-tuned on 5,000 of the insurer's own labelled emails reaches 96% at about a twentieth of the cost, and runs on the insurer's own servers. The large model is kept for the 4% of emails that the small one marks as uncertain.

Most often confused with

SLM vs. Large language model (LLM)

SLMSmaller, cheaper, faster; good at narrow tasks
Large language model (LLM)Larger, costlier; better at open-ended reasoning

No official line separates them. “Small” today usually means somewhere from under a billion to roughly ten billion parameters, and the boundary moves every year. The technology is the same. The difference is an engineering trade-off between capability and cost, and good systems use both: a small model for the routine volume and a large one for the hard cases.

Under the hood

Small models are made in three ways: by training a compact architecture directly, by distillation from a larger model, and by quantisation, which stores weights at lower precision to save memory. Examples include Microsoft's Phi family, Google's Gemma, the smaller Llama and Mistral models, and the small tiers of commercial model families. They can be deployed on devices (phones, laptops, embedded systems), on a single GPU server, or through a low-priced API. They respond well to fine-tuning on narrow tasks, where a tuned small model can match a general large one. Common patterns: routing (send easy requests to the small model and escalate the rest), cascades, and small models as subagents or classifiers inside a larger system. Limits: a shorter usable context, weaker multi-step reasoning, less world knowledge and greater sensitivity to the wording of prompts.

Written by Mehmet Erkek · Last updated: