What is a small language model? SLMs vs LLMs
A small language model (SLM) is a language model with far fewer parameters than a large one, which makes it cheaper and faster to run and able to work on a single machine or even a phone or edge device. SLMs are often fine-tuned for a specific task, trading some general breadth for efficiency. This article explains how they compare with large language models, when a small model is the better choice, and why “small” is a relative and shifting term.
What makes a language model “small”
A language model’s size is usually described by its parameters, the internal values it adjusts during training. A small language model has far fewer of them than a large language model (LLM), and that single difference drives most of the practical trade-offs: less memory, less computing power and lower cost per answer.
Crucially, “small” is relative and moves over time. A model that counted as large a few years ago can be considered small today, so it is better to think of SLMs as the efficient end of a spectrum than as a fixed category. We deliberately avoid quoting sizes here, because they date quickly; the provider’s documentation is the place for them.
SLM versus LLM: the trade-off
Small and large models make opposite trade-offs, and neither is better in every situation:
- Breadth — an LLM generally handles a wider range of tasks and more open-ended questions; an SLM is usually stronger when focused on a narrower job.
- Cost — an SLM costs less to run per request; an LLM costs more for the extra capability.
- Speed — an SLM typically responds faster because it has less to compute.
- Where it runs — an SLM can run on a laptop, phone or edge device; an LLM usually needs substantial server infrastructure.
- Privacy — running an SLM locally can keep data on the device rather than sending it elsewhere.
How SLMs get good at specific jobs
Because a small model has less general capacity, it is commonly adapted to a defined task through fine-tuning: further training on examples from one domain, such as support tickets or product questions. Our explainer on what fine-tuning is describes the process.
The result is a model that does one thing well at modest cost. It will not match a large model on everything, but for a bounded task that gap often does not matter. Both kinds are part of the wider family covered in our explainer on what generative AI is.
When a small model is the better choice
Choose an SLM when the task is narrow and well defined, when low latency matters, when cost per request needs to stay low at high volume, when the model must run offline or on a device, or when data should not leave the user’s hardware. Typical examples include on-device text suggestions, classifying or routing messages, and extracting fields from documents.
Choose a large model when the task is open-ended, needs broad knowledge or complex reasoning, or you cannot predict what users will ask. Many real systems use both, sending simple requests to a small model and escalating harder ones to a large model.
Where this fits in your learning
Choosing the right model size is a standard practical decision in generative-AI work, and foundational certifications touch on it when they discuss cost, latency and model selection. Our explainer on large language models covers the other end of the spectrum.
Cloud providers offer both small and large models through managed services, and their catalogues change often, so rely on current documentation for specifics. Our /revision library covers the exam syllabus lesson by lesson, including how model choice affects cost and performance.
Original practice questions, timed mock exams and revision notes. No card, nothing to pay.
Questions, answered
Sources
Exam details in this post come from the vendor's published exam guide, which is the authority on what is tested and how.
- AWS Certified AI Practitioner (AIF-C01) exam guide — Amazon Web Services
- Google Cloud Generative AI Leader exam guide — Google Cloud