SaveMyCert
AI & ML

What is a diffusion model? Image generation explained

A diffusion model is a type of generative AI that creates images, and other media, by learning to reverse a step-by-step process of adding random noise. It starts from pure static and gradually removes the noise until a coherent result emerges, steered by your text prompt. Diffusion models power most of the text-to-image tools you will have seen, and they belong to the same generative-AI family as the chatbots built on large language models, even though they are built differently. This article explains the noising and denoising idea at a beginner level, how diffusion differs from transformer-based language models, and where the technique is used.

The core idea: learn to remove noise

A diffusion model learns to undo damage. During training, real images are deliberately ruined: random noise is added a little at a time until nothing but static remains. The model is then taught the reverse, so that given a noisy image it can predict what a slightly cleaner version looks like. Repeat that across a vast number of examples and it becomes very good at spotting structure inside noise.

The word “diffusion” comes from physics, where particles spread out gradually, like ink dispersing in water. The model learns to run that spreading backwards, pulling order out of randomness.

How an image is generated from your prompt

Generation starts with an image of pure random noise. The model removes a little of that noise, then a little more, over many small steps, and at each step your prompt nudges the result towards what you described. A prompt such as “a lighthouse at dusk” steers the denoising so the emerging shapes and colours match that description rather than any random picture.

Because each run starts from different random noise, the same prompt can produce many different images. That variety is a feature, and it is why generated images are not copies pulled from a stored library but fresh output shaped by learned patterns.

How diffusion models differ from language models

Both are generative AI, but they are built differently. The large language models behind chatbots use a transformer architecture and generate text one piece at a time, each piece predicted from what came before. A diffusion model instead refines a whole image from noise over repeated steps. Different architecture, different process, same family: both learn patterns from data and then produce new content.

The two also combine. Many tools use a language model to understand your prompt and a diffusion model to draw it, which is part of how multimodal systems work. Our explainers on what generative AI is and what multimodal AI is cover that wider picture.

Where diffusion models are used

The most familiar use is text-to-image generation, but the same idea extends well beyond it:

  • Image generation and editing — creating pictures from text, filling in missing regions, changing a style or extending a scene.
  • Video generation — applying denoising across a sequence of frames so motion stays coherent.
  • Audio and music — generating sound by denoising audio representations.
  • Design and product work — rapid mock-ups, concept art and marketing visuals that a person then refines.

Limits worth knowing

Diffusion models inherit the habits of their training data, including its biases and gaps, and they can produce images that look convincing but contain errors, such as distorted details. Questions about copyright, consent and misuse of generated media are active concerns, which is why responsible-AI practice matters here as much as for text.

Which diffusion models exist, and how capable each is, changes constantly, so check the provider’s own documentation for specifics. The concepts above are the durable part. For how this is examined, our /revision library covers the generative-AI syllabus lesson by lesson, and our computer vision explainer shows how generating images differs from analysing them.

Ready to start studying — free?

Original practice questions, timed mock exams and revision notes. No card, nothing to pay.

Jump straight into an exam
AIF-C01GenAI Leader

Questions, answered

A diffusion model is a generative AI that creates images by starting from random noise and removing it step by step until a picture appears. It learns this by being trained to reverse the process of gradually adding noise to real images, and your text prompt guides what it draws.

Sources

Exam details in this post come from the vendor's published exam guide, which is the authority on what is tested and how.

Keep reading

AI & ML
What is a small language model? SLMs vs LLMs
AI & ML
What is AI governance? A plain-English explanation
AI & ML
What is semantic search? Search by meaning, explained
AI & ML
Why does AI need GPUs? Parallel maths explained