SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
Identify AI concepts and capabilities

How Generative AI Models Work: Tokens, Model Choice and Deployment (AI-901)

15 min readAI-901 · Identify AI concepts and capabilitiesUpdated

A generative AI model is a neural network, usually a transformer-based language model, that creates new content by predicting one token at a time from patterns it learned in training. For AI-901 you need three things on top of that definition: how the model works inside (tokens, embeddings, attention and next-token prediction), how to pick the right model for a capability in the Microsoft Foundry model catalog, and how deployment types and request parameters such as temperature and max output tokens change cost, throughput, data location and the output you get back.

What you’ll learn
  • Describe how a language model turns a prompt into a completion through tokens, embeddings, attention and next-token prediction
  • Explain what tokens, embeddings and the context window are, and why limits and billing are counted in tokens
  • Choose between chat, reasoning, multimodal, embeddings, image-generation and small language models from a stated requirement
  • Use the Foundry model catalog, model cards and leaderboards to compare models before deploying
  • Match a deployment type (Global, Data Zone or Standard; standard, provisioned, batch or Developer) to residency, throughput and cost needs
  • Diagnose throttling, input-too-long errors and truncated replies, and pick the setting that fixes each
  • Set temperature, top_p and max output tokens appropriately, including the different parameters reasoning models accept

How does a generative AI model turn a prompt into a completion?

A generative AI model turns a prompt into a completion by predicting a probable next token, appending it to the text, and repeating that step until it predicts the end of the response. The input text is the prompt; the generated text is the completion (in a chat app, the assistant's reply).

The pipeline behind every request has four stages:

  1. Tokenize — the prompt is split into tokens, each mapped to an integer ID.
  2. Embed — each token ID becomes a vector of numbers, combined with a positional encoding that records where the token sits.
  3. Attend — transformer layers use attention to adjust each token's vector according to the tokens around it, producing contextual representations.
  4. Predict — the model outputs a probability for every token in its vocabulary, one token is sampled (temperature and top_p shape this choice; at very low temperature the most probable token almost always wins), and the loop repeats.

Two consequences follow. First, the model does not look anything up: it does not search its training documents, pick a prewritten answer or run rules against a database. It continues text using statistical patterns it learned during training. Second, because it produces a plausible-sounding continuation, it can write fluent, confident text that is not true — for example a citation to a report that doesn't exist. Its knowledge is limited to its training data, and it has no built-in sense of where that knowledge ends. Applications reduce this by grounding the prompt with current, trusted data (covered with retrieval and agents in the generative AI workload lessons).

What is a token?

A token is the unit of text a language model reads and writes: a whole word, part of a word, a punctuation mark or another frequent character sequence, each mapped to an integer ID by the model's tokenizer. Common short words are usually one token; longer or rarer words are split into several sub-word pieces (for example a prefix such as un plus the rest of the word); punctuation marks are tokens in their own right.

Tokens matter because everything on a model is counted in them, never in words or characters:

  • the context window — how much a single request can hold;
  • max output tokens — how long one response can be;
  • the deployment's tokens-per-minute (TPM) allocation — how much traffic it accepts;
  • pricing — input and output tokens are billed (output usually at a higher rate).

That is why a cost estimate built by counting words comes out low: sub-word splits and punctuation make the token count higher than the word count.

What are embeddings and why do they capture meaning?

An embedding is a vector of floating-point numbers that represents the meaning of a piece of text. Texts with similar meanings produce vectors that point in similar directions, so you can compare meaning mathematically — most often with cosine similarity (the closer to 1, the more similar).

Embeddings appear in two places:

  • Inside every language model, token IDs are converted to embedding vectors before the transformer layers process them.
  • As a product in their own right, an embeddings model (such as text-embedding-3-small or text-embedding-3-large) takes text and returns one vector per input. These vectors power semantic search, recommendations, clustering and the retrieval step of retrieval-augmented generation: a query for a waterproof jacket can match a product described as a rain shell even though the two share no words.

The properties of an embedding to know:

PropertyWhat it means
FormatAn array of floating-point numbers, one array per input
LengthFixed for a given model and dimensions setting, whatever the input length (for example 1,536 dimensions by default for text-embedding-3-small; the text-embedding-3 models let you request fewer)
GeometryRelated texts sit close together; distance or cosine similarity measures how alike they are
UseComparison — search, ranking, clustering, retrieval

Embeddings also differ from keyword approaches: keyword extraction or keyword search still needs shared words, while embeddings compare meaning.

How do transformers use attention and positional encoding?

A transformer is the neural network architecture behind today's language models; it uses attention to work out how much each token should be influenced by every other token, and a positional encoding to keep track of word order.

  • Attention. For each token, attention layers weigh the other tokens in the sequence by relevance. In the river burst its bank after the storm, the representation of bank draws more on river and storm than on the or its — which is how the model knows this is a riverside, not a financial institution. Multi-head attention runs several attention calculations in parallel, each picking up a different kind of relationship. The output is a contextual embedding: the same word gets a different vector in a different sentence.
  • Positional encoding. Attention on its own treats the input as a set, not a sequence. The teacher thanked the student and the student thanked the teacher contain exactly the same tokens with the same IDs, so the model adds a positional encoding to each token's vector to mark where it sits. Order changes meaning, and positional encoding is what carries it.
  • Encoder and decoder blocks. The original transformer has an encoder block that builds contextual representations of the input and a decoder block that generates output one token at a time, attending to what it has produced so far. Generative chat models (the GPT family, for example) are decoder-based; models built for understanding text, such as BERT, are encoder-based.
ComponentWhat it contributes
TokenizerSplits text into tokens with integer IDs
EmbeddingTurns each token into a vector of numbers
Positional encodingAdds each token's position, so order counts
Attention layersWeigh how strongly other tokens influence each token
Sampling (temperature, top_p)Chooses the next token from the predicted probabilities — only at output time

What is a context window?

A context window is the maximum number of tokens a model can handle in a single request, and the input and the output share it. Everything the model sees and writes for one call — system message, conversation history, grounding data, the user's question, the generated reply and, for reasoning models, the hidden reasoning tokens — must fit inside that one budget.

Language models are stateless: a chat app keeps a conversation going by resending the history with every request. In a long conversation that history keeps growing until the request no longer fits and the call fails with an input-too-long error. The fix is to keep each request within the window — trim older turns, summarize them, or keep a sliding window of recent messages — or to choose a model with a larger context window.

Keep two different limits apart: the context window caps the size of one request, while a deployment's throughput settings (tokens per minute, provisioned throughput) cap how much traffic it accepts per minute across all requests.

How do you choose a model in the Microsoft Foundry model catalog?

You choose a model in the Microsoft Foundry model catalog, the hub where you discover, compare and deploy models from Microsoft, OpenAI and many other providers. You can filter by capability, task, provider and deployment option, and every model has a model card describing what it does, its inputs and outputs, limits and benchmark results.

To compare candidates before you deploy anything, use the model leaderboards, which rank models on quality, safety, cost and throughput, and the side-by-side comparison view. Once you have picked a model, you deploy it and can try it in the playground.

The catalog groups models into two families:

Foundry Models sold directly by AzureModels from partners and community
ExamplesAll Azure OpenAI models plus selected models from other leading providersA much wider range of partner, open-weight and community models (for example from Hugging Face)
Hosted and supported byMicrosoft, under Microsoft Product TermsThe model provider (many offered through Azure Marketplace)
SLA and billingCovered by Azure SLAs, billed through your Azure subscriptionProvider's own terms; support and SLA come from the provider
Best forOrganizations that need Microsoft hosting, support and enterprise termsVariety, specialised or open models, niche tasks

Which type of model fits which capability?

Pick the model type from the input and output the solution needs, then from its constraints (latency, cost, where it runs).

NeedModel typeTypical examples
Fast conversational answers, drafting, summarizing, high-volume chatGeneral-purpose chat (completion) modelGPT-4.1, GPT-4o families
Multi-step problems: math, planning, complex analysis, codeReasoning modelo-series, GPT-5 reasoning models
Questions about an image with a written answerMultimodal chat model (image input, text output)GPT-4o, GPT-4.1
Semantic search, similarity, clustering, retrievalEmbeddings modeltext-embedding-3-small / -large
New still images from a text descriptionImage-generation model (text in, image out)gpt-image-1 family
New video from a descriptionVideo-generation modelSora
Small device, offline, tight memory or low costSmall language model (SLM)Phi family

The near-twins to keep apart: a multimodal chat model reads an image and answers in text, an image-generation model goes the other way (text to picture), and an image-analysis service such as Azure AI Vision tags or captions existing pictures. Spoken input or output is a separate capability again: it needs an audio-capable model variant, such as the GPT-4o audio or realtime models, rather than a standard text-and-image chat deployment.

Reasoning models are not always the better choice. They work through a problem before answering, which raises accuracy on multi-step tasks but adds latency and generates reasoning tokens — never shown in the reply, yet billed as output tokens and counted against the context window. For simple, high-volume, real-time chat where speed and cost decide, a standard chat model such as GPT-4.1 is faster and cheaper.

Large vs small language models. Large language models (LLMs) are broad and capable but typically need data-centre-class hardware, so they are usually consumed as a cloud endpoint. Small language models such as Microsoft's Phi family are compact models designed for constrained scenarios: running on-device or offline, with limited memory, or where low cost and latency matter more than breadth. What an SLM gives up is breadth of knowledge and depth of reasoning; what it gains is lower cost, lower latency and the ability to run on modest hardware. A longer context window is not one of its advantages: a small model such as Phi-4-mini supports about 128,000 tokens, while large models such as GPT-4.1 accept far more.

What deployment types control where data is processed?

When you deploy a model sold directly by Azure, the deployment type decides where prompts and responses may be processed and how capacity is billed. Location comes in three scopes:

ScopeWhere inference may be processedWhen to choose it
GlobalAny Azure region where the model is availableNo residency requirement; lowest price, best availability, highest default quota, newest models first
Data ZoneOnly within a data zone — the US, the EU or Asia-PacificData must stay within one of those zones
Standard / RegionalOnly within the Azure geography of the resourceData must stay in a single geography

In every scope, data at rest stays in the geography you chose for the resource; the scope governs where processing happens. Microsoft recommends Global Standard as the starting point for most workloads: pay-per-token, the lowest price, broadest region coverage, the highest default quota and early access to new models. Move away from it only for a reason — residency, reserved capacity or batch work.

What deployment types control billing, throughput and throttling?

The second half of a deployment type's name is the capacity model, and it combines with the location scope (Global Standard, Data Zone Provisioned, Global Batch and so on).

Capacity modelHow it worksFits
StandardPay per token; shared capacity, best effortMost workloads, variable traffic, prototypes
ProvisionedReserved throughput bought in provisioned throughput units (PTUs)Steady, high-volume production that needs predictable, low-variance latency
Batch (Global or Data Zone)Submit a file of requests; processed asynchronously against a separate enqueued-token quota, targeting completion within 24 hours, at 50% less than Global StandardLarge, non-urgent jobs: bulk summarization, classification, evaluation
DeveloperPay per token for fine-tuned models; no SLA, no data-residency guarantee; each deployment is deleted automatically after 24 hoursCheap, short evaluation of a fine-tuned model before a production decision

Combine the two axes to answer a requirement. The location scope answers where processing may happen and the capacity model answers how capacity is bought, so read each requirement against its own axis. For example, US-only processing for a large job whose results can wait means Data Zone Batch: Data Zone for the location, Batch for the asynchronous, discounted processing. Reserved capacity is only ever provided by a Provisioned type, and residency only by a Data Zone or Standard/Regional scope.

Deployment options: serverless API vs managed compute. Foundry chooses how a model can be hosted based on the model itself. Serverless API is the path for Foundry Models; it offers the standard, provisioned, batch and Developer types, billed per token or by PTU, and Microsoft runs the infrastructure. Managed compute (in preview in the current Foundry) hosts open-source, partner and custom-weight models — for example open-weight models from the Hugging Face collection — on dedicated virtual machines that Foundry manages, billed hourly per accelerator (GPU) SKU whether or not requests arrive.

Tokens-per-minute quota and throttling

Tokens per minute (TPM) is the throughput allocation you assign to a deployment from your subscription's quota for that model and region; it also implies a requests-per-minute limit. When traffic exceeds it, the deployment throttles and returns rate-limit errors (HTTP 429), even though each individual request is perfectly valid.

TPM is estimated from both the prompt size and the max output tokens setting of each request, so there are two levers when throttling hits at peak: raise the TPM allocation (up to your available quota), or make each request count for fewer tokens — shorter prompts and a lower max output tokens value. Spreading requests out and retrying with back-off also help.

SymptomCauseFix
Rate-limit errors at peak traffic, each request validTPM / requests-per-minute exceededRaise TPM, lower max output tokens, smooth traffic
Input-too-long error on one large requestContext window exceededTrim or summarize the input, or use a model with a larger window
A response is cut short and marked incomplete rather than failingMax output tokens reachedRaise max output tokens

Which configuration parameters shape a model's output?

Request parameters control how varied and how long each response is; they are set per request (or in the playground), separately from the deployment.

  • Temperature (0 to 2) sharpens or flattens the probability distribution the next token is sampled from. Low values such as 0.2 keep output focused and consistent — good for legal, factual or extraction tasks. High values give more varied, creative text. Temperature 0 makes output close to deterministic.
  • Top P (top_p, nucleus sampling) limits the choice to the smallest set of tokens whose probabilities add up to the chosen value. It also controls variety, so Microsoft recommends adjusting temperature or top_p, not both.
  • Max output tokens caps how many tokens one response may generate. When a long answer reaches the cap, generation simply stops — often mid-sentence — and the response is marked incomplete instead of returning an error. It also bounds cost and TPM consumption. It is called max_output_tokens in the Responses API and max_tokens or max_completion_tokens in Chat Completions.
  • Stop sequences end generation when a given string appears.
  • Presence and frequency penalties discourage the model from repeating tokens it has already used.

Reasoning models take different parameters. Support varies by model, so check the model card: the o-series and GPT-5 reasoning models don't accept temperature, top_p, presence or frequency penalties, logit_bias or max_tokens. Cap length with max_completion_tokens (Chat Completions) or max_output_tokens (Responses API) — a limit that includes the hidden reasoning tokens — and tune thinking depth with reasoning_effort. Some newer reasoning models, such as the GPT-6 family, do accept temperature and top_p.

Tip. AI-901 tests this task with short scenarios that state one deciding constraint — no residency requirement, processing only inside the EU, reserved capacity, results that can wait a day, a short trial of a fine-tuned model, an offline device, consistent output, an image as input — and ask for the model type, deployment type or parameter that meets it. Expect near-twins side by side: tokens versus embeddings versus keywords, attention versus positional encoding, multimodal versus image-generation models, chat versus reasoning models, Global versus Data Zone versus Standard, standard versus provisioned versus batch versus Developer, serverless API versus managed compute, and temperature versus top_p versus max output tokens. Some items describe a symptom (rate-limit errors, an input-too-long error, a truncated reply) and ask for the cause or fix, and a few ask you to select two changes.

Key takeaways
  • A language model generates text by predicting a probable next token over and over; it doesn't look facts up, so it can sound confident while being wrong.
  • Tokens are words, sub-words and punctuation; context windows, max output tokens, TPM quota and pricing are all counted in tokens, so token counts exceed word counts.
  • Embeddings are vectors (fixed length for a given model and dimensions setting) of floating-point numbers; texts with similar meaning sit close together, measured with cosine similarity.
  • Attention weighs how much each token influences the others; positional encoding carries word order.
  • Input, output and reasoning tokens share one context window per request; trim or summarize history when it overflows.
  • Pick the model type from input and output: multimodal reads images, image-generation creates them, reasoning models trade latency and cost for accuracy, SLMs such as Phi run on constrained devices.
  • Global Standard is the default deployment; Data Zone and Standard/Regional restrict processing location; provisioned reserves capacity; batch is 50% cheaper with a 24-hour target; Developer is for short fine-tuned evaluation.
  • Rate-limit errors mean TPM is exceeded; replies cut off with no error mean max output tokens was reached.
  • Lower temperature (or top_p, not both) for consistent output; o-series and GPT-5 reasoning models reject temperature and top_p and cap length with max_completion_tokens or max_output_tokens.

Frequently asked questions

What is the difference between a token and an embedding?

A token is a unit of text — a word, part of a word or a punctuation mark — that a tokenizer maps to an integer ID. An embedding is a vector of floating-point numbers that represents meaning, so texts with similar meanings have vectors that are close together. Language models turn tokens into embeddings internally, and embeddings models return them for search and similarity tasks.

Why do generative AI models make up facts?

A generative AI model predicts a probable next token from patterns learned in training; it has no built-in fact lookup and doesn't know where its knowledge ends. When asked about something it never saw, or asked for a source it cannot check, it still produces fluent, plausible text that can contain invented details. Grounding prompts with current, trusted data reduces this.

When should I use a reasoning model instead of a chat model?

Use a reasoning model for multi-step problems such as math, planning, complex analysis or code, where accuracy matters more than speed. Reasoning models think before answering, which adds latency and produces reasoning tokens billed as output. For simple, high-volume, real-time chat, a general-purpose chat model such as GPT-4.1 is faster and cheaper.

What does Global Standard mean for a Microsoft Foundry model deployment?

Global Standard is a Microsoft Foundry deployment type that is billed per token and may process requests in any Azure region where the model is available. It has the broadest availability, the highest default quota and gets new models first, which is why Microsoft recommends it as the starting point when there is no data residency requirement.

What is the difference between Data Zone and Standard (regional) deployments?

A Data Zone deployment keeps inference processing within one data zone — the US, the EU or Asia-Pacific — while a Standard (regional) deployment keeps processing within the single Azure geography of the resource. Both restrict where prompts and responses are processed; Data Zone is the broader boundary and Standard the narrower one.

What causes rate-limit (429) errors on a model deployment?

Rate-limit errors occur when traffic exceeds the deployment's tokens-per-minute (TPM) allocation or its implied requests-per-minute limit. Each request may be valid on its own. You can raise the TPM allocation within your quota, reduce tokens per request (for example a lower max output tokens value), or spread requests out with retries and back-off.

What does the temperature parameter do?

Temperature, from 0 to 2, controls how random a model's output is by sharpening or flattening the probabilities it samples the next token from. Low values such as 0.2 give focused, consistent output; higher values give more varied, creative output. Microsoft recommends changing temperature or top_p, not both, and the o-series and GPT-5 reasoning models don't accept temperature at all.

What happens when a response reaches the max output tokens limit?

When a response reaches the max output tokens limit set for the request, generation stops at the cap — often mid-sentence — and the response is marked incomplete rather than returning an error. Short answers are unaffected; raising the limit (max_output_tokens, max_tokens or max_completion_tokens, depending on the API and model) lets long answers finish.

Source

This lesson covers the "Identify AI concepts and capabilities" domain of the official AI-901 exam guide. Vendors revise their guides — check the source for the current version.

Test yourself on this topic
Practice now

Sign up free to mark lessons complete, bookmark topics and track your exam readiness.

Spotted a mistake in this lesson?