AI-901 study guide — every exam topic on one page
68 key facts across 2 exam domains — every topic on the AI-901 blueprint, with the exam pattern behind each. A condensed cheat sheet, distilled from the full AI-901 revision notes. Skim it the week of your exam.
Updated
Identify AI concepts and capabilities
42% of the examMicrosoft's Responsible AI Principles for Azure AI Fundamentals (AI-901)
- The six principles are fairness, reliability and safety, privacy and security, inclusiveness, transparency and accountability; classify a measure by the problem it mainly solves.
- Fairness: similar people get similar outcomes. Bias usually comes from historical or unrepresentative training data, and removing a sensitive column does not fix it because proxy features carry the same pattern.
- Fairness is evidenced by comparing selection and error rates across groups, and improved by rebalancing or reweighting data; unequal quality of service for a group is a fairness harm.
- Reliability and safety starts from AI being probabilistic: test widely, use confidence thresholds, send weak or high-stakes cases to people, and fail safely.
- For generative AI, content filters on Microsoft Foundry deployments block harmful content in prompts and responses; adversarial testing and monitoring complete the safety picture.
- Privacy and security covers the data (minimize, de-identify, restrict, encrypt, get consent) and the model's outputs (it must not reveal private personal or organizational details).
- Inclusiveness means nobody is excluded from using or benefiting: accessibility features and diverse people in design and testing from the start.
- Transparency means people understand the system: disclose AI use, describe the data, state limitations, and explain predictions globally (overall) or locally (one decision).
- Accountability means people answer for the system: governance boards, named owners, and humans as the final authority with an appeal route.
How the exam tests this
AI-901 tests this task mostly with short scenarios that describe a measure, a design choice or a problem with an AI system and ask which responsible AI principle it relates to, or which action best addresses a given principle. Distractors are usually the near-twin principle (fairness vs inclusiveness, transparency vs accountability, reliability and safety vs accountability or privacy) or a sound practice that belongs to a different principle. Some items ask you to select two measures for one principle, or to name both principles raised by a scenario, so know the concrete practices behind each principle, not just its definition.
How Generative AI Models Work: Tokens, Model Choice and Deployment (AI-901)
- A language model generates text by predicting a probable next token over and over; it doesn't look facts up, so it can sound confident while being wrong.
- Tokens are words, sub-words and punctuation; context windows, max output tokens, TPM quota and pricing are all counted in tokens, so token counts exceed word counts.
- Embeddings are vectors (fixed length for a given model and dimensions setting) of floating-point numbers; texts with similar meaning sit close together, measured with cosine similarity.
- Attention weighs how much each token influences the others; positional encoding carries word order.
- Input, output and reasoning tokens share one context window per request; trim or summarize history when it overflows.
- Pick the model type from input and output: multimodal reads images, image-generation creates them, reasoning models trade latency and cost for accuracy, SLMs such as Phi run on constrained devices.
- Global Standard is the default deployment; Data Zone and Standard/Regional restrict processing location; provisioned reserves capacity; batch is 50% cheaper with a 24-hour target; Developer is for short fine-tuned evaluation.
- Rate-limit errors mean TPM is exceeded; replies cut off with no error mean max output tokens was reached.
- Lower temperature (or top_p, not both) for consistent output; o-series and GPT-5 reasoning models reject temperature and top_p and cap length with max_completion_tokens or max_output_tokens.
How the exam tests this
AI-901 tests this task with short scenarios that state one deciding constraint — no residency requirement, processing only inside the EU, reserved capacity, results that can wait a day, a short trial of a fine-tuned model, an offline device, consistent output, an image as input — and ask for the model type, deployment type or parameter that meets it. Expect near-twins side by side: tokens versus embeddings versus keywords, attention versus positional encoding, multimodal versus image-generation models, chat versus reasoning models, Global versus Data Zone versus Standard, standard versus provisioned versus batch versus Developer, serverless API versus managed compute, and temperature versus top_p versus max output tokens. Some items describe a symptom (rate-limit errors, an input-too-long error, a truncated reply) and ask for the cause or fix, and a few ask you to select two changes.
AI Workloads: Text Analysis, Speech, Vision and Information Extraction (AI-901)
- Identify a workload by its input and output: new content = generative AI; completed actions through tools = agentic AI; insights from existing text = text analysis; audio ↔ text = speech; labels, boxes, pixels or text from images = computer vision; structured fields from content = information extraction.
- An agent is a generative model plus instructions plus tools it decides to call; generating text or grounding answers in a knowledge base does not make an app an agent.
- Key phrases = list of topics; NER = words labelled with types; PII detection = find and mask personal data; entity linking = links to a knowledge base (the Azure Language Entity Linking and Key Phrase Extraction features are scheduled to retire in 2028 and 2029).
- Sentiment analysis gives positive, neutral or negative with confidence scores per document and sentence; opinion mining gives sentiment per aspect.
- Extractive summaries copy source sentences word for word; abstractive summaries write new sentences.
- Speech recognition is audio to text (custom speech for jargon and noise, diarization for who said what); speech synthesis is text to audio (SSML for pronunciation, rate, pitch and volume; custom voice for a unique brand voice).
- Classification = what; object detection = what and where (boxes, counts); segmentation = which pixels; OCR = what the image says.
- A multimodal model answers questions about an image; an image-generation model creates new images, and an edit with a mask targets the area the mask marks.
- OCR returns all the text; field extraction returns labelled values. Confidence scores drive straight-through processing; grounding lets reviewers check values against the source.
- Field methods: Extract copies verbatim values with source location and confidence (documents only, because only document analyzers support that grounding); Classify picks from fixed categories; Generate produces open-ended values, and is the method for open-ended fields from audio and video.
How the exam tests this
This task tests recognition, not implementation: short business scenarios where you name the workload, technique or feature that meets a stated need. Expect near-twins side by side — key phrase extraction versus entity recognition versus entity linking, document sentiment versus opinion mining, extractive versus abstractive summaries, speech recognition versus synthesis, custom speech versus SSML versus custom voice, classification versus detection versus segmentation, OCR versus field extraction, image generation versus editing with a mask, and extract versus classify versus generate fields. A single deciding detail in the stem (word for word, where each object is, a fixed set of labels, audio rather than a document, a unique brand voice) usually settles it, and some questions ask you to select two capabilities that together cover a combined requirement.
Implement AI solutions by using Microsoft Foundry
58% of the examGenerative AI Apps and Agents in Microsoft Foundry (AI-901)
- The system message carries standing rules (role, audience, tone, scope, out-of-scope behaviour, safety, format); the user prompt carries this turn's request; assistant messages are history or examples. Guardrails (content filters) screen harmful content; they do not define role or scope.
- Few-shot examples demonstrate the output and only affect the current request; in chat APIs they are example user/assistant turns after the system message. Fine-tuning is the permanent alternative.
- Reduce made-up answers by supplying grounding data, telling the model to answer only from it, giving it an out such as 'not found', and asking for citations.
- Get parseable output by describing the structure and showing an example, and express missing values inside it (for example null), never as text outside it. Token limits and stop sequences only cut a reply off. Fence pasted content with delimiters and say what it is. Against recency bias, state the task first and test repeating it after long content.
- Deploy a catalog model before using it; code passes the deployment name as model, while the resource and project names form the project endpoint.
- The model playground compares up to three models with synchronized input, offers code samples on the Code tab, and can save the setup as an agent.
- The Foundry SDK is azure-ai-projects: AIProjectClient(endpoint, DefaultAzureCredential()), then get_openai_client() and responses.create. Auth is Entra ID only: az login plus a role such as Foundry User.
- Chain turns with previous_response_id or a conversation; instructions are not carried over, so send them on every call.
- A prompt agent is a model deployment plus instructions plus optional tools, versioned in the project; hosted agents run your own container code.
- Code interpreter runs code, file search retrieves from uploaded files, web search brings current public data, and a function tool runs in your app, which returns function_call_output with the matching call_id.
How the exam tests this
Task 2.1 of AI-901 tests whether you can turn a requirement into the right prompt, portal action or line of code across the Foundry workflow: where a rule or example belongs in a prompt, which prompt change fixes a symptom such as invented answers, unparseable output or drifting replies, what must exist before a catalog model can be used and which name code passes, which playground feature fits a task, how a Python client authenticates and keeps multi-turn state, which built-in tool an agent needs, and what a client app does with a function call. Questions are short scenarios, some with a code fragment to complete, and some are multiple response.
Text Analysis and Speech Apps in Microsoft Foundry (AI-901)
- Azure Language returns consistent, structured fields; a deployed model returns flexible generated text — pick by the shape of result the app needs.
- TextAnalyticsClient takes the Foundry resource endpoint and AzureKeyCredential(key); no separate Language resource and no deployment name.
- recognize_pii_entities() is the method in the Language SDK table that returns redacted_text — the masked copy to store; its entities list describes each finding.
- Language results come back one per document in order; a bad document is a DocumentError with is_error True, not an exception.
- The OpenAI client for a Foundry deployment uses https://<resource>.openai.azure.com/openai/v1/ and the deployment name as model.
- Spoken input goes in an input_audio part as base64; spoken output needs modalities ['text','audio'] plus audio voice/format, and its words are in message.audio.transcript.
- SpeechConfig holds key, endpoint and language/voice settings; AudioConfig is recognizer input, AudioOutputConfig is synthesizer output.
- Use the recognized event and wait for session_stopped in continuous recognition; NoMatch means no speech was recognized, Canceled with Error means the request failed (most often a key or endpoint problem).
- SSML with prosody and break goes to speak_ssml_async(); Voice Live is the managed speech-to-speech API for real-time voice agents.
- The Azure Language MCP tool gives an agent text analyzers (language detection, sentiment, entities, key phrases, summarization, PII) and works on text only; audio, scanned documents and images need Speech, Content Understanding or vision.
How the exam tests this
This task tests building, not just naming: short practical scenarios with brief Python snippets, where you pick the method, argument, object or result field that completes a text-analysis or speech app. Expect answer options drawn from the same SDK, so you need to know what each similar-looking method, setting, event and result field actually does, and some items ask you to select two lines or checks that together meet a requirement.
Computer Vision and Image Generation in Microsoft Foundry (AI-901)
- Image in, words out needs a multimodal (vision-enabled) chat model; words in, image or video out needs a GPT-image, FLUX or Sora model.
- An 'image input not supported' error means a text-only deployment: deploy a vision-capable model.
- Test image understanding in the model playground chat; test image creation in the image playground.
- Send text and image together as input_text and input_image parts of one user message; compare images by putting all of them in one request.
- Image URLs must be publicly reachable; for private or local images send a base64 data URL (data:image/<type>;base64,...) or a Files API file_id.
- detail low cuts image tokens for coarse tasks; detail high is needed for small text and fine features; a smaller vision model is the other cost lever.
- Animated GIFs are read as their first frame only; send the frames you need as separate images.
- GPT-image models return base64 in data[i].b64_json (no URLs); decode it before saving. images.edit uses real input images and an optional same-size PNG mask whose transparent pixels are edited.
- The Responses image_generation tool returns an image_generation_call item and supports multi-turn refinement with previous_response_id; Sora video is create, poll, then download; for image-to-video, pass the still as input_reference at exactly the requested video size.
- A prompt that fails with an error and no image while others work has usually been blocked by content filtering.
How the exam tests this
This task tests building, not just naming: short practical scenarios and brief Python snippets where you pick the model, playground, message part, image source, parameter or response field that makes a vision or image-generation app work. Expect near-identical options drawn from the same API (input_image versus input_text, detail low versus high, images.generate versus images.edit, data[0].b64_json versus a URL), constraints that decide the answer (the storage must stay private, higher cost is acceptable, the app must not rebuild history), and some items that ask you to select two settings that together meet a requirement.
Information Extraction with Azure Content Understanding in Foundry (AI-901)
- An analyzer is a reusable configuration (content type, settings, models, optional field schema) that returns the same JSON shape every time; fields are matched by meaning, so one well-described field covers every label variant.
- prebuilt-read, prebuilt-layout and prebuilt-digitalParse need no model deployments or generative tokens; layout adds tables, sections and figures.
- RAG analyzers (documentSearch, imageSearch, audioSearch, videoSearch) return markdown plus summaries or descriptions for search; domain analyzers (invoice, receipt, idDocument for passports and driver licences, contract, callCenter) return ready-made fields, and prebuilt-invoice also covers utility bills, sales orders and purchase orders.
- prebuilt-procurement is a composed analyzer that classifies a procurement document and routes it to the matching analyzer; the utility analyzers documentFields (generic key-value pairs) and documentFieldSchema (a proposed schema) are not tuned domain schemas.
- Custom analyzers derive from one of four bases via baseAnalyzerId: prebuilt-document, prebuilt-image, prebuilt-audio or prebuilt-video; text in an image file is document content.
- Field methods: extract (verbatim, documents only), classify (fixed list in enum), generate (free-form, any content type).
- estimateFieldSourceAndConfidence adds, per field value, a confidence score (0 to 1) and its source (page and bounding region); enableLayout gives text positions without confidence, and omitContent drops the content object.
- enableSegment plus contentCategories (each with an analyzerId) splits a mixed file and routes each part; segmentPerPage splits by page only.
- Set default model deployments once per resource and override per request with modelDeployments; an analyzer's models property holds model names, not deployment names.
- ContentUnderstandingClient takes an endpoint and credential; begin_analyze returns a poller, and over REST you poll Operation-Location until Succeeded.
How the exam tests this
This task tests implementation in Microsoft Foundry rather than theory: short practical scenarios in which you choose the analyzer ID, base analyzer, field method, configuration property, model deployment setting, SDK method or result property that meets a stated requirement for documents, forms, images, audio or video. It also checks that you can read a Content Understanding result and know how a lightweight app calls the service, including the asynchronous analyze pattern. Some items ask you to select two settings that together meet a requirement, or to complete a short Python snippet.
Where these AI-901 facts came from
AI-901 cheat sheet: your questions
Source
Exam structure, domain weights and scoring on this page come from the official AI-901 exam guide.
- Microsoft AI-901 study guide — Microsoft Learn