SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
AI-901 · Domain 2

Implement AI solutions by using Microsoft Foundry practice questions

Implement AI solutions by using Microsoft Foundry is worth 58% of the AI-901 exam — the heaviest of the 2 domains. Hands-on AI in Microsoft Foundry: deploying models, prompting, building lightweight chat and agent clients with the Foundry SDK, and text, speech, vision and Content Understanding applications. Official weighting 55–60%. 6 fully worked examples are further down this page, answers included.

Exam weight
58%
the heaviest of the 2 domains
Questions
120
across 4 topics
Free, no account
5/day
sign up free to remove the cap
Explanations
Every option
right and wrong

Build a practice session

5 free questions left today.

Domains

How many?

Mode

Ready when you are

10 fresh questions drawn across 1 of 2 domains, in Learn mode.

Focused review

Every question you answer incorrectly, and every question you flag while practising, is saved here automatically. Finish a session and you can come back to re-drill just those.

6 sample Implement AI solutions by using Microsoft Foundry questions, fully explained

Questions from the AI-901 bank mapped to domain 2, with the answer key and the reasoning behind every option. None of them repeat the examples on the main AI-901 practice page.

Question 1Implement AI solutions by using Microsoft Foundry

A classification prompt works well in the Foundry playground when three labelled examples are included. The developer now builds the same request in Python with a list of role-tagged chat messages. How should the three examples be supplied so that the model treats them as demonstrations rather than as the new request?

Choose one.

  • a
    As extra system messages placed after the final user message

    Instructions placed after the real question are read as part of the current request, and system messages are not meant to carry worked example turns.

  • b
    As a fine-tuning dataset uploaded before the model is deployed

    Fine-tuning permanently adapts a model and needs a training job; few-shot examples condition only the current request and need no training.

  • c
    As a separate request sent just before the user's real question

    Each request is independent unless state is linked, so examples sent in an earlier, separate request are not seen when the real question is answered.

  • d
    As example user and assistant turns after the system message Correct

    In chat APIs, few-shot examples are given as a series of example user and assistant messages after the system message and before the real user message.

The concept

Placing few-shot examples in a chat-style request.

Why that’s the answer

With chat-style APIs, few-shot examples are supplied as example user/assistant message pairs that follow the system message and come before the real user message; the model reads them as earlier turns that show the expected answer pattern. Fine-tuning is a different, permanent customization. Examples sent in an unrelated earlier request are not part of the context, and system messages after the user's question blur the request.

How to reason it out
  1. Recall that few-shot examples are part of the prompt for this one request.
  2. In a role-tagged message list, a demonstration looks like a past user turn and the assistant's ideal reply.
  3. Place those pairs after the system message and before the real user message.
  4. Fine-tuning and separate requests do not put the examples into this request's context.

Exam tip: Few-shot in chat = example user/assistant pairs between the system message and the real question.

Generative AI Apps and Agents in Microsoft Foundry (AI-901) — the lesson that teaches this.

Question 2Implement AI solutions by using Microsoft Foundry

A company wants its chat app to answer staff questions about a travel policy that was written last month, after the model was trained. Which prompt change most directly reduces made-up answers?

Choose one.

  • a
    Add a system message line telling it not to invent any facts

    An instruction like this alone is often not an adequate mitigation; without the policy text the model still has nothing reliable to answer from.

  • b
    Lower the max output tokens so the answers stay short and focused

    A token limit truncates replies; a short answer can still be invented.

  • c
    Ask for a formal tone that matches the writing style of the policy

    Tone changes how the answer sounds, not whether its content is accurate.

  • d
    Add the policy text to the prompt and answer only from that text Correct

    Supplying the source text and instructing the model to answer from it (grounding) gives the model current, reliable data to draw on.

The concept

Grounding prompts in supplied data.

Why that’s the answer

When answers must reflect information the model was not trained on, the most effective prompt technique is to provide grounding data, here the policy text, and instruct the model to answer from it. Telling a model not to invent facts is not enough on its own because it still lacks the facts. Output length and tone do not affect accuracy.

How to reason it out
  1. Note that the policy is newer than the model's training data.
  2. The model cannot know facts it has never seen, so they must be supplied.
  3. Put the source text in the prompt and tell the model to answer from it.
  4. Length and tone settings do not make content more accurate.

Exam tip: Fresh or private facts: put the source in the prompt and tell the model to use it.

Generative AI Apps and Agents in Microsoft Foundry (AI-901) — the lesson that teaches this.

Question 3Implement AI solutions by using Microsoft Foundry

A prompt includes a product manual and tells the model to answer customer questions from it. When the manual does not cover a question, the model still produces a confident but invented answer. Which TWO prompt changes help? (Select TWO.)

Choose TWO.

  • a
    Tell the model to give longer answers with more background detail

    Longer answers give more room for invented content; length does not improve grounding.

  • b
    Tell the model to reply 'not found' when the manual lacks the answer Correct

    Giving the model an alternative path (an 'out') when the answer is not in the source helps it avoid generating a false response.

  • c
    Ask the model to cite the passage that supports each statement Correct

    Requiring citations to the source makes unsupported statements less likely, because a fabricated answer would also need a fabricated citation.

  • d
    Switch to a model whose context window fits a longer manual

    The manual already fits in the prompt; a bigger context window does not stop the model filling gaps with invented answers.

  • e
    Raise the temperature so the model weighs more possible answers

    Higher temperature makes output more random and creative, which tends to increase, not reduce, invented content.

The concept

Reducing fabrication in grounded prompts: an 'out' and citations.

Why that’s the answer

Microsoft's prompt guidance recommends giving the model an 'out', such as replying 'not found' when the answer is not in the provided text, and asking it to cite the source passage for its statements; both make it harder for the model to present invented content as fact. Longer answers, a bigger context window and a higher temperature do not address why the model fills gaps.

How to reason it out
  1. The source is already in the prompt; the failure is what happens when it is silent.
  2. Give the model a permitted response for missing answers.
  3. Require citations so each statement has to point at real source text.
  4. Discard changes to length, context size or randomness.

Exam tip: Grounded prompts need an 'out' for missing answers and citations for the rest.

Generative AI Apps and Agents in Microsoft Foundry (AI-901) — the lesson that teaches this.

Question 4Implement AI solutions by using Microsoft Foundry

An app parses each model reply as JSON with the fields name and priority. Some replies arrive as plain sentences, so parsing fails. Which change most directly fixes this?

Choose one.

  • a
    A lower max output tokens value so that just the two fields fit

    A token limit truncates the reply; a short plain sentence still fits, and a cut-off JSON reply would be invalid. It does not tell the model what structure to produce.

  • b
    A system message instruction asking for short, concise replies

    Brevity is not structure: a short plain sentence still fails JSON parsing.

  • c
    A temperature of 0 so that the model's replies stay consistent

    Temperature 0 makes replies more deterministic, but it does not define a format; the model can consistently answer in sentences.

  • d
    A description of the JSON structure plus an example of the output Correct

    Specifying the output structure, with an example, tells the model exactly what shape of reply to produce so the app can parse it.

The concept

Specifying the output structure in a prompt.

Why that’s the answer

When software consumes the reply, the prompt should state the exact output structure, here the JSON fields, and ideally show an example. Explicit structure makes the format far more consistent. A token limit truncates rather than shapes the reply, asking for brevity still allows plain sentences, and temperature 0 makes the model consistent without telling it which format to be consistent in.

How to reason it out
  1. The failure is about format, not length or randomness.
  2. The model needs to know the exact structure the app expects.
  3. Describe the fields and show an example reply.

Exam tip: If code reads the reply, specify the structure and show an example.

Generative AI Apps and Agents in Microsoft Foundry (AI-901) — the lesson that teaches this.

Question 5Implement AI solutions by using Microsoft Foundry

A user prompt starts with the instruction 'List each party's obligations', followed by a 6,000-word contract pasted between --- separators. Replies often drift into general commentary on the contract. Following Microsoft's prompt-engineering guidance, which change is worth testing first?

Choose one.

  • a
    Repeat the instruction once more after the end of the contract text Correct

    Models can show recency bias, so restating the instruction after long content often keeps the reply on task.

  • b
    Remove the --- separators placed around the pasted contract text

    Clear separators help the model tell instructions from content; removing them makes the prompt less clear.

  • c
    Lower the max output tokens so replies have no room for commentary

    A token cap truncates the reply rather than refocusing it: the obligations list can be cut off while commentary still appears first.

  • d
    Send the contract in several separate requests, one section each

    Unlinked requests do not share context, so each part would be processed without the others and without a combined result.

The concept

Recency bias and repeating instructions after long content.

Why that’s the answer

Microsoft's guidance notes that models can be susceptible to recency bias: information at the end of a prompt can influence the output more than information at the start. With a long document between the instruction and the end of the prompt, repeating the instruction after the content (doubling down) is a recommended first experiment. Separators help rather than hurt, a token cap truncates the reply instead of refocusing it, and splitting into unlinked requests loses context.

How to reason it out
  1. Note the layout: a short instruction, then a very long block of content.
  2. Recall that later text in a prompt can weigh more heavily (recency bias).
  3. Restate the instruction after the content and compare the results.
  4. Keep delimiters; they make the prompt clearer.

Exam tip: Long content pushes your instruction out of focus: say it again at the end.

Generative AI Apps and Agents in Microsoft Foundry (AI-901) — the lesson that teaches this.

Question 6Implement AI solutions by using Microsoft Foundry

A developer finds a chat model in the Microsoft Foundry model catalog and wants to test it in the playground and then call it from a Python app. Following the standard workflow, what must the developer create first?

Choose one.

  • a
    A fine-tuning job that adapts the model to the app's own data

    Fine-tuning is optional customization; a base model can be used as soon as it is deployed.

  • b
    An agent version that references the model by its catalog name

    A prompt agent itself needs a deployed model; creating an agent is not the first step for using the model.

  • c
    A deployment of the model in the Foundry resource Correct

    Deploying the model creates the named deployment that the playground and API calls use.

  • d
    A hosted agent container image that wraps the chosen model

    Hosted agents run your own agent code; a container is not needed to chat with a catalog model.

The concept

Deploying a model before using it.

Why that’s the answer

In Foundry, you deploy a model from the catalog to get a named deployment; the playground opens on that deployment and application code calls it by name. Fine-tuning, agents and hosted agent containers are later or optional steps. (Foundry also previews 'instant' models that skip deployment in limited regions, but deploying is the standard path.)

How to reason it out
  1. Picking a model in the catalog does not make it callable yet.
  2. Deploy it to create a named endpoint target.
  3. Use the deployment in the playground and from code.

Exam tip: Catalog model → deployment → playground and code.

Generative AI Apps and Agents in Microsoft Foundry (AI-901) — the lesson that teaches this.

What AI-901 domain 2 tests, topic by topic

The official exam guide breaks Implement AI solutions by using Microsoft Foundry into 4 topics. The question bank follows the same split, so a weak topic shows up as a cluster of misses you can go back and read.

Published AI-901 practice questions per topic in Implement AI solutions by using Microsoft Foundry
TopicWhat it coversQuestions
Implement generative AI apps and agents by using FoundrySkills outline section (AI-901, as of April 15, 2026). Create effective system and user prompts for generative AI models; deploy a model and interact with it in the Foundry portal; create a lightweight chat client application by using the Foundry SDK; create and test a single-agent solution in the Foundry portal; create a lightweight client application for an agent.30
Implement AI solutions for text and speech by using FoundrySkills outline section (AI-901, as of April 15, 2026). Build a lightweight application that includes text analysis; respond to spoken prompts by using a deployed multimodal model; build a lightweight application by using Azure Speech in Foundry Tools.30
Implement AI solutions with computer vision and image-generation capabilities by using FoundrySkills outline section (AI-901, as of April 15, 2026). Interpret visual input in prompts by using a deployed multimodal model; create new visual outputs by using generative models; build a lightweight application that includes vision capabilities.30
Implement AI solutions for information extraction by using FoundrySkills outline section (AI-901, as of April 15, 2026). Extract information from documents and forms, from images, and from audio and video by using Azure Content Understanding in Foundry Tools; build a lightweight application with information extraction capabilities by using Content Understanding.30
Total120

Revise Implement AI solutions by using Microsoft Foundry before you drill it

Other AI-901 domains

Implement AI solutions by using Microsoft Foundry: your questions

Implement AI solutions by using Microsoft Foundry is domain 2 of the AI-901 exam guide and carries 58% of the scored content — the heaviest of the 2 domains. On a 45-question paper that works out to roughly 26 questions, though Microsoft Azure does not publish an exact per-domain count and individual exam forms vary.

Source

The domain weight and topic list on this page come from the official AI-901 exam guide.