Generative AI Apps and Agents in Microsoft Foundry (AI-901)
Building generative AI apps and agents in Microsoft Foundry means writing effective prompts, deploying a model from the Foundry model catalog, calling it from a small Python client through the Foundry SDK, and packaging instructions, a model and tools into a prompt agent that a client app can chat with. This AI-901 topic covers the whole path: what goes in the system message versus the user prompt, few-shot examples, grounding and output structure, model deployments and the model playground, the azure-ai-projects client and the Responses API, agent tools and the agents playground, and the function-call round trip a client app handles.
On this page10 sections
- What belongs in the system message and what belongs in the user prompt?
- How do zero-shot, one-shot and few-shot prompts differ?
- How do you ground a prompt to reduce made-up answers?
- How do you get output an app can parse, and write clear prompts?
- How do you deploy a model and try it in the Foundry portal?
- How do you build a lightweight chat client with the Foundry SDK?
- How does a chat client keep the conversation across turns?
- What is a prompt agent, and which tools can it use?
- How do you test a single agent in the agents playground?
- How does a lightweight client app chat with an agent?
- Write system messages and user prompts that put standing rules, the specific request and examples in the right place.
- Use few-shot examples, grounding data, an explicit output structure and clear syntax to make replies accurate and parseable.
- Deploy a model from the Foundry model catalog and test it in the model playground, including compare mode and the Code tab.
- Build a lightweight Python chat client with azure-ai-projects, Entra ID authentication and the Responses API, including multi-turn state.
- Create a prompt agent from a model deployment, instructions and the right built-in tools, and test it in the agents playground.
- Build a client app that chats with an agent through a conversation and completes the function-tool round trip.
What belongs in the system message and what belongs in the user prompt?
The system message holds the standing rules that apply to every turn; the user prompt holds the specific request for this turn. A chat request to a generative AI model is a list of role-tagged messages, and each role has a different job:
| Message role | What it carries | Example |
|---|---|---|
| System | The assistant's role and task, its audience and tone, scope and boundaries (what to do when a request is out of scope), safety rules, and optionally the tools and data it may use and the output format | "You are a patient booking assistant for Fabrikam Travel. Answer only questions about trips and bookings, and point medical questions to a doctor." |
| User | The person's request and any content to work on in this turn | "Can I change my flight to Friday?" |
| Assistant | The model's earlier replies (conversation history), or example replies in a few-shot prompt | "Yes. Friday has two seats left on…" |
Because the app sends the system message with every request and the model treats it as high-priority guidance, it is the place for anything that must hold whatever the user types: tone, staying on topic, refusing certain requests. Putting those rules in user prompts depends on each user repeating them, which never happens.
Safety screening is a separate layer. Content filters, which the current Foundry portal presents as guardrails (built on Azure AI Content Safety), detect and block harmful categories of input and output. They are a safety control, not a way to describe the assistant's role, tone or subject area; those behaviours are written as instructions.
How do zero-shot, one-shot and few-shot prompts differ?
A zero-shot prompt gives instructions only; a one-shot or few-shot prompt adds one or several worked examples, usually input and output pairs, so the model can copy the exact behaviour and format. Examples are the most reliable way to get a precise reply shape, such as a single classification label with no explanation, because they demonstrate the output instead of describing it. Describing a format in words leaves the model to interpret the description; an example removes the guesswork.
Few-shot examples condition the model for the current request only; nothing is learned permanently. That is the difference from fine-tuning, which trains a customized model on a dataset before it is deployed.
In a chat API, examples are supplied as example user and assistant turns placed after the system message and before the real user message. The model reads them as earlier conversation that shows the pattern, then answers the final user message in the same way:
messages = [
{"role": "system", "content": "Rate the sentiment of each product review as Positive, Negative or Mixed. Reply with the rating only."},
{"role": "user", "content": "Arrived quickly and works perfectly."},
{"role": "assistant", "content": "Positive"},
{"role": "user", "content": "Great screen, but the battery dies by lunchtime."},
{"role": "assistant", "content": "Mixed"},
{"role": "user", "content": "<the next review>"}
]
The examples work because they sit inside the same request as the real input: a model only sees what the current request contains, plus anything the request is explicitly chained to.
How do you ground a prompt to reduce made-up answers?
You ground a prompt by putting the facts the answer needs into the prompt and telling the model to answer only from them. A model knows nothing that happened after its training data was collected and nothing about your private documents, so a policy written last month or an internal manual must be supplied as grounding data. Telling the model "do not invent facts" on its own does not help much: it still lacks the facts.
Microsoft's prompt-engineering guidance adds two techniques that make invented content harder to slip through:
- Give the model an "out". Tell it what to reply when the supplied text does not contain the answer, for example "not found" or "I don't have that information". Without an allowed alternative, a model tends to fill the gap with a confident guess.
- Ask for citations. Require it to quote or reference the passage that supports each statement, which ties every claim to the source and makes unsupported claims visible.
The underlying principle: fabrication is a missing-facts problem. Generation parameters and style requests change how the model writes, not what it knows, so the fix is always to supply the facts and constrain the answer to them. For large or changing document sets, grounding is automated with retrieval, for example the file search or Azure AI Search tool on an agent (covered below).
How do you get output an app can parse, and write clear prompts?
To get parseable output, state the exact structure in the prompt and show an example of it. If an app reads each reply as JSON with the fields task and owner, the prompt should describe those fields and include a sample object; you can also "prime" the output by starting it, for example ending the prompt with the opening of the expected format. Generation parameters are not a substitute: a max output tokens limit truncates a reply and a stop sequence ends it as soon as a chosen string appears, so both only decide where the reply stops, not what it contains or how the model reads the prompt; a low temperature makes replies more repeatable in whatever format the model happened to choose. Only an explicit format, ideally demonstrated, defines the structure.
The "out" must fit the structure too. Giving the model an allowed answer for missing information (see grounding above) still applies when an app parses the reply, but that answer has to be expressed inside the format. If a value is absent from the source, the prompt should say how to represent it in the structure, for example null for that JSON field. A plain-text "not found" in place of the object, or any extra line written before or after the object (a note, a source line), breaks the parser just as a free-text reply does. Put supporting detail, such as a citation, in a field of the structure if the app needs it.
Clear prompts follow a few habits from the same guidance:
- Be specific about the task, audience and length instead of leaving the model to guess.
- Use clear syntax: delimiters such as
---, headings or quotation marks mark where pasted content starts and ends, and the prompt should also say what the delimited block is, for example "the text between the --- lines is a document to summarize". Without both, a sentence inside the pasted content (a request in an email, say) can be read and obeyed as if it were part of your instructions. - Break the task down into smaller steps when one prompt asks for too much.
- Mind the order (recency bias). Information at the end of a prompt can influence the output more than information at the start. State the task clearly at the start; when a long document then sits between the instruction and the end of the prompt and replies drift, repeating the instruction after the content is worth testing, and evaluating the effect.
One caveat from the same Microsoft guidance: these techniques are written for standard chat models and are not recommended for reasoning models (such as the o-series and gpt-5 models), which plan their own steps and respond best to short, direct instructions.
How do you deploy a model and try it in the Foundry portal?
You pick a model in the Microsoft Foundry model catalog and deploy it to your Foundry resource; the result is a named deployment, and both the playground and application code use that deployment. The standard order is: choose a model, check it is available in your region, deploy it, then test it. Fine-tuning, agents and hosted agents are later, optional steps that all build on a deployed model. (Foundry also previews "instant" models that can be used without a deployment in limited regions; deploying remains the standard workflow.)
The deployment name is what code passes as the model. When you deploy, you name the deployment, and that name, shown in the Name column of the project's deployed models, is the value of the model argument. Tutorials often name the deployment after the model (gpt-4o deployed as "gpt-4o"), which hides the difference: if gpt-4o is deployed as trip-helper, code passes model="trip-helper". The resource and project names go into the project endpoint instead:
https://<resource-name>.services.ai.azure.com/api/projects/<project-name>
The model playground is where you interact with a deployment without code:
- Write a system message and adjust parameters such as temperature and max output tokens, then chat to see the effect.
- Compare up to three models side by side: the input is synchronized, so each receives the same prompt, system message and parameters, and the responses stream next to each other.
- The Code tab gives sample code in several languages for the deployment you are testing, and Open in VS Code (for the Web) imports the sample and endpoint into an editor.
- Save as agent turns the current model, instructions and tools into an agent (see below).
How do you build a lightweight chat client with the Foundry SDK?
A Foundry SDK chat client in Python creates an AIProjectClient from the azure-ai-projects package, gets an OpenAI client from it with get_openai_client(), and calls responses.create with the deployment name:
from azure.identity import DefaultAzureCredential
from azure.ai.projects import AIProjectClient
project = AIProjectClient(
endpoint="https://fabrikam-ai.services.ai.azure.com/api/projects/travel-project",
credential=DefaultAzureCredential(),
)
client = project.get_openai_client()
response = client.responses.create(
model="trip-helper", # the deployment name
instructions="You are a patient booking assistant.",
input="Can I change my flight to Friday?",
)
print(response.output_text)
The pieces and their jobs:
| Package / object | Job |
|---|---|
azure-ai-projects / AIProjectClient | The Foundry SDK: connects to a project; its get_openai_client() returns an authenticated OpenAI client for chatting with models and agents |
azure-identity / DefaultAzureCredential | Supplies the Microsoft Entra ID credential; it is not a client |
Text generation always goes through the OpenAI client that the project returns; the project client's own operation groups manage project resources rather than answering prompts. Libraries for individual Azure AI services, and older standalone inference clients, are separate packages from the Foundry project SDK.
Authentication is Microsoft Entra ID only for the project client; you do not pass an API key. DefaultAzureCredential needs a signed-in identity, which on a developer laptop usually comes from az login, and that identity needs a role on the Foundry resource or project. Foundry User (formerly named Azure AI User) is the least-privilege role for developers. So a new machine needs both: a signed-in identity for the credential to use, and a role assignment that authorizes that identity.
How does a chat client keep the conversation across turns?
Each responses.create call is independent: the model sees only what that call sends, so a follow-up such as "How long does it take?" makes no sense to it unless the call is linked to earlier turns. The Responses API offers two service-side ways to link them, so the client does not have to resend the history:
previous_response_id: pass theidof the last response, and the service supplies that exchange (and the chain before it) as context.- A conversation: create one with
conversations.create()and pass its id on every call; the service stores the items of each turn in it.
first = client.responses.create(model="trip-helper", instructions=RULES, input="Suggest a scenic train route in Switzerland.")
follow = client.responses.create(
model="trip-helper",
instructions=RULES, # sent again on every call
previous_response_id=first.id,
input="How long does it take?",
)
Two details decide many design choices:
instructionsapply only to the response they are sent with. They are not carried over when you chain withprevious_response_id, and a conversation stores message items, not instructions. To keep system rules on every turn, send them with every call.- Chaining needs stored responses. The service can only supply an earlier exchange it has kept. Responses are stored by default; a response created with storage switched off cannot be chained to later.
The older pattern, a system message at the top of a messages list with the whole history resent each turn, also works, but it keeps the history on the client. Only these explicit links carry context; nothing else about two calls (the same deployment, the same client object) connects them.
What is a prompt agent, and which tools can it use?
A prompt agent in Foundry Agent Service is an agent defined entirely by configuration: a model deployment, instructions and optional tools. You create it in the Foundry portal (or with the SDK) without writing agent code, and Foundry runs it for you. The definition is stored in the project, and each saved change creates a new version of the agent. A hosted agent is the code-first alternative: you package your own agent code in a container that Foundry runs.
Moving behaviour from a raw model call into a prompt agent changes where it lives: the instructions sit in the agent definition, so the client sends only the user's input. Because every client calls the agent by name, editing its instructions or tools in the portal changes how it behaves for all of those clients at once, with no change to or redeploy of the client code. The agent still uses a model deployment, and it still needs a conversation or previous response id to remember earlier turns.
Tools extend what the agent can do. Choose them by what the task needs:
| Tool | Use it when the agent must… |
|---|---|
| Code interpreter | Write and run Python in a sandbox: calculations, transforming uploaded files such as CSVs, drawing charts |
| File search | Answer from documents you upload to the agent (grounding in your own files) |
| Web search | Bring in current public information from the web |
| Azure AI Search | Ground answers in an existing Azure AI Search index |
| Function tool | Ask your application to run its own code, for example to reach an internal system only the app can access |
| OpenAPI tool | Call an external HTTP API described by an OpenAPI specification |
| MCP | Use tools exposed by a Model Context Protocol server |
The near-twins to keep apart: code interpreter runs code in an isolated sandbox, so it cannot reach your internal network; file search retrieves text from uploaded files; a function tool's code runs in your app. Match the tool to where the data lives and how fresh it must be: public and current, uploaded documents, an existing index, or a system only your application can reach.
How do you test a single agent in the agents playground?
You test an agent in the agents playground of the Foundry portal before writing any client code. There you can set and edit the agent's instructions, add tools and knowledge, hold multi-turn conversations, and inspect tracing for each response (which tools were called, with what input and output) and evaluation data. Tuning the instructions there and re-testing is the quickest loop; each saved change becomes a new agent version.
| Portal surface | What it is for |
|---|---|
| Model playground | One deployment: system message, parameters, compare up to three models, Code tab, Save as agent |
| Agents playground | One agent: instructions, tools, knowledge, multi-turn chat, tracing and evaluation |
How does a lightweight client app chat with an agent?
A client app chats with an existing agent by getting an OpenAI client from the project that is bound to the agent, creating one conversation, and passing that conversation's id with every responses.create call, so every turn shares the history:
client = project.get_openai_client(agent_name="TripPlanner")
conversation = client.conversations.create()
for question in ["Plan three days in Lisbon.", "Make day two cheaper."]:
response = client.responses.create(conversation=conversation.id, input=question)
print(response.output_text)
Two principles sit behind this pattern. The conversation is created once and reused, because it is the conversation that holds the history. And it is the agent binding that applies the agent's stored instructions and tools; managing the agent's definition (creating versions) is a separate, administrative task from chatting with it.
Function tools need a round trip through your app. The service never runs a function tool's code. When the model decides to use one, the response contains a function_call item, with the function name, its arguments and a call_id, instead of answer text. The app then:
- reads the function name and arguments from the
function_callitem; - runs its own function (for example, the flight-status lookup);
- sends the result back in the same conversation (or chained to the response) as a
function_call_outputitem carrying the samecall_id; - reads the model's final answer from the next response.
for item in response.output:
if item.type == "function_call" and item.name == "get_flight_status":
result = get_flight_status(**json.loads(item.arguments))
response = client.responses.create(
conversation=conversation.id,
input=[{"type": "function_call_output", "call_id": item.call_id, "output": json.dumps(result)}],
)
print(response.output_text)
Until the app sends that output back, the model has nothing to answer with, which is why the function tool is the way to reach systems that only your application can see.
Tip. Task 2.1 of AI-901 tests whether you can turn a requirement into the right prompt, portal action or line of code across the Foundry workflow: where a rule or example belongs in a prompt, which prompt change fixes a symptom such as invented answers, unparseable output or drifting replies, what must exist before a catalog model can be used and which name code passes, which playground feature fits a task, how a Python client authenticates and keeps multi-turn state, which built-in tool an agent needs, and what a client app does with a function call. Questions are short scenarios, some with a code fragment to complete, and some are multiple response.
- The system message carries standing rules (role, audience, tone, scope, out-of-scope behaviour, safety, format); the user prompt carries this turn's request; assistant messages are history or examples. Guardrails (content filters) screen harmful content; they do not define role or scope.
- Few-shot examples demonstrate the output and only affect the current request; in chat APIs they are example user/assistant turns after the system message. Fine-tuning is the permanent alternative.
- Reduce made-up answers by supplying grounding data, telling the model to answer only from it, giving it an out such as 'not found', and asking for citations.
- Get parseable output by describing the structure and showing an example, and express missing values inside it (for example null), never as text outside it. Token limits and stop sequences only cut a reply off. Fence pasted content with delimiters and say what it is. Against recency bias, state the task first and test repeating it after long content.
- Deploy a catalog model before using it; code passes the deployment name as model, while the resource and project names form the project endpoint.
- The model playground compares up to three models with synchronized input, offers code samples on the Code tab, and can save the setup as an agent.
- The Foundry SDK is azure-ai-projects: AIProjectClient(endpoint, DefaultAzureCredential()), then get_openai_client() and responses.create. Auth is Entra ID only: az login plus a role such as Foundry User.
- Chain turns with previous_response_id or a conversation; instructions are not carried over, so send them on every call.
- A prompt agent is a model deployment plus instructions plus optional tools, versioned in the project; hosted agents run your own container code.
- Code interpreter runs code, file search retrieves from uploaded files, web search brings current public data, and a function tool runs in your app, which returns function_call_output with the matching call_id.
Frequently asked questions
What is the difference between a system message and a user prompt in Microsoft Foundry?
The system message sets the standing rules the model follows on every turn, such as its role, audience, tone, scope and what to do with out-of-scope requests. The user prompt is the specific request a person makes in the current turn. Rules that must hold whatever the user types belong in the system message.
Which value do I pass as the model when calling a Foundry deployment from code?
Pass the deployment name, the name you gave the deployment when you deployed the model from the Foundry model catalog. It differs from the catalog model name whenever you chose a different name. The Foundry resource and project names are part of the project endpoint, not the model argument.
Which Python package is the Microsoft Foundry SDK?
The Foundry SDK for Python is azure-ai-projects. Code creates an AIProjectClient with the project endpoint and a DefaultAzureCredential from azure-identity, then calls get_openai_client() to get an authenticated OpenAI client for responses and conversations with models and agents.
Can the Foundry project client authenticate with an API key?
No. The Foundry project client supports Microsoft Entra ID only. On a developer machine, DefaultAzureCredential typically uses the identity from az login, and that identity needs a role on the Foundry resource or project, such as the least-privilege Foundry User role.
Are instructions kept when chaining responses with previous_response_id?
No. In the Responses API, the instructions parameter applies only to the response it is sent with and is not carried over when a later call chains with previous_response_id. A conversation stores message items, not instructions. Send the instructions on every call to keep the rules in force.
What is a prompt agent in Foundry Agent Service?
A prompt agent is an agent defined by configuration: a model deployment, instructions and optional tools such as code interpreter, file search or web search. It is stored and versioned in the Foundry project and run by Foundry with no code to host. A hosted agent, by contrast, runs your own code in a container.
Who runs the code of an agent's function tool?
Your application does. The model returns a function_call item with the function name, arguments and a call_id. The app runs the function, sends the result back as a function_call_output item with the same call_id in the same conversation, and the model then produces its final answer.
Where do I test an agent before writing a client app?
Use the agents playground in the Foundry portal. It lets you edit the agent's instructions and tools, hold multi-turn conversations, and review tracing and evaluation data for each response, all without writing code.
Source
This lesson covers the "Implement AI solutions by using Microsoft Foundry" domain of the official AI-901 exam guide. Vendors revise their guides — check the source for the current version.
- Microsoft AI-901 study guide — Microsoft Learn
Sign up free to mark lessons complete, bookmark topics and track your exam readiness.