SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
Implement AI solutions by using Microsoft Foundry

Computer Vision and Image Generation in Microsoft Foundry (AI-901)

16 min readAI-901 · Implement AI solutions by using Microsoft FoundryUpdated

Implementing computer vision and image generation in Microsoft Foundry means sending images to a deployed multimodal model so it can describe, compare or answer questions about them, and calling generative models to create or edit images and videos from a prompt. This lesson covers the AI-901 skills for that task: choosing the right model, testing in the Foundry portal playgrounds, building Responses API messages with input_text and input_image parts, supplying images as URLs, base64 data URLs or file IDs, controlling cost with the detail setting, generating and editing images with the images API, conversational image creation with the image_generation tool, asynchronous Sora video jobs, and the safety filters that block harmful prompts.

What you’ll learn
  • Choose between a multimodal chat model, an image generation model and a video generation model for a vision requirement.
  • Test image understanding in the model playground and image creation in the image playground before writing code.
  • Build a Responses API request that combines input_text and input_image parts, supplying the image as a public URL, base64 data URL or file ID.
  • Use the detail setting, model size, multiple image parts and structured output to make a vision app accurate, affordable and parseable.
  • Generate and edit images with images.generate and images.edit, including masks, and decode the base64 results.
  • Build multi-turn image creation with the Responses API image_generation tool and run Sora video generation as an asynchronous job.
  • Recognise when content filtering blocks an image or video prompt.

Which model do you need: one that understands images or one that creates them?

Pick by the direction of the data: if an image goes in and words come out, you need a multimodal (vision-enabled) chat model; if words go in and an image or video comes out, you need a generative image or video model. Microsoft Foundry hosts both kinds in its model catalog, and they are deployed and called differently.

RequirementModel typeExamples in FoundryTypical call
Describe, classify, compare or answer questions about a photoMultimodal chat model with image inputVision-capable GPT-4.1 / GPT-5 series modelsclient.responses.create() with an input_image part
Create a new still image from a text promptImage generation modelGPT-image models (for example GPT-Image-1 and GPT-Image-1-Mini); FLUX models from Black Forest Labsclient.images.generate()
Change an existing image (add, replace or repaint parts of it)Image generation model with editingGPT-image modelsclient.images.edit()
Create a short video clip from a promptVideo generation modelSoraAsynchronous /openai/v1/videos jobs

A text embedding model is neither: it turns input into vectors for search and similarity, and never describes or draws a picture. And not every chat model accepts images. If a request that includes an image part fails with an error saying image input is not supported, while the same request without the image works, the deployment is a text-only model. The fix is the model, not the message: deploy a vision-capable model and point the app at that deployment.

This lesson is about building with these models. What computer vision and image generation are as workloads is covered in the Task 1.3 lesson, and extracting fields or text from images into a fixed schema with Content Understanding is covered in the Task 2.4 lesson.

How do you try vision and image generation in the Foundry portal before writing code?

Use the playground that matches the model: the model (chat) playground for image understanding and the image playground for image generation. Both open from a deployment in the Foundry portal and let you experiment with no code.

  • Model playground (chat). With a vision-capable chat deployment selected, attach an image to the chat and type a question about it, such as asking what is damaged in a photo. You see exactly how the model interprets your images and can refine the system instructions and prompt before coding.
  • Image playground. When you deploy a text-to-image model, the portal offers an image playground where you type a prompt, choose options such as size and number of images, and review the generated pictures. Uploading a photo to a chat is not how you test image creation.
  • Video playground. Sora deployments have a matching playground for trying text-to-video prompts.

The playgrounds are the fastest way to see how a deployment behaves; heavier tools for improving or measuring a model at scale come later. The playgrounds can also show sample code for the settings you tried, which is a convenient starting point for an app.

How do you send an image to a multimodal model in code?

Send one user message whose content is a list of parts: an input_text part carrying the question and an input_image part carrying the picture. With the Responses API and the OpenAI Python SDK pointed at the Foundry resource's /openai/v1/ endpoint (client setup is covered in the Task 2.1 lesson), a lightweight vision app looks like this:

response = client.responses.create(
    model=deployment_name,          # a vision-capable chat deployment
    instructions="You inspect parcel photos for a shipping company.",
    input=[{
        "role": "user",
        "content": [
            {"type": "input_text", "text": "Which part of this parcel is damaged?"},
            {"type": "input_image", "image_url": photo_url}
        ]
    }]
)
print(response.output_text)
  • The part types are input_text and input_image. Putting a URL inside an input_text part only gives the model the characters of the link, not the picture; input_file is for documents such as PDFs; output_* types describe what the model returns, not what you send.
  • Don't mix up the two APIs' part names. The older Chat Completions API uses text and image_url as its content part types. In the Responses API those names are replaced by input_text and input_image, and image_url survives only as a field inside an input_image part (alongside the optional detail), not as a part type of its own.
  • System-level guidance (role, tone, what to look for) goes in instructions or a system/developer message; the image itself belongs in the user message.
  • The text answer is read from response.output_text, exactly as in a text-only chat app.

Where can the image come from: URL, base64 data URL or file ID?

The model can receive an image in three ways, and the right one depends on whether the model's service can fetch the image itself.

SourceWhat you sendUse when
Public URL"image_url": "https://…/photo.jpg"The image is hosted at an address anyone on the internet can read
Base64 data URL"image_url": f"data:image/jpeg;base64,{b64}"The app holds the bytes: a local file, an upload, or a private blob it has read itself
Uploaded file"file_id": file_id (from the Files API)You upload the image once and reference it by ID

The service fetches URLs itself, so a URL must be publicly reachable. Images behind a private endpoint, inside a virtual network, on localhost or addressed as file:/// paths cannot be fetched, and granting the Foundry resource's managed identity access to storage does not change how the image URL is fetched. When the image must stay private, have the app read the bytes and send them inline as a base64 data URL (or upload them as a file and pass the file_id).

A data URL has a fixed shape: data:, the image's MIME type (for example image/jpeg or image/png), ;base64,, then the base64 text. In Python that means base64.b64encode(data).decode("utf-8") and then f"data:image/jpeg;base64,{b64}". Base64 changes how the image travels, not how many tokens the model spends looking at it.

How do you balance accuracy and cost when sending images?

Use the detail setting on each input_image part, and choose the model size, to trade visual precision against tokens, latency and cost.

detailWhat the model doesGood for
lowLooks at a downscaled version of the image for a small, fixed token costCoarse tasks: sorting photos into a few categories, "is this photo taken indoors or outdoors?"
highExamines the image in more detail, spending more input tokens and timeFine detail: small printed text, gauge readings, tiny defects
auto (default)The model chooses based on the imageGeneral use; it may still pick the high-detail path, so it is not a cost control

Two practical rules follow. If a high-volume job only needs coarse labels and input tokens are the problem, setting detail to low is the direct fix. If the model misreads small text after you lowered detail, set it back to high. The principle: only a higher detail level adds visual information. Settings that shape the reply (prompt wording, sampling, output length) cannot restore pixels the model never saw, and they do not change the tokens spent reading the image.

The second lever is the model itself. A large flagship model is wasted on simple sorting; a smaller, cheaper vision-capable model often classifies just as accurately at a fraction of the cost. As a rule, cost falls when the model or the image input gets lighter, not when the model is asked to do more work.

How do you compare several images and get answers your code can parse?

To have the model reason across several pictures, put every image and the instruction in the same request, each image as its own input_image part. A model only sees what is in the current request: sending images in separate calls and asking the second call to "remember" the first does not work unless the earlier context is passed along, and one input_image part holds exactly one image, so you cannot list two URLs in it.

"content": [
    {"type": "input_text", "text": "List every change between last week's building-site photo and today's."},
    {"type": "input_image", "image_url": last_week_url},
    {"type": "input_image", "image_url": today_url}
]

Supported image formats include PNG, JPEG, WEBP and non-animated GIF. An animated GIF is processed as a single frame (the first one), so a model only ever "sees" what that frame shows. Because the other frames never reach the model, no setting or prompt can recover them; when several frames matter, extract those frames in the app and send each as a separate input_image part. (Analyzing real video content is a different task: Content Understanding, in the Task 2.4 lesson, handles video.)

Structured output for app code

Give the model clear instructions and, when code consumes the answer, request structured output with a JSON schema. Without a schema the model may describe a photo in a different layout each time, which breaks a parser even when the content is right.

  • Instructions set the task and constraints: what to look for, what to ignore, what to do if the image is unclear.
  • A JSON schema (passed as the structured output format of the request) fixes the shape of the reply, for example an array of objects with defect and location fields for an inspection photo. The model's reply then conforms to the schema, so the app can load it as JSON.
  • The principle: only a schema fixes the shape of the reply. Prompt wording and image settings change what the model notices and how it phrases things, not the structure your parser receives.

When the goal is to pull fixed fields out of many standard images or documents (invoices, IDs, receipts) with confidence scores, Content Understanding analyzers are the purpose-built option; see the Task 2.4 lesson.

How do you generate a new image with the images API?

Call client.images.generate() on an image generation deployment with a text prompt; GPT-image models return each picture as base64 in data[i].b64_json, which the app decodes and saves.

import base64

img = client.images.generate(
    model=image_deployment,
    prompt="A flat-style banner of a lighthouse at dawn",
    n=1,
    size="1024x1024",
    quality="medium",
    output_format="png"
)
with open("banner.png", "wb") as f:
    f.write(base64.b64decode(img.data[0].b64_json))
ParameterControls
promptWhat to draw; be specific about subject, style, composition and colours
nHow many images one request returns (each in its own data entry)
sizeOutput dimensions, such as square, portrait or landscape sizes the model supports
qualityRendering quality versus speed and cost (for example low, medium, high)
output_formatFile format of the returned image: png or jpeg
backgroundCan request a transparent background, which needs PNG output
partial_imagesStreams in-progress previews while the image renders; it does not change how many final images you get

Two common mistakes: GPT-image models do not return download URLs and do not accept response_format, so code that reads data[0].url fails; and the base64 string must be decoded before writing it to disk, because writing the encoded text produces a file no image viewer can open.

Choosing an image model is a fidelity-versus-cost decision. A lighter model such as GPT-Image-1-Mini suits high-volume, low-stakes work like rough thumbnails and drafts; the flagship GPT-image models suit brand-critical visuals where detail and prompt accuracy matter most; FLUX models are a third-party alternative in the same catalog. A cheap chat model that accepts images is not an image generator at all.

How do you edit an existing image, with or without a mask?

Call client.images.edit() with one or more input images and a prompt; add an optional mask to restrict which pixels may change. Use editing whenever the result must contain the real content of an existing picture. images.generate only knows your words, so describing a mascot in text produces a new mascot, not yours.

NeedCall
A brand-new picture from a descriptionimages.generate(prompt=…)
Combine real images (for example a team mascot and a stadium photo) into a new compositionimages.edit(image=[logo, product], prompt=…)
Change only one region of a photo and keep the rest pixel-for-pixelimages.edit(image=photo, mask=mask, prompt=…)

A mask must follow two rules:

  • It is a PNG with an alpha channel, and its fully transparent pixels mark the area the model may change; opaque pixels are kept. (It is the transparent region that gets edited, not the opaque one.)
  • It has exactly the same dimensions as the input image. A smaller or JPEG mask is rejected, since JPEG has no transparency.

A vision chat model cannot do this job: it reads images and replies in text, it does not return a composed picture.

How do you build conversational image creation with the Responses API?

Give a Responses API request the image_generation tool: the chat model decides when to draw, a GPT-image deployment renders the picture, and the result comes back as an image_generation_call item in response.output whose result holds the base64 image.

response = client.responses.create(
    model=chat_deployment,
    input="Design a poster for a jazz night",
    tools=[{"type": "image_generation"}]
)
images = [o.result for o in response.output if o.type == "image_generation_call"]
with open("poster.png", "wb") as f:
    f.write(base64.b64decode(images[0]))

follow_up = client.responses.create(
    model=chat_deployment,
    previous_response_id=response.id,
    input="Make the background deep blue and add the title 'Jazz Night'",
    tools=[{"type": "image_generation"}]
)

The tool needs a GPT-image deployment as well as the chat deployment. The snippet above is simplified: in Foundry the image deployment is configured for the tool separately (the Responses how-to passes its name in the x-ms-oai-image-generation-deployment request header), so make sure both deployments exist.

This pattern fits chat-style design assistants where each turn refines the last picture. Passing previous_response_id links the new request to the earlier one, so the service carries the conversation, including the previous image, and the app does not rebuild and resend the whole history each turn. Calling images.generate afresh on every turn loses the earlier image, and a vision chat model on its own can describe the last poster but cannot redraw it.

Note where the output lives: each API has its own response shape. The Responses tool returns the image inside an output item; output_text carries only text, and data[i].b64_json belongs to the images API.

How do you generate video with Sora, and what stops harmful images or videos?

Deploy a Sora model and treat video generation as an asynchronous job: create it, poll its status by ID, and download the content only when it has completed. Rendering a clip takes far longer than an HTTP request should stay open, so the create call returns a job, not a video.

  1. Create. POST /openai/v1/videos with the Sora deployment, the prompt and options such as size and duration. The response contains the video's id and an initial status (queued or in progress).
  2. Poll. GET /openai/v1/videos/{id} every few seconds until the status is completed (or failed).
  3. Download. GET /openai/v1/videos/{id}/content returns the MP4.

Image-to-video. Sora 2 can also start a clip from a still image: pass the picture as input_reference alongside the prompt and size, and it becomes the clip's starting frame. The rule to remember is that the reference image's resolution must match the requested video size exactly, and the size must be one of the landscape and portrait sizes the model supports, so a photo with different dimensions must be cropped or resized to the target size first. In the Python SDK the same create step is client.videos.create() (with input_reference as an argument), followed by the same poll and download steps. Remixing is a different feature: it modifies a clip you already generated and takes that completed video's ID, not an image.

The principle: the create call only ever returns a job. Every create starts a new job, and the MP4 is fetched separately once that job reports completed.

Responsible AI safeguards apply to every generated image and video. Foundry runs prompts (and outputs) through its content filtering system. When a prompt is judged harmful, for example a request for hateful or sexual imagery, the call fails with a content-filter error and no image is returned, while ordinary prompts keep working. That pattern, failing only for certain prompts, points to the safety system rather than to quota, an unsupported size or a wrong deployment, which would fail regardless of the prompt's content. The principles behind these safeguards are covered in the Task 1.1 lesson.

Tightening the filter is a per-deployment setting. The content filter scores categories such as violence, hate, sexual and self-harm content by severity, and blocks anything at or above each category's threshold. If the default lets milder content through that your audience should not see, create a custom content filter with a stricter (lower) threshold for that category and assign it to the image or video deployment. Don't confuse this with Prompt Shields: they detect jailbreak and prompt-injection attempts, meaning users trying to override the model's instructions, and they do not change the severity level at which a category is blocked.

Tip. This task tests building, not just naming: short practical scenarios and brief Python snippets where you pick the model, playground, message part, image source, parameter or response field that makes a vision or image-generation app work. Expect near-identical options drawn from the same API (input_image versus input_text, detail low versus high, images.generate versus images.edit, data[0].b64_json versus a URL), constraints that decide the answer (the storage must stay private, higher cost is acceptable, the app must not rebuild history), and some items that ask you to select two settings that together meet a requirement.

Key takeaways
  • Image in, words out needs a multimodal (vision-enabled) chat model; words in, image or video out needs a GPT-image, FLUX or Sora model.
  • An 'image input not supported' error means a text-only deployment: deploy a vision-capable model.
  • Test image understanding in the model playground chat; test image creation in the image playground.
  • Send text and image together as input_text and input_image parts of one user message; compare images by putting all of them in one request.
  • Image URLs must be publicly reachable; for private or local images send a base64 data URL (data:image/<type>;base64,...) or a Files API file_id.
  • detail low cuts image tokens for coarse tasks; detail high is needed for small text and fine features; a smaller vision model is the other cost lever.
  • Animated GIFs are read as their first frame only; send the frames you need as separate images.
  • GPT-image models return base64 in data[i].b64_json (no URLs); decode it before saving. images.edit uses real input images and an optional same-size PNG mask whose transparent pixels are edited.
  • The Responses image_generation tool returns an image_generation_call item and supports multi-turn refinement with previous_response_id; Sora video is create, poll, then download; for image-to-video, pass the still as input_reference at exactly the requested video size.
  • A prompt that fails with an error and no image while others work has usually been blocked by content filtering.

Frequently asked questions

Can any chat model in Microsoft Foundry read images?

No. Only multimodal, vision-capable chat models accept image input. Sending an input_image part to a text-only deployment returns an error saying image input is not supported; the fix is to deploy a vision-capable model and send the requests to that deployment.

Why does a vision model fail to read an image stored in private Azure Blob Storage?

The model service fetches image URLs itself, so a URL behind a private endpoint, inside a virtual network or on localhost cannot be retrieved. Keep the storage private by having the app read the image bytes and send them as a base64 data URL in the input_image part, or upload the image with the Files API and pass its file_id.

What does the detail setting on an image input do?

The detail setting on an input_image part controls how closely a vision-capable model examines the image. low uses a downscaled image and a small fixed number of tokens, which suits coarse classification; high examines finer detail at higher token cost, which is needed for small text and tiny features; auto, the default, lets the model choose.

How do I save an image generated by a GPT-image model in Foundry?

GPT-image models in Microsoft Foundry return each generated image as a base64 string in data[i].b64_json rather than as a URL. Decode it with base64.b64decode and write the bytes to a file whose extension matches the output_format you requested, png or jpeg.

What is the difference between images.generate and images.edit?

images.generate creates a new picture from a text prompt alone. images.edit starts from one or more existing images plus a prompt, so the real content of those images, such as a brand mascot, can appear in the result, and it accepts an optional mask: a PNG the same size as the image whose transparent pixels mark the area allowed to change.

When should I use the Responses API image_generation tool instead of the images API?

Use the image_generation tool for conversational, multi-turn image creation, where each request builds on the previous image. Chaining requests with previous_response_id lets the service carry the conversation, and the image arrives as an image_generation_call item in response.output with the base64 image in its result. The images API suits single, stateless generate or edit calls.

Why can't I download a Sora video straight after creating it?

Sora video generation in Microsoft Foundry is asynchronous. Creating a video returns a job ID and a queued or in-progress status, not the video. The app polls the video's status by that ID until it is completed and then downloads the MP4 from the video's content endpoint.

Why does an image generation request return an error and no image for only some prompts?

Microsoft Foundry applies content filtering to image and video generation. A prompt judged harmful, such as one asking for hateful or sexual imagery, is blocked with a content-filter error and no image is produced, while other prompts keep working. Errors caused by quota, unsupported sizes or the wrong deployment would occur regardless of the prompt's content.

Source

This lesson covers the "Implement AI solutions by using Microsoft Foundry" domain of the official AI-901 exam guide. Vendors revise their guides — check the source for the current version.

Sign up free to mark lessons complete, bookmark topics and track your exam readiness.

Spotted a mistake in this lesson?