Information Extraction with Azure Content Understanding in Foundry (AI-901)
Information extraction in Microsoft Foundry means using Azure Content Understanding in Foundry Tools to turn documents, forms, images, audio and video into markdown and named fields returned as predictable JSON. This AI-901 topic covers what an analyzer and a field schema are, choosing between prebuilt read, layout, search and domain analyzers, building custom analyzers from the four base analyzers, the extract, classify and generate field methods, confidence and grounding, splitting and routing mixed files, image, audio and video results, the model deployments the service relies on, and a lightweight Python app that runs an analysis and reads its result.
On this page10 sections
- What is Azure Content Understanding and what is an analyzer?
- Which prebuilt analyzer should you use?
- Which analyzers work without any model deployments?
- How does Content Understanding use your model deployments?
- How do you build a custom analyzer?
- What do the extract, classify and generate field methods do?
- How do you check confidence and split mixed files?
- How does Content Understanding extract information from images?
- How does Content Understanding extract information from audio and video?
- How do you build a lightweight extraction app in Python?
- Explain what a Content Understanding analyzer and field schema are, and why a schema matches values by meaning rather than by label.
- Choose between the content extraction, RAG, domain-specific (including composed), utility and base analyzers for a document, image, audio or video requirement.
- Define a custom analyzer with the right baseAnalyzerId and fields using the extract, classify and generate methods.
- Use confidence scores, grounding, segmentation and content categories to route values to review and split mixed files.
- Configure default and per-request model deployments, and know which analyzers need none.
- Build a lightweight Python app with ContentUnderstandingClient, begin_analyze and the result's markdown and fields, and explain the asynchronous REST pattern.
What is Azure Content Understanding and what is an analyzer?
Azure Content Understanding in Foundry Tools is the Microsoft Foundry service that turns documents, images, audio and video into structured output: markdown text plus named fields returned as predictable JSON. Every request names an analyzer, which is a reusable unit of configuration: what kind of content it expects, which processing options are switched on, which models it uses, and (optionally) a field schema listing the values to return.
Because the analyzer is fixed, every file sent to it is processed with the same settings and comes back in the same JSON shape. That is the key difference from simply prompting a chat model with a file: a prompt can return a different structure each time and is not a reusable extraction configuration.
A schema is applied semantically. Each field has a name and a description, and the service finds the value that matches the meaning of the field rather than a fixed label or position. One PolicyNumber field described as "the insurance policy reference" therefore finds the value whether a form labels it "Policy ref", "Pol. no." or prints it unlabelled beside the insurer's logo: a single well-described field covers every layout and label variant.
Analyzers come in two kinds:
- Prebuilt analyzers: ready to call by ID, for common content (invoices, receipts, call recordings, search-ready markdown).
- Custom analyzers: you define them, derived from a base analyzer, when you need your own list of fields (for example the warranty codes and technician IDs on your repair tickets).
Which prebuilt analyzer should you use?
Pick a prebuilt analyzer by what you need back: plain content, search-ready content with summaries, or a ready-made set of fields for a known document type. Prebuilt analyzers fall into five groups.
| Group | Examples | What you get | Needs model deployments? |
|---|---|---|---|
| Content extraction | prebuilt-read, prebuilt-layout, prebuilt-digitalParse | OCR text (read); text plus tables, sections, headings and figures (layout); text parsed straight from digital files without OCR (digitalParse) | No |
| RAG (search) analyzers | prebuilt-documentSearch, prebuilt-imageSearch, prebuilt-audioSearch, prebuilt-videoSearch | Markdown plus a generated summary or description, optimised for search indexes and retrieval-augmented generation | Yes |
| Domain-specific | prebuilt-invoice, prebuilt-receipt, prebuilt-idDocument, prebuilt-contract, prebuilt-callCenter and others | A ready-made field schema for a known document or recording type | Yes |
| Utility | prebuilt-documentFields, prebuilt-documentFieldSchema | Generic key-value pairs found in a document (documentFields); a proposed field schema for a new document type (documentFieldSchema) | Yes |
| Base analyzers | prebuilt-document, prebuilt-image, prebuilt-audio, prebuilt-video | The parents you derive custom analyzers from | Yes, for fields, segmentation or classification, and figure descriptions |
- Known document type, no schema design wanted: use the domain analyzer. Supplier invoices go to
prebuilt-invoice(vendor, invoice number, line items, amount due); sales receipts go toprebuilt-receipt, its near-twin.prebuilt-invoicecovers more than invoices: utility bills, sales orders and purchase orders use the same schema (dedicated analyzers such asprebuilt-utilityBillandprebuilt-purchaseOrderalso exist when you want one tuned to that type), while receipts from shops and restaurants belong toprebuilt-receipt. - An identity document: use
prebuilt-idDocument. It covers passports, driver licences, national ID cards and residency permits from many countries and returns identity fields such as first and last name, date of birth, document number and expiry date, so a check-in or onboarding app can read an ID without a custom schema. Legal agreements go toprebuilt-contract(parties, dates, terms). - A procurement document of unknown type: use
prebuilt-procurement, a composed prebuilt analyzer. It classifies each document (invoice, receipt, purchase order and so on) and routes it to the matching procurement analyzer for extraction, so you get tuned named fields without building or maintaining a classifier of your own. Other composed prebuilts work the same way for their domain, such as US tax forms and US mortgage documents. - Utility analyzers are not domain schemas:
prebuilt-documentFieldsreturns whatever key-value pairs it finds, with generic keys rather than a tuned, named schema (domain analyzers fall back to it when a document matches none of their schemas), andprebuilt-documentFieldSchemaanalyses sample documents and proposes a field schema, which helps you start designing a custom analyzer. When a domain analyzer exists for the document type, it returns more reliable, consistently named fields than either utility. - Content for a search index or chat assistant: use the matching
…Searchanalyzer.prebuilt-documentSearchkeeps headings and tables in markdown, describes figures such as diagrams and adds a summary of the whole document. - Just the text, or text plus table structure, at the lowest cost: use
prebuilt-readorprebuilt-layout.
Which analyzers work without any model deployments?
The content extraction analyzers, prebuilt-read, prebuilt-layout and prebuilt-digitalParse, run without a generative (chat completion) or embedding model deployment. They perform OCR, layout analysis or digital-file parsing only, so they consume no generative tokens and work on a brand-new Foundry resource before you have deployed anything.
prebuilt-readreturns the text (printed and handwritten) with no structure beyond lines and words.prebuilt-layoutadds structure: tables with their rows and columns, sections, headings and figures. Choose it when a requirement mentions tables but forbids generative cost.prebuilt-digitalParseparses born-digital files (for example a PDF exported from a word processor) directly, without running OCR.
Every analyzer that produces something the file does not literally contain, such as named fields, a summary, a figure description, sentiment or topics, relies on a connected chat completion deployment (and usually an embedding deployment). That covers the RAG analyzers, the domain analyzers such as prebuilt-invoice and prebuilt-callCenter, and custom analyzers with fields, segmentation or classification.
How does Content Understanding use your model deployments?
Content Understanding runs its generative steps on your own chat completion and embedding deployments in the Foundry resource, so you deploy those models first and then tell the service which deployment serves each model. There are two places to do that:
- Resource-level defaults: map each model name to a deployment once, in the Foundry portal or Content Understanding Studio (its resource settings can also deploy the models for you), or with a
PATCHto/contentunderstanding/defaults. Every analyzer on the resource then resolves its models through that mapping. Prebuilt analyzers refer to their models through aliases such asprebuilt-analyzer-completionandprebuilt-analyzer-embedding, which are resolved through the samemodelDeploymentsmapping. So if the deployments exist but no mapping does, an analyzer such asprebuilt-invoicefails because its required models cannot be resolved to a deployment. - Per request: pass a
modelDeploymentsobject in an analyze request to override the defaults for that call only, for example a one-off evaluation that tries a newer chat deployment while production traffic keeps the default.
An analyzer definition also has a models property (for example models.completion), but it holds catalog model names, not deployment names. The service maps those names to deployments at run time using the defaults or the request's modelDeployments.
How do you build a custom analyzer?
You build a custom analyzer by deriving it from one of exactly four base analyzers, named in its baseAnalyzerId property, and adding a fieldSchema with your fields. Only these four can be parents; RAG, layout and domain analyzers cannot. That is different from copying: you can fetch any prebuilt analyzer's definition (GET /analyzers/{id}) and use it as a template for a new custom analyzer, but the copy still declares one of the four as its baseAnalyzerId.
baseAnalyzerId | Use it for |
|---|---|
prebuilt-document | Documents and forms, including scans and phone photos of paper (PDF, images of text, handwriting) |
prebuilt-image | Images without meaningful text: product photos, scenes, charts treated as pictures |
prebuilt-audio | Audio recordings: calls, voicemails, meetings |
prebuilt-video | Video, where fields can depend on both what is shown and what is said |
Choose the base by the content, not the file extension. A phone photo of a handwritten inspection checklist is a document that happens to be stored as an image, so it derives from prebuilt-document, which reads printed and handwritten text. A field that depends on what is both shown and said, such as "the dishes cooked and named in a recipe video", derives from prebuilt-video, which combines frames with the transcript.
A minimal definition looks like this:
{
"baseAnalyzerId": "prebuilt-document",
"config": { "estimateFieldSourceAndConfidence": true },
"fieldSchema": {
"fields": {
"InvoiceNumber": { "type": "string", "method": "extract",
"description": "The supplier's invoice number" },
"DocumentType": { "type": "string", "method": "classify",
"enum": ["Invoice", "Receipt", "Contract"] }
}
}
}What do the extract, classify and generate field methods do?
extract returns a value verbatim (documents only), classify picks one value from a fixed list in enum, and generate writes free-form text from the content (any content type). The method is one of three properties on every field in a fieldSchema, alongside its type and description.
| Method | Produces | Supported for | Example |
|---|---|---|---|
extract | A value exactly as it appears in the content | Documents only | Policy number on a claim form, a tax ID |
classify | One value from a fixed list given in enum | All content types | Region: North, South or West; Priority: P1, P2 or P3 |
generate | Free-form text the model writes from the content | All content types | Summary, headline, an order number a caller reads out |
- A value restricted to a fixed list is a classification: use
classifywith the categories inenum. Theenumlist is what constrains the output. extractworks only for document content. In audio, image or video analyzers, even a literal value such as a booking reference spoken in a recording is produced withgenerate.extractalso needs source and confidence estimation for that field: either setestimateFieldSourceAndConfidence: truein the analyzer config, as in the example above, or setestimateSourceAndConfidence: trueon the field itself (the field-level setting overrides the analyzer-level one).
How do you check confidence and split mixed files?
Confidence scores and grounding tell an app how far to trust each extracted value. When estimateFieldSourceAndConfidence is set to true in a document analyzer's config, each extracted field value carries a confidence score between 0 and 1 and its source: the page number and bounding region on that page where the value was found, which a review tool can use to highlight it. An app checks each field's confidence: high values flow straight into the business system, low values go to a human reviewer, who uses the source location to verify them.
Three document settings are easy to confuse because all of them change what comes back:
| Setting | What it adds or removes |
|---|---|
estimateFieldSourceAndConfidence | Per field value: a 0 to 1 confidence score and its source (page and bounding region) |
enableLayout | Layout of the content: positions of paragraphs, lines and words, tables and sections; no per-field confidence |
omitContent | Drops the content object (markdown and layout) from the response, leaving only the fields |
A single file often contains several documents, such as a mortgage pack that bundles a contract, several bank statements and an ID scan in one PDF. Two settings handle this:
enableSegment: truesplits the content into segments, one per logical document.contentCategoriesdescribes each category (for example Contract and BankStatement), so every segment is classified, and each category can name ananalyzerId(such asprebuilt-contractor a custom statement analyzer) to route its segments to.
For procurement documents the routing step already exists as a prebuilt: prebuilt-procurement classifies each document and hands it to the right procurement analyzer. Build your own contentCategories analyzer when your categories or target analyzers are not covered by a composed prebuilt, or when one file holds several documents that must be split first.
By contrast, segmentPerPage splits strictly by page, which suits files where every page is its own document but cuts apart any document that runs over several pages.
How does Content Understanding extract information from images?
For images without text, Content Understanding describes or analyses what is shown; for images that contain text, it treats them as documents. The choice is:
- A ready-made description of each photo, for example to put in a search index:
prebuilt-imageSearch, which returns a one-paragraph description with no field design. - Your own fields from photos (for example the vehicle type or whether a dent is visible in an insurance photo): a custom analyzer derived from
prebuilt-image, usingclassifyorgeneratefields.prebuilt-imageis a base, not a finished analyzer. - Photos or scans of documents, including handwriting:
prebuilt-documentSearch, a domain analyzer, or a custom analyzer derived fromprebuilt-document. These read the printed and handwritten text and can extract fields from it.
Analysing images for captions, tags or objects with vision models, and generating images, are covered in the computer vision lesson for this exam.
How does Content Understanding extract information from audio and video?
Audio analyzers transcribe speech and add generated fields; video analyzers combine sampled frames with the transcript and return results per segment.
| Need | Analyzer | Result |
|---|---|---|
| Transcript and summary of any recording | prebuilt-audioSearch | WebVTT transcript with speakers labelled generically (Speaker 1, Speaker 2) plus a summary |
| Contact-centre calls analysed out of the box | prebuilt-callCenter | Transcript with speaker roles (Agent, Customer), summary, sentiment and topics |
| Your own fields from audio | Custom analyzer on prebuilt-audio | Transcript plus classify/generate fields |
| Searchable descriptions of video | prebuilt-videoSearch | Segments with descriptions, transcripts and key frames |
| Your own fields from video | Custom analyzer on prebuilt-video | Fields generated per segment |
- Timing: audio content includes
transcriptPhrases; each phrase has its ownstartTimeMsandendTimeMs, which lets an app line text up with the audio, for example to highlight each line of a subtitle track as it is spoken. The prebuilt audio analyzers turn onreturnDetails, which is why they return these phrases; a custom analyzer based onprebuilt-audioneedsreturnDetails: truein its config to get them. ThestartTimeMsandendTimeMson the audio content itself only bound the whole recording (or segment), so they cannot locate an individual sentence.cameraShotTimesMs, the list of shot boundaries, is a video-only detail and never appears in audio results. - Video segmentation: by default
prebuilt-videoSearchcuts video into scene-based segments. In a custom video analyzer,enableSegment: falsetreats the whole video as one segment;enableSegment: truewith acontentCategoriesdescription (for example "one product reviewed in a vlog, excluding the sponsor segments") creates custom segments, and fields are then generated for each one. A video analyzer supports one content category. - Frame sampling: video is sampled at roughly one frame per second and frames are resized to about 512 x 512 pixels, so fine detail and anything on screen for less than a second can be lost from the visual input, while the audio track is transcribed separately.
How do you build a lightweight extraction app in Python?
Try the analyzer in the Foundry portal first, then call it from code with the azure-ai-contentunderstanding package. In the new Foundry portal, Content Understanding lets you choose a prebuilt or custom analyzer, run it on a sample or uploaded file and inspect the fields and the JSON result, with no code, which is the quickest way to check that an analyzer returns what you expect.
from azure.ai.contentunderstanding import ContentUnderstandingClient
from azure.ai.contentunderstanding.models import AnalysisInput
from azure.core.credentials import AzureKeyCredential
client = ContentUnderstandingClient(
endpoint="https://<resource>.services.ai.azure.com/",
credential=AzureKeyCredential(key))
poller = client.begin_analyze(
analyzer_id="prebuilt-receipt",
inputs=[AnalysisInput(url=receipt_url)])
result = poller.result()
for content in result.contents:
print(content.markdown) # structured text: headings, tables
print(content.fields) # named values (with confidence when estimation is on)
- The client takes only the Foundry resource endpoint and a credential: an
AzureKeyCredentialfor key authentication or a Microsoft Entra ID token credential such asDefaultAzureCredential()(the identity then needs the Cognitive Services User role on the resource). The analyzer ID and input files are per-request arguments ofbegin_analyze, and model deployments come from the resource defaults or a request'smodelDeployments. - The result:
result.contentsholds one item per content (or segment). Readmarkdownto chunk text for a vector index while keeping headings and tables; readfieldsfor schema values such as a generated summary or an invoice total.
Analysis is a long-running operation. In the SDK, begin_analyze returns a poller and result() waits for the finished analysis. Over REST, a POST to /contentunderstanding/analyzers/{analyzerId}:analyze returns at once with no fields and an Operation-Location header; the app sends GET requests to that URL until the status is Succeeded, then reads the result.
Tip. This task tests implementation in Microsoft Foundry rather than theory: short practical scenarios in which you choose the analyzer ID, base analyzer, field method, configuration property, model deployment setting, SDK method or result property that meets a stated requirement for documents, forms, images, audio or video. It also checks that you can read a Content Understanding result and know how a lightweight app calls the service, including the asynchronous analyze pattern. Some items ask you to select two settings that together meet a requirement, or to complete a short Python snippet.
- An analyzer is a reusable configuration (content type, settings, models, optional field schema) that returns the same JSON shape every time; fields are matched by meaning, so one well-described field covers every label variant.
- prebuilt-read, prebuilt-layout and prebuilt-digitalParse need no model deployments or generative tokens; layout adds tables, sections and figures.
- RAG analyzers (documentSearch, imageSearch, audioSearch, videoSearch) return markdown plus summaries or descriptions for search; domain analyzers (invoice, receipt, idDocument for passports and driver licences, contract, callCenter) return ready-made fields, and prebuilt-invoice also covers utility bills, sales orders and purchase orders.
- prebuilt-procurement is a composed analyzer that classifies a procurement document and routes it to the matching analyzer; the utility analyzers documentFields (generic key-value pairs) and documentFieldSchema (a proposed schema) are not tuned domain schemas.
- Custom analyzers derive from one of four bases via baseAnalyzerId: prebuilt-document, prebuilt-image, prebuilt-audio or prebuilt-video; text in an image file is document content.
- Field methods: extract (verbatim, documents only), classify (fixed list in enum), generate (free-form, any content type).
- estimateFieldSourceAndConfidence adds, per field value, a confidence score (0 to 1) and its source (page and bounding region); enableLayout gives text positions without confidence, and omitContent drops the content object.
- enableSegment plus contentCategories (each with an analyzerId) splits a mixed file and routes each part; segmentPerPage splits by page only.
- Set default model deployments once per resource and override per request with modelDeployments; an analyzer's models property holds model names, not deployment names.
- ContentUnderstandingClient takes an endpoint and credential; begin_analyze returns a poller, and over REST you poll Operation-Location until Succeeded.
Frequently asked questions
What is the difference between Azure Content Understanding and prompting a chat model with a file?
Content Understanding runs a fixed analyzer, with a defined field schema and settings, so every file returns the same JSON structure with confidence scores and source locations. Prompting a chat model with a file can work, but its output structure can vary from call to call and it is not a reusable extraction configuration.
Which Content Understanding analyzers work without deploying any models?
The content extraction analyzers: prebuilt-read, prebuilt-layout and prebuilt-digitalParse, which parses digital files without OCR. They perform OCR, layout analysis or parsing without a chat completion or embedding model, so they consume no generative tokens. Analyzers that produce fields, summaries, descriptions, sentiment or topics need connected model deployments.
Should a photo of a handwritten form use prebuilt-image or prebuilt-document?
prebuilt-document. Content Understanding chooses the analyzer by content, not file type: a photo or scan that contains text is a document, and the document analyzers read printed and handwritten text. prebuilt-image is meant for images without meaningful text.
Can a Content Understanding audio or video analyzer use the extract method?
No. The extract method, which returns a value verbatim, is supported for document content only. Audio, image and video analyzers use classify for values from a fixed list and generate for everything else, including literal values such as a reference number spoken in a recording.
How do I tell Content Understanding which model deployments to use?
Map model names to your chat completion and embedding deployments as resource-level defaults, in the Foundry portal, in Content Understanding Studio or with a PATCH to the defaults endpoint, so every analyzer uses them. To use a different deployment for one call, pass a modelDeployments object in that analyze request.
Why does a Content Understanding analyze call return no fields at first?
Analysis is asynchronous. The REST POST returns immediately with an Operation-Location header, and the app polls that URL until the status is Succeeded before reading the result. In the Python SDK, begin_analyze returns a poller and its result() method waits for the analysis to finish.
What is the difference between prebuilt-audioSearch and prebuilt-callCenter?
Both transcribe and summarize recordings. prebuilt-audioSearch labels speakers generically (Speaker 1, Speaker 2), while prebuilt-callCenter identifies Agent and Customer roles and also returns sentiment and the main topics of the call.
What does prebuilt-procurement do in Content Understanding?
prebuilt-procurement is a composed prebuilt analyzer for procurement documents. It classifies each document, for example as an invoice, receipt or purchase order, and routes it to the matching procurement analyzer, so the result contains that document type's tuned fields without you building a custom classifier. prebuilt-invoice itself covers invoices, utility bills, sales orders and purchase orders, while sales receipts use prebuilt-receipt.
Source
This lesson covers the "Implement AI solutions by using Microsoft Foundry" domain of the official AI-901 exam guide. Vendors revise their guides — check the source for the current version.
- Microsoft AI-901 study guide — Microsoft Learn
Sign up free to mark lessons complete, bookmark topics and track your exam readiness.