AI Workloads: Text Analysis, Speech, Vision and Information Extraction (AI-901)
AI workloads are the kinds of job AI systems do — generating content, acting as an agent, analysing text, recognising and synthesizing speech, interpreting and generating images, and extracting structured information from documents, audio and video. The AI-901 exam's "Identify AI workloads" task tests whether you can match a business scenario to the right workload and technique: key phrase extraction versus entity detection, sentiment versus opinion mining, extractive versus abstractive summaries, custom speech versus custom voice, object detection versus segmentation, OCR versus field extraction. This lesson walks through each workload by what goes in and what comes out, which is how you tell near-twins apart.
On this page10 sections
- What are the common AI workloads on the AI-901 exam?
- How is agentic AI different from generative AI?
- What do key phrase extraction and entity detection return?
- What do sentiment analysis and opinion mining measure?
- What is the difference between extractive and abstractive summarization?
- What can speech recognition do?
- What can speech synthesis do?
- Which computer vision task fits which requirement?
- How do multimodal models and image-generation models differ?
- How do you extract information from documents, images, audio and video?
- Match a business scenario to the AI workload it needs — generative AI, agentic AI, text analysis, speech, computer vision or information extraction — including scenarios that combine two.
- Explain what an AI agent adds to a generative AI app: tools, actions in other systems and multi-step work toward a goal.
- Describe key phrase extraction, entity detection (NER, PII, entity linking), sentiment analysis with opinion mining, and extractive versus abstractive summarization.
- Identify speech recognition features (real-time and batch transcription, custom speech, diarization, language identification, translation) and speech synthesis features (neural voices, SSML, custom voice).
- Choose between image classification, object detection, semantic segmentation, OCR, captioning and tagging, and between multimodal image understanding and image generation or editing.
- Explain how information is extracted from documents, images, audio and video: OCR versus field extraction, extract/classify/generate field methods, confidence scores and grounding.
What are the common AI workloads on the AI-901 exam?
An AI workload is the kind of job an AI system does, named by what goes in and what comes out. AI-901 groups them into six families: generative AI, agentic AI, text analysis, speech, computer vision and information extraction. Most scenario questions are solved by naming the input (text, audio, an image, a document, a video) and the output the business wants (new content, an action, a label, a list, audio, structured fields).
| Workload | Input | Output | Typical scenario |
|---|---|---|---|
| Generative AI | A prompt (text, and with multimodal models images or audio too) | New content: text, code, images | Draft a product description, answer questions in a chat app, create an image from a description |
| Agentic AI | A goal or request | Actions completed in other systems, over several steps | Reschedule a delivery through a courier's API and update the order record; open a support ticket and email the customer |
| Text analysis | Existing text | Insights about the text: key phrases, entities, sentiment, a summary, its language | Tag documents by topic, track how customers feel, mask personal data before sharing |
| Speech | Audio or text | Text from audio (recognition) or audio from text (synthesis) | Dictation into a notes app, an audiobook narrated from a manuscript |
| Computer vision | Images or video frames | Labels, boxes, pixel masks, text read from the image, captions | Sort product photos by category, locate defects on a circuit board, read a shipping label |
| Information extraction | Documents, forms, images, audio, video | Structured fields, segments and descriptions ready for another system | Pull the policy number and start date from insurance certificates; index a training-video library by topic |
Watch the direction of speech workloads (audio to text is recognition, text to audio is synthesis) and the shape of the output for vision and text (one label, a list, boxes, per-pixel labels, a paragraph). Many real solutions combine two workloads: an in-car voice assistant needs speech recognition to hear a request and speech synthesis to answer aloud; a shelf-monitoring camera needs object detection to find each product and OCR to read its price label. When a scenario asks for two capabilities, check that each part of the requirement is covered by one of them.
How is agentic AI different from generative AI?
Generative AI creates new content in response to a prompt; agentic AI pursues a goal over several steps and uses tools to take actions in other systems. An agent is built on a generative model, so it still writes text, but it adds three things: instructions that define its job, tools it can decide to call (APIs, search, code execution, functions that create records) and the ability to plan and repeat steps until the goal is reached, with little human input.
| Generative AI app (for example a chat app) | AI agent | |
|---|---|---|
| Main output | Content: an answer, a draft, an image | A completed task, plus messages about it |
| Steps | Usually one prompt, one response | Several steps, choosing what to do next |
| Tools and actions | None needed (it may be grounded in your data) | Calls tools and changes things in other systems |
| Model | A pretrained model | The same kind of pretrained model, not one trained from scratch |
Two distinctions are often blurred. Generating new text is not what makes something an agent — every generative model does that. Grounding (answering only from an approved knowledge base) is not agentic either; a plain chat app can be grounded. What an agent adds is tool use and action. In Microsoft Foundry, agents are built and tested with Foundry Agent Service; building one is covered in the lesson on generative AI apps and agents in Foundry.
What do key phrase extraction and entity detection return?
Key phrase extraction returns the main concepts of a text as a list of short phrases; entity detection finds specific items in the text and tells you what each one is. Both work on text that already exists and both are available as prebuilt features (no model training) in Azure Language in Foundry Tools.
- Key phrase extraction (the exam guide calls it keyword extraction): "battery life", "screen brightness", "return policy". The phrases come from the text itself and carry no category and no sentiment. Typical uses: tags under articles, topic lists for search, what a set of reviews talks about.
- Named entity recognition (NER): finds entities and assigns each a predefined category such as Person, Organization, Location, DateTime, Quantity or Product. In "Fabrikam's chief executive, Maria Lopez, visited Nairobi on 12 May" it returns Fabrikam as Organization, Maria Lopez as Person, Nairobi as Location and 12 May as DateTime. Custom NER can learn your own entity types.
- PII detection: a specialised form of entity detection for personally identifiable information — names, phone numbers, email addresses, ID numbers, addresses. It can return a redacted version of the text with those values masked, which is what you need before sharing documents.
- Entity linking: disambiguates well-known entities and links each to a knowledge-base entry (Wikipedia), so "Mars" the planet and "Mars" the company are told apart. It returns links, not category labels, and private people are not in the knowledge base.
- Text analytics for health: extracts medical concepts such as diagnoses, medications and dosages from clinical text, and the relationships between them (which dosage belongs to which medication).
| Need | Technique |
|---|---|
| Main topics as a list of phrases | Key phrase extraction |
| Words labelled with types (Person, Location, DateTime) | Named entity recognition |
| Find and mask personal data | PII detection (with redaction) |
| Link a name to its encyclopedia entry | Entity linking |
| One category for the whole document | Text classification (not entity detection) |
Language detection is a related prebuilt feature: it reports which language a text is written in.
Feature status: the techniques above are what the exam tests, but two of the Azure Language features that deliver them are on a retirement path. Entity Linking retires on 1 September 2028 and Key Phrase Extraction on 31 March 2029; Microsoft directs new entity-linking projects to named entity recognition or to generative models deployed in Microsoft Foundry, and new key-phrase projects to Foundry models. Existing solutions keep working until those dates, and the techniques themselves (listing a text's main topics, linking a name to the entity it refers to) remain valid ideas whichever tool delivers them.
What do sentiment analysis and opinion mining measure?
Sentiment analysis measures how the writer feels: it labels a document and each sentence as positive, neutral or negative (a document can also be mixed) and returns a confidence score for each label. It is the technique for tracking how customers feel at scale, across survey comments, support chats or social posts.
Opinion mining (aspect-based sentiment) goes one level deeper. It links each target in the text to the opinion expressed about it and gives that pair its own sentiment. In the restaurant review "The pasta was delicious, but the service was slow and the music was far too loud", document-level sentiment returns one overall label (likely mixed), while opinion mining returns pasta → delicious (positive), service → slow (negative) and music → too loud (negative).
| Technique | Answers | Does not answer |
|---|---|---|
| Sentiment analysis | Is this text positive, neutral or negative overall and per sentence? | Which aspect was praised or criticised |
| Opinion mining | What was said about each aspect, and was it positive or negative? | — |
| Key phrase extraction | What does the text talk about? | Whether the writer liked it |
What is the difference between extractive and abstractive summarization?
Summarization condenses a long text or conversation into a shorter version, and it comes in two forms. Extractive summarization selects the most important existing sentences and returns them unchanged, so every sentence appears word for word in the source. Abstractive summarization writes new sentences that paraphrase the source, producing a fluent plain-language recap.
| Extractive | Abstractive | |
|---|---|---|
| Sentences | Copied from the source | Newly written |
| Strength | Wording is never changed, so a lower risk of a change in meaning; easy to trace back to the source | Readable, concise recap that joins ideas together |
| Fits | Text whose exact wording matters, such as policies or quotations that must be traceable | Readable recaps, such as meeting notes, chat transcripts or a news digest |
Summarization works on documents and on conversations (chat or call transcripts), where it can recap the issue and the resolution. A generative chat model asked to "summarise" is producing an abstractive summary: its wording is new. Key phrase extraction is not a summary — it returns a list of phrases, not readable sentences.
What can speech recognition do?
Speech recognition (speech to text) converts spoken audio into text. In Azure it is provided by Azure Speech in Foundry Tools, and the features you need to recognise are:
- Real-time transcription: text appears while someone is speaking — dictation, voice commands, subtitles for a live broadcast.
- Fast and batch transcription: transcribe recorded audio files. Fast transcription returns a result for a file quickly; batch transcription processes large volumes of stored recordings asynchronously.
- Custom speech: adapt the recognition model with your own text and audio so it handles specialist vocabulary (product names, legal terms, part numbers) and difficult acoustic conditions (background noise, accents). It improves transcription accuracy; it has nothing to do with how synthesized speech sounds.
- Speaker diarization: separates the speakers in the audio and labels each phrase with a speaker, so the transcript shows who said what.
- Language identification: detects which language is being spoken by comparing the audio against a set of candidate languages you supply and returning the best match. It can run once at the start of the audio (the language is assumed not to change) or continuously, for audio in which the speaker may switch language. It names the language; it does not change it — that is translation.
- Speech translation: recognises speech in one language and returns text (or synthesized speech) in another. It is only needed when the speaker and the listener use different languages.
- Keyword recognition: listens for a wake word, as on a smart speaker, before processing anything else.
- Pronunciation assessment: scores how accurately someone pronounces words against reference text, for language learning. It grades speech; it does not fix a transcript.
Speech recognition is about audio. Reading text that appears in an image is OCR (computer vision), and analysing text that already exists is text analysis.
What can speech synthesis do?
Speech synthesis (text to speech) turns text into spoken audio. It is the workload whenever the information already exists as text — a delivery status from an order system, a news article, a chatbot's reply — and someone needs to hear it.
- Neural voices: prebuilt, natural-sounding voices in many languages, used for narration, IVR phone systems and voice assistants without recording voice actors.
- Speech Synthesis Markup Language (SSML): an XML markup sent with the text that controls how individual words or phrases are spoken — pronunciation, speaking rate, pitch and volume, pauses and speaking style. It changes how a voice speaks, not whose voice it is.
- Custom voice (Microsoft now calls it professional voice): trains a synthetic voice from recordings of a chosen voice talent, giving a brand a private voice that other companies cannot use. It is a limited-access feature that requires an application to Microsoft and the voice talent's consent.
| Requirement | Feature | Workload direction |
|---|---|---|
| Transcripts get specialist terms wrong, or audio is noisy | Custom speech | Audio → text (recognition) |
| Control how particular words or phrases are spoken (pronunciation, rate, pitch, volume, pauses) | SSML | Text → audio (synthesis) |
| A synthetic voice that belongs to one organisation | Custom voice | Text → audio (synthesis) |
| Natural narration from scripts, no special voice needed | Prebuilt neural voice | Text → audio (synthesis) |
Each feature controls one thing: choosing a prebuilt voice sets which standard voice speaks, SSML sets how the words are spoken, and custom voice sets whose voice it is. The near-twin names matter: custom speech improves recognition, custom voice creates a synthetic voice.
Which computer vision task fits which requirement?
Computer vision tasks differ in what they return about an image: one label, a box per object, a label per pixel, the text in it, or a description. Pick the task whose output answers the business question.
| Task | Returns | Use it to | Cannot |
|---|---|---|---|
| Image classification | One label (or a few) for the whole image | Sort photos: "ripe" or "unripe", "cat" or "dog" | Count objects or say where they are |
| Object detection | Each object's class plus a bounding box | Count items and locate each one | Measure exact irregular areas (boxes include background pixels) |
| Semantic segmentation | A class for every pixel | Measure the area of irregular regions (flooded land, road surface), find exact outlines | Separate touching objects of the same class or read text |
| Optical character recognition (OCR) | Printed or handwritten text found in the image | Read road signs, product labels, scanned pages | Detect the objects themselves |
| Image captioning | A sentence describing the scene | Alt text, searchable descriptions | Give exact counts or positions |
| Image tagging | A list of things present ("tree", "beach", "person") | Index and search an image library | Say which pixels or regions each tag covers |
Requirements that combine questions need combined tasks: "how many products are on this shelf and what does each price label say" is object detection (count and locate the products) plus OCR (read the labels). A rule of thumb: what = classification, what and where = detection, which pixels = segmentation, what does it say = OCR.
How do multimodal models and image-generation models differ?
A multimodal model understands images: it accepts an image and text in the same prompt and replies in natural language. An image-generation model creates images: it takes a text prompt (and optionally an existing image) and returns a new or edited picture. Both are generative models you deploy in Microsoft Foundry, but their outputs point in opposite directions.
Multimodal (vision-capable) models, such as the GPT-4o and later GPT model families, can describe a photo, answer questions about it (visual question answering: "what does this warning light on my dashboard mean?"), compare two images or read and reason about a chart. A text-only language model never receives the photo, so it can only give generic advice. How to send an image in a prompt is covered in domain 2.
Image-generation models (in Foundry, the GPT-image series; DALL-E 3 has been retired) support two operations:
- Generate: create a brand-new image from a text description. Every pixel is new, so nothing from an existing photo is carried over.
- Edit: send an existing image, a prompt and optionally a mask. The fully transparent pixels of the mask mark the area to edit; the model regenerates that area to match the prompt and aims to leave the rest of the image unchanged. This is inpainting, used to remove an object or change one region of a picture.
| Need | Model and operation |
|---|---|
| Answer a question about an uploaded photo | Multimodal model with image input |
| Create a new illustration from a description | Image generation |
| Change one region of an existing image and leave the rest as unchanged as possible | Image edit with a mask |
How do you extract information from documents, images, audio and video?
Information extraction turns unstructured content into structured data another system can use, and for documents the key difference is between OCR, which returns all the text, and field extraction, which returns named values. OCR on a delivery note gives you every word on the page but does not say which value is the order number; field extraction returns labelled key-value pairs such as OrderNumber, DeliveryDate and RecipientName that can be written straight into an order system.
In Microsoft Foundry this is done with Azure Content Understanding in Foundry Tools (one service for documents, images, audio and video, using analyzers that you configure with a schema of fields) and Azure Document Intelligence in Foundry Tools (prebuilt and custom models for forms such as invoices and receipts). Useful outputs to know:
- Content and layout: the text, plus the structure of each page (paragraphs, tables, selection marks), and a Markdown rendering of the whole document — useful for search and retrieval-augmented generation.
- Fields: the values defined in your schema.
- Confidence scores: a value from 0 to 1 per extracted field. Workflows use a threshold: high-confidence values flow straight through (straight-through processing), low-confidence ones go to a person.
- Grounding: the location in the source where each value was found, so a reviewer can check it against the original quickly.
- Classification: identifying the document type (invoice, claim form, ID) so it can be routed to the right analyzer.
Extract, classify and generate field methods
Each field in a Content Understanding schema is produced by one of three methods:
| Method | What it does | Example field | Limits |
|---|---|---|---|
| Extract | Returns a value exactly as it appears in the content | Order number, policy start date | Documents only; cannot produce a value that is not written in the source |
| Classify | Picks one value from a predefined list of categories | Priority: High, Medium or Low; document type | Only returns the listed categories |
| Generate | Produces a free-form value with a generative model | A short meeting recap, a product description | Wording is new, so it is not a verbatim quote |
Text extraction techniques from earlier in this lesson (entities, PII, key phrases) are also information extraction from text; the difference is that a field schema returns exactly the values your process needs.
Audio and video
Extraction from audio and video starts with a transcript and then adds structure: segments, speakers and generated descriptions or summaries. For audio, such as recorded meetings or interviews, an analyzer transcribes the speech (with speaker labels) and then produces fields such as a recap (a Generate field), a label chosen from a fixed list (a Classify field) or the topics discussed.
For video, an analyzer such as a Content Understanding video analyzer segments the timeline into scenes or chapters and generates a description of each segment from both what is shown and what is said, along with the transcript and keyframes. That makes a video library searchable by what happens in it. The principle: a useful description of a segment draws on both the visuals and the speech, and covers the whole segment rather than one frame or one signal.
Why Extract works on documents only
An Extract field must be traceable: alongside each value, the analyzer returns where it was found (page and bounding region) and a confidence score — the field-source-and-confidence estimation (estimateFieldSourceAndConfidence) that grounds the value in the source. Only document analyzers support that grounding, so Extract is a documents-only method. Audio and video analyzers therefore use the other two methods: Classify when the answer must be one of a fixed set of categories, and Generate for any open-ended value — including a value that was spoken word for word, which Generate returns as a free-form field guided by the field's description.
Tip. This task tests recognition, not implementation: short business scenarios where you name the workload, technique or feature that meets a stated need. Expect near-twins side by side — key phrase extraction versus entity recognition versus entity linking, document sentiment versus opinion mining, extractive versus abstractive summaries, speech recognition versus synthesis, custom speech versus SSML versus custom voice, classification versus detection versus segmentation, OCR versus field extraction, image generation versus editing with a mask, and extract versus classify versus generate fields. A single deciding detail in the stem (word for word, where each object is, a fixed set of labels, audio rather than a document, a unique brand voice) usually settles it, and some questions ask you to select two capabilities that together cover a combined requirement.
- Identify a workload by its input and output: new content = generative AI; completed actions through tools = agentic AI; insights from existing text = text analysis; audio ↔ text = speech; labels, boxes, pixels or text from images = computer vision; structured fields from content = information extraction.
- An agent is a generative model plus instructions plus tools it decides to call; generating text or grounding answers in a knowledge base does not make an app an agent.
- Key phrases = list of topics; NER = words labelled with types; PII detection = find and mask personal data; entity linking = links to a knowledge base (the Azure Language Entity Linking and Key Phrase Extraction features are scheduled to retire in 2028 and 2029).
- Sentiment analysis gives positive, neutral or negative with confidence scores per document and sentence; opinion mining gives sentiment per aspect.
- Extractive summaries copy source sentences word for word; abstractive summaries write new sentences.
- Speech recognition is audio to text (custom speech for jargon and noise, diarization for who said what); speech synthesis is text to audio (SSML for pronunciation, rate, pitch and volume; custom voice for a unique brand voice).
- Classification = what; object detection = what and where (boxes, counts); segmentation = which pixels; OCR = what the image says.
- A multimodal model answers questions about an image; an image-generation model creates new images, and an edit with a mask targets the area the mask marks.
- OCR returns all the text; field extraction returns labelled values. Confidence scores drive straight-through processing; grounding lets reviewers check values against the source.
- Field methods: Extract copies verbatim values with source location and confidence (documents only, because only document analyzers support that grounding); Classify picks from fixed categories; Generate produces open-ended values, and is the method for open-ended fields from audio and video.
Frequently asked questions
What is the difference between generative AI and agentic AI?
Generative AI creates new content, such as text or images, in response to a prompt. Agentic AI uses a generative model plus instructions and tools to pursue a goal over several steps, calling APIs and taking actions in other systems, such as updating an order or creating a support ticket, with little human input.
What is the difference between key phrase extraction and named entity recognition?
Key phrase extraction returns the main concepts of a text as a plain list of short phrases, such as "battery life" or "delivery time". Named entity recognition finds specific items in the text and assigns each a category such as Person, Organization, Location or DateTime. Topics are key phrases; named, typed things are entities.
What is opinion mining in sentiment analysis?
Opinion mining, also called aspect-based sentiment analysis, links each target mentioned in a text (for example food, service, price) to the opinion expressed about it and gives each its own positive or negative sentiment. Standard sentiment analysis returns one label per document and per sentence, which hides which aspects were praised or criticised.
What is the difference between extractive and abstractive summarization?
Extractive summarization selects the most important sentences from the source and returns them word for word, so the wording is never changed and the risk of altering the meaning is low (although choosing sentences out of context can still shift emphasis). Abstractive summarization writes new sentences that paraphrase the source, which reads more naturally but means the wording is not the original's.
What is the difference between custom speech and custom voice in Azure Speech?
Custom speech adapts speech recognition (speech to text) with your own text and audio so transcripts handle specialist vocabulary and noisy conditions more accurately. Custom voice is a text-to-speech feature that trains a unique synthetic voice from a voice talent's recordings; it is a limited-access feature. One improves listening, the other changes how the system speaks.
What is speaker diarization?
Speaker diarization is a speech recognition feature that separates the different speakers in an audio recording and labels each part of the transcript with a speaker, so the transcript shows who said what, for example interviewer versus interviewee in a recorded interview.
What is the difference between object detection and semantic segmentation?
Object detection finds each object, returns its class and a rectangular bounding box, so objects can be counted and located. Semantic segmentation assigns a class to every pixel of the image, so it can measure the exact area or outline of irregular regions, but it does not separate touching objects of the same class into individual items.
What is the difference between OCR and field extraction for documents?
OCR (optical character recognition) returns all the text on a page but does not say what any value means. Field extraction returns named values defined in a schema, such as order number, delivery date and recipient, with a confidence score and the source location for each, so they can go straight into another system or to a reviewer.
Source
This lesson covers the "Identify AI concepts and capabilities" domain of the official AI-901 exam guide. Vendors revise their guides — check the source for the current version.
- Microsoft AI-901 study guide — Microsoft Learn
Sign up free to mark lessons complete, bookmark topics and track your exam readiness.