SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
Identify AI concepts and capabilities

AI Workloads: Text Analysis, Speech, Vision and Information Extraction (AI-901)

18 min readAI-901 · Identify AI concepts and capabilitiesUpdated

AI workloads are the kinds of job AI systems do — generating content, acting as an agent, analysing text, recognising and synthesizing speech, interpreting and generating images, and extracting structured information from documents, audio and video. The AI-901 exam's "Identify AI workloads" task tests whether you can match a business scenario to the right workload and technique: key phrase extraction versus entity detection, sentiment versus opinion mining, extractive versus abstractive summaries, custom speech versus custom voice, object detection versus segmentation, OCR versus field extraction. This lesson walks through each workload by what goes in and what comes out, which is how you tell near-twins apart.

What you’ll learn
  • Match a business scenario to the AI workload it needs — generative AI, agentic AI, text analysis, speech, computer vision or information extraction — including scenarios that combine two.
  • Explain what an AI agent adds to a generative AI app: tools, actions in other systems and multi-step work toward a goal.
  • Describe key phrase extraction, entity detection (NER, PII, entity linking), sentiment analysis with opinion mining, and extractive versus abstractive summarization.
  • Identify speech recognition features (real-time and batch transcription, custom speech, diarization, language identification, translation) and speech synthesis features (neural voices, SSML, custom voice).
  • Choose between image classification, object detection, semantic segmentation, OCR, captioning and tagging, and between multimodal image understanding and image generation or editing.
  • Explain how information is extracted from documents, images, audio and video: OCR versus field extraction, extract/classify/generate field methods, confidence scores and grounding.

What are the common AI workloads on the AI-901 exam?

An AI workload is the kind of job an AI system does, named by what goes in and what comes out. AI-901 groups them into six families: generative AI, agentic AI, text analysis, speech, computer vision and information extraction. Most scenario questions are solved by naming the input (text, audio, an image, a document, a video) and the output the business wants (new content, an action, a label, a list, audio, structured fields).

WorkloadInputOutputTypical scenario
Generative AIA prompt (text, and with multimodal models images or audio too)New content: text, code, imagesDraft a product description, answer questions in a chat app, create an image from a description
Agentic AIA goal or requestActions completed in other systems, over several stepsReschedule a delivery through a courier's API and update the order record; open a support ticket and email the customer
Text analysisExisting textInsights about the text: key phrases, entities, sentiment, a summary, its languageTag documents by topic, track how customers feel, mask personal data before sharing
SpeechAudio or textText from audio (recognition) or audio from text (synthesis)Dictation into a notes app, an audiobook narrated from a manuscript
Computer visionImages or video framesLabels, boxes, pixel masks, text read from the image, captionsSort product photos by category, locate defects on a circuit board, read a shipping label
Information extractionDocuments, forms, images, audio, videoStructured fields, segments and descriptions ready for another systemPull the policy number and start date from insurance certificates; index a training-video library by topic

Watch the direction of speech workloads (audio to text is recognition, text to audio is synthesis) and the shape of the output for vision and text (one label, a list, boxes, per-pixel labels, a paragraph). Many real solutions combine two workloads: an in-car voice assistant needs speech recognition to hear a request and speech synthesis to answer aloud; a shelf-monitoring camera needs object detection to find each product and OCR to read its price label. When a scenario asks for two capabilities, check that each part of the requirement is covered by one of them.

How is agentic AI different from generative AI?

Generative AI creates new content in response to a prompt; agentic AI pursues a goal over several steps and uses tools to take actions in other systems. An agent is built on a generative model, so it still writes text, but it adds three things: instructions that define its job, tools it can decide to call (APIs, search, code execution, functions that create records) and the ability to plan and repeat steps until the goal is reached, with little human input.

Generative AI app (for example a chat app)AI agent
Main outputContent: an answer, a draft, an imageA completed task, plus messages about it
StepsUsually one prompt, one responseSeveral steps, choosing what to do next
Tools and actionsNone needed (it may be grounded in your data)Calls tools and changes things in other systems
ModelA pretrained modelThe same kind of pretrained model, not one trained from scratch

Two distinctions are often blurred. Generating new text is not what makes something an agent — every generative model does that. Grounding (answering only from an approved knowledge base) is not agentic either; a plain chat app can be grounded. What an agent adds is tool use and action. In Microsoft Foundry, agents are built and tested with Foundry Agent Service; building one is covered in the lesson on generative AI apps and agents in Foundry.

What do key phrase extraction and entity detection return?

Key phrase extraction returns the main concepts of a text as a list of short phrases; entity detection finds specific items in the text and tells you what each one is. Both work on text that already exists and both are available as prebuilt features (no model training) in Azure Language in Foundry Tools.

  • Key phrase extraction (the exam guide calls it keyword extraction): "battery life", "screen brightness", "return policy". The phrases come from the text itself and carry no category and no sentiment. Typical uses: tags under articles, topic lists for search, what a set of reviews talks about.
  • Named entity recognition (NER): finds entities and assigns each a predefined category such as Person, Organization, Location, DateTime, Quantity or Product. In "Fabrikam's chief executive, Maria Lopez, visited Nairobi on 12 May" it returns Fabrikam as Organization, Maria Lopez as Person, Nairobi as Location and 12 May as DateTime. Custom NER can learn your own entity types.
  • PII detection: a specialised form of entity detection for personally identifiable information — names, phone numbers, email addresses, ID numbers, addresses. It can return a redacted version of the text with those values masked, which is what you need before sharing documents.
  • Entity linking: disambiguates well-known entities and links each to a knowledge-base entry (Wikipedia), so "Mars" the planet and "Mars" the company are told apart. It returns links, not category labels, and private people are not in the knowledge base.
  • Text analytics for health: extracts medical concepts such as diagnoses, medications and dosages from clinical text, and the relationships between them (which dosage belongs to which medication).
NeedTechnique
Main topics as a list of phrasesKey phrase extraction
Words labelled with types (Person, Location, DateTime)Named entity recognition
Find and mask personal dataPII detection (with redaction)
Link a name to its encyclopedia entryEntity linking
One category for the whole documentText classification (not entity detection)

Language detection is a related prebuilt feature: it reports which language a text is written in.

Feature status: the techniques above are what the exam tests, but two of the Azure Language features that deliver them are on a retirement path. Entity Linking retires on 1 September 2028 and Key Phrase Extraction on 31 March 2029; Microsoft directs new entity-linking projects to named entity recognition or to generative models deployed in Microsoft Foundry, and new key-phrase projects to Foundry models. Existing solutions keep working until those dates, and the techniques themselves (listing a text's main topics, linking a name to the entity it refers to) remain valid ideas whichever tool delivers them.

What do sentiment analysis and opinion mining measure?

Sentiment analysis measures how the writer feels: it labels a document and each sentence as positive, neutral or negative (a document can also be mixed) and returns a confidence score for each label. It is the technique for tracking how customers feel at scale, across survey comments, support chats or social posts.

Opinion mining (aspect-based sentiment) goes one level deeper. It links each target in the text to the opinion expressed about it and gives that pair its own sentiment. In the restaurant review "The pasta was delicious, but the service was slow and the music was far too loud", document-level sentiment returns one overall label (likely mixed), while opinion mining returns pasta → delicious (positive), service → slow (negative) and music → too loud (negative).

TechniqueAnswersDoes not answer
Sentiment analysisIs this text positive, neutral or negative overall and per sentence?Which aspect was praised or criticised
Opinion miningWhat was said about each aspect, and was it positive or negative?—
Key phrase extractionWhat does the text talk about?Whether the writer liked it

What is the difference between extractive and abstractive summarization?

Summarization condenses a long text or conversation into a shorter version, and it comes in two forms. Extractive summarization selects the most important existing sentences and returns them unchanged, so every sentence appears word for word in the source. Abstractive summarization writes new sentences that paraphrase the source, producing a fluent plain-language recap.

ExtractiveAbstractive
SentencesCopied from the sourceNewly written
StrengthWording is never changed, so a lower risk of a change in meaning; easy to trace back to the sourceReadable, concise recap that joins ideas together
FitsText whose exact wording matters, such as policies or quotations that must be traceableReadable recaps, such as meeting notes, chat transcripts or a news digest

Summarization works on documents and on conversations (chat or call transcripts), where it can recap the issue and the resolution. A generative chat model asked to "summarise" is producing an abstractive summary: its wording is new. Key phrase extraction is not a summary — it returns a list of phrases, not readable sentences.

What can speech recognition do?

Speech recognition (speech to text) converts spoken audio into text. In Azure it is provided by Azure Speech in Foundry Tools, and the features you need to recognise are:

  • Real-time transcription: text appears while someone is speaking — dictation, voice commands, subtitles for a live broadcast.
  • Fast and batch transcription: transcribe recorded audio files. Fast transcription returns a result for a file quickly; batch transcription processes large volumes of stored recordings asynchronously.
  • Custom speech: adapt the recognition model with your own text and audio so it handles specialist vocabulary (product names, legal terms, part numbers) and difficult acoustic conditions (background noise, accents). It improves transcription accuracy; it has nothing to do with how synthesized speech sounds.
  • Speaker diarization: separates the speakers in the audio and labels each phrase with a speaker, so the transcript shows who said what.
  • Language identification: detects which language is being spoken by comparing the audio against a set of candidate languages you supply and returning the best match. It can run once at the start of the audio (the language is assumed not to change) or continuously, for audio in which the speaker may switch language. It names the language; it does not change it — that is translation.
  • Speech translation: recognises speech in one language and returns text (or synthesized speech) in another. It is only needed when the speaker and the listener use different languages.
  • Keyword recognition: listens for a wake word, as on a smart speaker, before processing anything else.
  • Pronunciation assessment: scores how accurately someone pronounces words against reference text, for language learning. It grades speech; it does not fix a transcript.

Speech recognition is about audio. Reading text that appears in an image is OCR (computer vision), and analysing text that already exists is text analysis.

What can speech synthesis do?

Speech synthesis (text to speech) turns text into spoken audio. It is the workload whenever the information already exists as text — a delivery status from an order system, a news article, a chatbot's reply — and someone needs to hear it.

  • Neural voices: prebuilt, natural-sounding voices in many languages, used for narration, IVR phone systems and voice assistants without recording voice actors.
  • Speech Synthesis Markup Language (SSML): an XML markup sent with the text that controls how individual words or phrases are spoken — pronunciation, speaking rate, pitch and volume, pauses and speaking style. It changes how a voice speaks, not whose voice it is.
  • Custom voice (Microsoft now calls it professional voice): trains a synthetic voice from recordings of a chosen voice talent, giving a brand a private voice that other companies cannot use. It is a limited-access feature that requires an application to Microsoft and the voice talent's consent.
RequirementFeatureWorkload direction
Transcripts get specialist terms wrong, or audio is noisyCustom speechAudio → text (recognition)
Control how particular words or phrases are spoken (pronunciation, rate, pitch, volume, pauses)SSMLText → audio (synthesis)
A synthetic voice that belongs to one organisationCustom voiceText → audio (synthesis)
Natural narration from scripts, no special voice neededPrebuilt neural voiceText → audio (synthesis)

Each feature controls one thing: choosing a prebuilt voice sets which standard voice speaks, SSML sets how the words are spoken, and custom voice sets whose voice it is. The near-twin names matter: custom speech improves recognition, custom voice creates a synthetic voice.

Which computer vision task fits which requirement?

Computer vision tasks differ in what they return about an image: one label, a box per object, a label per pixel, the text in it, or a description. Pick the task whose output answers the business question.

TaskReturnsUse it toCannot
Image classificationOne label (or a few) for the whole imageSort photos: "ripe" or "unripe", "cat" or "dog"Count objects or say where they are
Object detectionEach object's class plus a bounding boxCount items and locate each oneMeasure exact irregular areas (boxes include background pixels)
Semantic segmentationA class for every pixelMeasure the area of irregular regions (flooded land, road surface), find exact outlinesSeparate touching objects of the same class or read text
Optical character recognition (OCR)Printed or handwritten text found in the imageRead road signs, product labels, scanned pagesDetect the objects themselves
Image captioningA sentence describing the sceneAlt text, searchable descriptionsGive exact counts or positions
Image taggingA list of things present ("tree", "beach", "person")Index and search an image librarySay which pixels or regions each tag covers

Requirements that combine questions need combined tasks: "how many products are on this shelf and what does each price label say" is object detection (count and locate the products) plus OCR (read the labels). A rule of thumb: what = classification, what and where = detection, which pixels = segmentation, what does it say = OCR.

How do multimodal models and image-generation models differ?

A multimodal model understands images: it accepts an image and text in the same prompt and replies in natural language. An image-generation model creates images: it takes a text prompt (and optionally an existing image) and returns a new or edited picture. Both are generative models you deploy in Microsoft Foundry, but their outputs point in opposite directions.

Multimodal (vision-capable) models, such as the GPT-4o and later GPT model families, can describe a photo, answer questions about it (visual question answering: "what does this warning light on my dashboard mean?"), compare two images or read and reason about a chart. A text-only language model never receives the photo, so it can only give generic advice. How to send an image in a prompt is covered in domain 2.

Image-generation models (in Foundry, the GPT-image series; DALL-E 3 has been retired) support two operations:

  • Generate: create a brand-new image from a text description. Every pixel is new, so nothing from an existing photo is carried over.
  • Edit: send an existing image, a prompt and optionally a mask. The fully transparent pixels of the mask mark the area to edit; the model regenerates that area to match the prompt and aims to leave the rest of the image unchanged. This is inpainting, used to remove an object or change one region of a picture.
NeedModel and operation
Answer a question about an uploaded photoMultimodal model with image input
Create a new illustration from a descriptionImage generation
Change one region of an existing image and leave the rest as unchanged as possibleImage edit with a mask

How do you extract information from documents, images, audio and video?

Information extraction turns unstructured content into structured data another system can use, and for documents the key difference is between OCR, which returns all the text, and field extraction, which returns named values. OCR on a delivery note gives you every word on the page but does not say which value is the order number; field extraction returns labelled key-value pairs such as OrderNumber, DeliveryDate and RecipientName that can be written straight into an order system.

In Microsoft Foundry this is done with Azure Content Understanding in Foundry Tools (one service for documents, images, audio and video, using analyzers that you configure with a schema of fields) and Azure Document Intelligence in Foundry Tools (prebuilt and custom models for forms such as invoices and receipts). Useful outputs to know:

  • Content and layout: the text, plus the structure of each page (paragraphs, tables, selection marks), and a Markdown rendering of the whole document — useful for search and retrieval-augmented generation.
  • Fields: the values defined in your schema.
  • Confidence scores: a value from 0 to 1 per extracted field. Workflows use a threshold: high-confidence values flow straight through (straight-through processing), low-confidence ones go to a person.
  • Grounding: the location in the source where each value was found, so a reviewer can check it against the original quickly.
  • Classification: identifying the document type (invoice, claim form, ID) so it can be routed to the right analyzer.

Extract, classify and generate field methods

Each field in a Content Understanding schema is produced by one of three methods:

MethodWhat it doesExample fieldLimits
ExtractReturns a value exactly as it appears in the contentOrder number, policy start dateDocuments only; cannot produce a value that is not written in the source
ClassifyPicks one value from a predefined list of categoriesPriority: High, Medium or Low; document typeOnly returns the listed categories
GenerateProduces a free-form value with a generative modelA short meeting recap, a product descriptionWording is new, so it is not a verbatim quote

Text extraction techniques from earlier in this lesson (entities, PII, key phrases) are also information extraction from text; the difference is that a field schema returns exactly the values your process needs.

Audio and video

Extraction from audio and video starts with a transcript and then adds structure: segments, speakers and generated descriptions or summaries. For audio, such as recorded meetings or interviews, an analyzer transcribes the speech (with speaker labels) and then produces fields such as a recap (a Generate field), a label chosen from a fixed list (a Classify field) or the topics discussed.

For video, an analyzer such as a Content Understanding video analyzer segments the timeline into scenes or chapters and generates a description of each segment from both what is shown and what is said, along with the transcript and keyframes. That makes a video library searchable by what happens in it. The principle: a useful description of a segment draws on both the visuals and the speech, and covers the whole segment rather than one frame or one signal.

Why Extract works on documents only

An Extract field must be traceable: alongside each value, the analyzer returns where it was found (page and bounding region) and a confidence score — the field-source-and-confidence estimation (estimateFieldSourceAndConfidence) that grounds the value in the source. Only document analyzers support that grounding, so Extract is a documents-only method. Audio and video analyzers therefore use the other two methods: Classify when the answer must be one of a fixed set of categories, and Generate for any open-ended value — including a value that was spoken word for word, which Generate returns as a free-form field guided by the field's description.

Tip. This task tests recognition, not implementation: short business scenarios where you name the workload, technique or feature that meets a stated need. Expect near-twins side by side — key phrase extraction versus entity recognition versus entity linking, document sentiment versus opinion mining, extractive versus abstractive summaries, speech recognition versus synthesis, custom speech versus SSML versus custom voice, classification versus detection versus segmentation, OCR versus field extraction, image generation versus editing with a mask, and extract versus classify versus generate fields. A single deciding detail in the stem (word for word, where each object is, a fixed set of labels, audio rather than a document, a unique brand voice) usually settles it, and some questions ask you to select two capabilities that together cover a combined requirement.

Key takeaways
  • Identify a workload by its input and output: new content = generative AI; completed actions through tools = agentic AI; insights from existing text = text analysis; audio ↔ text = speech; labels, boxes, pixels or text from images = computer vision; structured fields from content = information extraction.
  • An agent is a generative model plus instructions plus tools it decides to call; generating text or grounding answers in a knowledge base does not make an app an agent.
  • Key phrases = list of topics; NER = words labelled with types; PII detection = find and mask personal data; entity linking = links to a knowledge base (the Azure Language Entity Linking and Key Phrase Extraction features are scheduled to retire in 2028 and 2029).
  • Sentiment analysis gives positive, neutral or negative with confidence scores per document and sentence; opinion mining gives sentiment per aspect.
  • Extractive summaries copy source sentences word for word; abstractive summaries write new sentences.
  • Speech recognition is audio to text (custom speech for jargon and noise, diarization for who said what); speech synthesis is text to audio (SSML for pronunciation, rate, pitch and volume; custom voice for a unique brand voice).
  • Classification = what; object detection = what and where (boxes, counts); segmentation = which pixels; OCR = what the image says.
  • A multimodal model answers questions about an image; an image-generation model creates new images, and an edit with a mask targets the area the mask marks.
  • OCR returns all the text; field extraction returns labelled values. Confidence scores drive straight-through processing; grounding lets reviewers check values against the source.
  • Field methods: Extract copies verbatim values with source location and confidence (documents only, because only document analyzers support that grounding); Classify picks from fixed categories; Generate produces open-ended values, and is the method for open-ended fields from audio and video.

Frequently asked questions

What is the difference between generative AI and agentic AI?

Generative AI creates new content, such as text or images, in response to a prompt. Agentic AI uses a generative model plus instructions and tools to pursue a goal over several steps, calling APIs and taking actions in other systems, such as updating an order or creating a support ticket, with little human input.

What is the difference between key phrase extraction and named entity recognition?

Key phrase extraction returns the main concepts of a text as a plain list of short phrases, such as "battery life" or "delivery time". Named entity recognition finds specific items in the text and assigns each a category such as Person, Organization, Location or DateTime. Topics are key phrases; named, typed things are entities.

What is opinion mining in sentiment analysis?

Opinion mining, also called aspect-based sentiment analysis, links each target mentioned in a text (for example food, service, price) to the opinion expressed about it and gives each its own positive or negative sentiment. Standard sentiment analysis returns one label per document and per sentence, which hides which aspects were praised or criticised.

What is the difference between extractive and abstractive summarization?

Extractive summarization selects the most important sentences from the source and returns them word for word, so the wording is never changed and the risk of altering the meaning is low (although choosing sentences out of context can still shift emphasis). Abstractive summarization writes new sentences that paraphrase the source, which reads more naturally but means the wording is not the original's.

What is the difference between custom speech and custom voice in Azure Speech?

Custom speech adapts speech recognition (speech to text) with your own text and audio so transcripts handle specialist vocabulary and noisy conditions more accurately. Custom voice is a text-to-speech feature that trains a unique synthetic voice from a voice talent's recordings; it is a limited-access feature. One improves listening, the other changes how the system speaks.

What is speaker diarization?

Speaker diarization is a speech recognition feature that separates the different speakers in an audio recording and labels each part of the transcript with a speaker, so the transcript shows who said what, for example interviewer versus interviewee in a recorded interview.

What is the difference between object detection and semantic segmentation?

Object detection finds each object, returns its class and a rectangular bounding box, so objects can be counted and located. Semantic segmentation assigns a class to every pixel of the image, so it can measure the exact area or outline of irregular regions, but it does not separate touching objects of the same class into individual items.

What is the difference between OCR and field extraction for documents?

OCR (optical character recognition) returns all the text on a page but does not say what any value means. Field extraction returns named values defined in a schema, such as order number, delivery date and recipient, with a confidence score and the source location for each, so they can go straight into another system or to a reviewer.

Source

This lesson covers the "Identify AI concepts and capabilities" domain of the official AI-901 exam guide. Vendors revise their guides — check the source for the current version.

Test yourself on this topic
Practice now

Sign up free to mark lessons complete, bookmark topics and track your exam readiness.

Spotted a mistake in this lesson?