SaveMyCert
Log in
5 of 5 free questions left today·for 30 a day
Implement AI solutions by using Microsoft Foundry

Text Analysis and Speech Apps in Microsoft Foundry (AI-901)

14 min readAI-901 · Implement AI solutions by using Microsoft FoundryUpdated

Building text and speech solutions in Microsoft Foundry means calling Azure Language in Foundry Tools or a deployed model to analyze text, sending spoken prompts to an audio-capable multimodal model, and using the Azure Speech SDK to turn speech into text and text into speech from a small Python app. This AI-901 topic covers when a purpose-built analyzer beats a prompt, the TextAnalyticsClient and its results, the OpenAI client for a Foundry deployment, the Azure Language tool for agents, audio input and output on multimodal models, SpeechConfig and audio configuration, single-shot and continuous recognition, SSML synthesis, result reasons, and Voice Live for real-time voice agents.

What you’ll learn
  • Choose between a deployed general-purpose model and Azure Language in Foundry Tools for a text-analysis requirement.
  • Create a TextAnalyticsClient from a Foundry resource, call the right analyzer method and read its per-document results and errors.
  • Call a deployed model from Python with the OpenAI library and the Foundry /openai/v1/ endpoint, and add the Azure Language tool to a Foundry agent.
  • Send a spoken prompt to an audio-capable multimodal model and get back a spoken answer and its transcript.
  • Configure the Azure Speech SDK with SpeechConfig, AudioConfig and AudioOutputConfig, including recognition language and voice selection.
  • Build speech-to-text with single-shot or continuous recognition and diagnose results with ResultReason and CancellationReason.
  • Synthesize speech with plain text or SSML, and pick between the Speech SDK, a multimodal model and Voice Live.

Should you analyze text with a deployed model or with Azure Language?

Use Azure Language in Foundry Tools when you need consistent, structured values from a well-defined analysis, and use a deployed general-purpose model when you need flexible, instruction-driven work that combines several tasks. Both are available from the same Microsoft Foundry resource, so the choice is about the shape of the result, not about extra infrastructure.

Deployed general-purpose model (for example a GPT model)Azure Language in Foundry Tools
How you askA natural-language promptCall a specific analyzer method (sentiment, PII, language detection…)
What comes backGenerated text; wording and layout can vary between runsA structured result object with fixed fields (labels, codes, scores, entity lists)
StrengthCombining tasks freely in one request (extract the main points, then rewrite them in a friendly tone); no configurationRepeatable, predictable output that a pipeline, database or compliance step can store directly
Typical fitAd-hoc analysis, multi-step instructions, conversational appsHigh-volume batches, PII redaction before storage, language codes and sentiment scores written to columns

A model's output can be constrained (for example with structured outputs or few-shot examples), but a purpose-built analyzer returns those values without any prompt design. Azure Language also offers summarization (extractive and abstractive summaries) through the same client. Which analysis technique matches which business need (key phrases, entities, sentiment, summarization) is covered in the Task 1.3 lesson on AI workloads; this lesson is about building the app.

How do you create an Azure Language client and pick the right method?

Install azure-ai-textanalytics, create a TextAnalyticsClient from your Foundry resource's endpoint and a credential, then call the method that matches the analysis. A Foundry resource already includes Azure Language, so you do not create a separate Language resource.

from azure.ai.textanalytics import TextAnalyticsClient
from azure.core.credentials import AzureKeyCredential

client = TextAnalyticsClient(
    endpoint="https://<resource>.cognitiveservices.azure.com/",
    credential=AzureKeyCredential(key))
  • Endpoint: the Foundry resource endpoint (the cognitiveservices.azure.com form), not the OpenAI /openai/v1/ path.
  • Credential: for key authentication, wrap the resource key in AzureKeyCredential. For Microsoft Entra ID authentication you pass a token credential such as DefaultAzureCredential() instead, and the client obtains tokens for the signed-in identity.
  • No deployment name: Language analyzers are prebuilt and called directly by method, so the client needs only the endpoint and credential. Deployment names belong to calls to a deployed generative model.
MethodWhat it doesMain result fields
detect_language()Identifies the language of each documentprimary_language.name, .iso6391_name, .confidence_score
analyze_sentiment()Labels each document positive, negative, neutral or mixedsentiment, confidence_scores, sentences
extract_key_phrases()Lists the main talking points or topicskey_phrases
recognize_entities()Named entity recognition: people, places, organizations, dates, quantitiesentities (text, category, confidence)
recognize_linked_entities()Entity linking: matches well-known entities to a knowledge base such as Wikipediaentities (name, URL, data source)
recognize_pii_entities()Finds personal data (phone numbers, addresses, emails…) and returns a redacted copyentities, redacted_text
begin_analyze_actions()Runs several analyses (for example PII, sentiment, summarization) as one long-running operationOne result per action per document

Each method does one job. PII detection is the privacy analyzer: it both lists personal details and produces a masked copy of the document. Named entity recognition classifies the people, places and organizations mentioned; entity linking ties well-known ones to a knowledge-base entry; key phrase extraction summarizes what a document is about; sentiment analysis labels opinion. An app that needs several kinds of output calls one method per analysis (or bundles them with begin_analyze_actions()).

How do you read the results the Language SDK returns?

The Language SDK takes a list of documents and returns a list of results in the same order, one per document, so result n always belongs to input n. Each list item is analyzed as a separate document with its own result.

Errors are reported per document, not raised. If the service rejects one document (an empty string, for example), that position holds a DocumentError whose is_error is True, while the other documents succeed normally. Well-written code checks is_error before reading fields:

for doc in client.analyze_sentiment(reviews):
    if doc.is_error:
        print(doc.error.code, doc.error.message)
    else:
        print(doc.sentiment, doc.confidence_scores.positive)

Know where each value lives:

  • Language detection: primary_language.name is the English name (for example French); primary_language.iso6391_name is the two-letter ISO 639-1 code (for example fr), which is what you use as a key in code; primary_language.confidence_score is a 0–1 certainty, useful for flagging doubtful detections.
  • PII detection: redacted_text is the whole document with each detected personal detail replaced by asterisks — the copy that is safe to store. Each item in entities describes one finding, with the detected value, its type (such as PhoneNumber) and a confidence score; because entities holds the detected values as they appear in the source, it is for inspection, while redacted_text is for storage.
  • Sentiment: sentiment is the label; confidence_scores holds the positive, neutral and negative scores.
  • Warnings: warnings lists non-fatal processing notices for a document (for example, text that was truncated).

How do you call a deployed model for text analysis from Python?

Use the OpenAI Python library pointed at your Foundry resource: set base_url to https://<resource>.openai.azure.com/openai/v1/, pass the Foundry key as api_key, and name your deployment in the model argument of a Responses API call.

from openai import OpenAI

client = OpenAI(base_url="https://<resource>.openai.azure.com/openai/v1/",
                api_key=key)
response = client.responses.create(
    model="my-gpt-deployment",
    input="List the product features this customer mentions and rewrite the complaint politely: ...")
print(response.output_text)

Three endpoints are easy to confuse, and each goes with a different client:

EndpointUsed by
https://<resource>.openai.azure.com/openai/v1/The OpenAI library calling models deployed in your Foundry resource
https://<resource>.cognitiveservices.azure.com/Foundry Tools SDKs such as Azure Language and Azure Speech
https://api.openai.com/v1/OpenAI's own public service — it never reaches a model in your Foundry resource

Keep endpoints and keys in configuration (a .env file or environment variables) rather than in code. Chat clients built with the Foundry SDK (azure-ai-projects) and multi-turn conversations are taught in the Task 2.1 lesson; here the model is simply another way to analyze text.

How do you give a Foundry agent Azure Language capabilities?

Add the Azure Language in Foundry Tools MCP server to the agent as a tool. In the Foundry agents playground you search the tool catalog for Azure Language in Foundry Tools and connect it using your Foundry resource name; the agent can then call Azure Language analyzers through the Model Context Protocol (MCP) with no integration code of your own.

The MCP server exposes Azure Language's text analyzers as tools, including language detection, sentiment analysis, named entity recognition, key phrase extraction, summarization and PII detection and redaction, among others. Because it reaches Azure Language, it works on text the agent already has. Other kinds of input belong to other Foundry capabilities:

Input and jobFoundry capability that handles it
Text: language, sentiment, entities, key phrases, summaries, personal dataAzure Language (this MCP tool)
Audio: turning recorded or live speech into textAzure Speech
Documents and forms: pulling named fields out of scanned filesContent Understanding (information extraction)
Images: describing or tagging what a picture showsA vision-capable model or vision service

With the tool connected, the agent decides when to call an analyzer, the analyzer does the analysis and returns its structured result, and the model uses that result in its reply. The analysis itself (what counts as personal data, how it is masked, which sentiment label applies) is done by Azure Language, so it is the same every time the tool is called.

Keep the two addresses apart: the deployed model is the agent's reasoning engine, while the Language tool connection identifies the Foundry resource. The Foundry resource already includes Azure Language, so that one connection is all the tool needs.

How do you respond to a spoken prompt with a deployed multimodal model?

Send the recording to an audio-capable chat model (for example a gpt-4o-audio-preview or gpt-4o-mini-audio-preview deployment) as base64-encoded data in an input_audio content part of the user message, and the model answers the spoken question directly. To also get a spoken answer, add modalities=["text", "audio"] and an audio argument that names the voice and output format.

audio_b64 = base64.b64encode(open("question.wav", "rb").read()).decode()

completion = client.chat.completions.create(
    model="my-audio-deployment",
    modalities=["text", "audio"],
    audio={"voice": "alloy", "format": "wav"},
    messages=[{"role": "user", "content": [
        {"type": "text", "text": "Answer the question in this recording."},
        {"type": "input_audio",
         "input_audio": {"data": audio_b64, "format": "wav"}}]}])

reply = completion.choices[0].message
print(reply.audio.transcript)                         # the words of the spoken answer
open("answer.wav", "wb").write(base64.b64decode(reply.audio.data))
PieceDirectionWhat it does
input_audio content partInputCarries the user's recording itself as base64 data plus its format
modalities=["text", "audio"]OutputAsks the model to produce audio as well as text
Top-level audio={"voice", "format"}OutputChooses the voice and file format of the spoken reply
message.audio.dataOutputThe spoken reply as base64 audio — decode it for playback
message.audio.transcriptOutputThe words of the spoken reply as text — what you log or display

When the reply comes back as audio, its words are carried in audio.transcript rather than in message.content. Spoken output is switched on by the pair modalities + audio; without them the model answers in text only. You can also try an audio-capable deployment without code in the Foundry chat playground by recording or uploading audio.

Distinguish audio-capable chat models from transcription models such as gpt-4o-transcribe: a transcription model produces a transcript of the audio it receives. An audio-capable chat model understands the spoken prompt and generates the answer, so one call covers what would otherwise be a three-step pipeline (recognize, generate, synthesize).

How do you set up the Azure Speech SDK?

Install azure-cognitiveservices-speech, create a SpeechConfig from the Foundry resource key and endpoint, then pass it — together with an audio configuration — to a SpeechRecognizer (speech to text) or a SpeechSynthesizer (text to speech).

import azure.cognitiveservices.speech as speechsdk

speech_config = speechsdk.SpeechConfig(
    subscription=os.environ.get("SPEECH_KEY"),
    endpoint=os.environ.get("ENDPOINT"))

SpeechConfig is the only object that carries the key and endpoint. It also holds service settings for each direction, and the two sets are easy to mix up:

SettingDirectionEffect
speech_recognition_languageSpeech to textThe locale to transcribe (for example fr-FR); the default is en-US, so other languages must be set or transcripts are poor
speech_synthesis_languageText to speechThe output locale; the service uses that locale's default voice
speech_synthesis_voice_nameText to speechOne specific prebuilt neural voice (for example de-DE-KatjaNeural); takes precedence over the synthesis language

Where the audio comes from or goes to is a separate object, and the direction is in the class name:

ObjectUsed withOptions
speechsdk.audio.AudioConfigRecognizer inputuse_default_microphone=True to listen live, or filename="interview.wav" to read a stored recording
speechsdk.audio.AudioOutputConfigSynthesizer outputuse_default_speaker=True to play, or filename="greeting.wav" to save the speech to a file

The same filename= argument means opposite things in the two classes: AudioConfig reads a file, AudioOutputConfig writes one. A recognizer created with no audio configuration listens on the default microphone. set_speech_synthesis_output_format(...) on SpeechConfig chooses the encoding of synthesized audio (for example a RIFF/WAV format), while AudioOutputConfig chooses its destination.

How do you transcribe speech, and how do you know why it failed?

Call recognize_once_async() for a single utterance and start_continuous_recognition() for long or live audio; then check result.reason to learn whether speech was recognized, nothing was heard, or the request failed.

ApproachBehaviourFits
recognize_once_async() / recognize_once()Returns after one utterance (it stops at the first pause, up to a short maximum)Voice commands, short spoken questions
start_continuous_recognition()Keeps listening and raises events until you stop itLive captions, long recordings, dictation
Batch transcription (REST)Asynchronous jobs over large sets of stored audio files; results arrive laterCall-center archives, overnight processing

Continuous recognition is event-driven: you connect handlers, start it, and results arrive through those handlers:

  • recognizing fires repeatedly with partial, growing hypotheses — good for showing live text, wrong for building a transcript (you would repeat words).
  • recognized fires once per finalized phrase — append evt.result.text here.
  • session_stopped and canceled signal that recognition ended; wait for them, then call stop_continuous_recognition(), so the program does not exit before the whole file is processed.

Speech results report their outcome through result.reason (a ResultReason value).

result.reasonMeaningWhat to read
RecognizedSpeechSpeech was recognizedresult.text
NoMatchAudio was processed but no speech could be recognized (silence, noise)result.no_match_details
CanceledThe request was stopped; with cancellation_details.reason == CancellationReason.Error, the request failed — most often a wrong, missing or rotated key or endpoint, though other service-side failures (such as a bad request or throttling) report the same wayresult.cancellation_details.error_details
RecognizingSpeechAn intermediate hypothesis during recognition (seen in recognizing events)Partial text only
SynthesizingAudioCompletedA text-to-speech success, returned by a synthesizerresult.audio_data

Whenever result.text is empty, result.reason and its details explain why, so a robust app logs the reason alongside each transcript. Recognition quality also depends on speech_recognition_language matching the audio's locale.

How do you turn text into speech in code?

Create a SpeechSynthesizer from the SpeechConfig and an AudioOutputConfig, then call speak_text_async() for plain text or speak_ssml_async() for SSML that controls how the text is spoken. A successful call returns ResultReason.SynthesizingAudioCompleted.

speech_config.speech_synthesis_voice_name = "en-US-AvaMultilingualNeural"
audio_out = speechsdk.audio.AudioOutputConfig(filename="reminder.wav")
synth = speechsdk.SpeechSynthesizer(speech_config=speech_config, audio_config=audio_out)
result = synth.speak_text_async("Your appointment is tomorrow at nine.").get()

SSML (Speech Synthesis Markup Language) is how you change delivery while keeping the same voice: a voice element names the voice, prosody adjusts rate, pitch and volume, and break inserts a pause of a set length. SSML must go to speak_ssml_async(); passing markup to speak_text_async() treats it as plain text to read.

NeedUse
Speak plain text in a chosen voicespeech_synthesis_voice_name + speak_text_async()
Slow down, change pitch, add pausesSSML with prosody / break + speak_ssml_async()
Play the speech nowAudioOutputConfig(use_default_speaker=True)
Keep the speech as a file for laterAudioOutputConfig(filename=...)

When should you use the Speech SDK, a multimodal model or Voice Live?

Use the Speech SDK when you need only one direction (transcribe or speak) with fine control, an audio-capable multimodal model when a short spoken question needs an answer in one call, and Azure Speech Voice Live when you are building a real-time, two-way voice agent. Voice Live is a fully managed speech-to-speech API that combines speech recognition, a generative model of your choice and speech synthesis in one service, and handles conversational behaviour such as callers interrupting. You can use it from code with the azure-ai-voicelive package, or switch on Voice mode for an agent in the Foundry portal.

DesignServices and callsFits
Speech to text → model → text to speechTwo services, three calls you wire togetherFull control of each step (custom voice, SSML, a specific recognition locale)
Audio-capable chat modelOne deployment, one call per turnAnswering a short spoken prompt, optionally with a spoken reply
Voice Live APIOne managed service with a live sessionReal-time voice agents and assistants with natural turn-taking

Each design trades control for simplicity: the separate pipeline exposes every step, while Voice Live manages the conversational loop for you. Azure Language works on text, so in a voice app it applies to transcripts rather than to audio.

Tip. This task tests building, not just naming: short practical scenarios with brief Python snippets, where you pick the method, argument, object or result field that completes a text-analysis or speech app. Expect answer options drawn from the same SDK, so you need to know what each similar-looking method, setting, event and result field actually does, and some items ask you to select two lines or checks that together meet a requirement.

Key takeaways
  • Azure Language returns consistent, structured fields; a deployed model returns flexible generated text — pick by the shape of result the app needs.
  • TextAnalyticsClient takes the Foundry resource endpoint and AzureKeyCredential(key); no separate Language resource and no deployment name.
  • recognize_pii_entities() is the method in the Language SDK table that returns redacted_text — the masked copy to store; its entities list describes each finding.
  • Language results come back one per document in order; a bad document is a DocumentError with is_error True, not an exception.
  • The OpenAI client for a Foundry deployment uses https://<resource>.openai.azure.com/openai/v1/ and the deployment name as model.
  • Spoken input goes in an input_audio part as base64; spoken output needs modalities ['text','audio'] plus audio voice/format, and its words are in message.audio.transcript.
  • SpeechConfig holds key, endpoint and language/voice settings; AudioConfig is recognizer input, AudioOutputConfig is synthesizer output.
  • Use the recognized event and wait for session_stopped in continuous recognition; NoMatch means no speech was recognized, Canceled with Error means the request failed (most often a key or endpoint problem).
  • SSML with prosody and break goes to speak_ssml_async(); Voice Live is the managed speech-to-speech API for real-time voice agents.
  • The Azure Language MCP tool gives an agent text analyzers (language detection, sentiment, entities, key phrases, summarization, PII) and works on text only; audio, scanned documents and images need Speech, Content Understanding or vision.

Frequently asked questions

Do I need a separate Azure Language or Speech resource to use them with Microsoft Foundry?

No. A Microsoft Foundry resource already includes Azure Language and Azure Speech in Foundry Tools, so the Language and Speech SDKs use the Foundry resource's endpoint and key, and the Azure Language MCP tool for agents is connected with the Foundry resource name.

Why use Azure Language instead of prompting a GPT model for sentiment or PII?

Azure Language analyzers return structured, consistent results — fixed fields such as a sentiment label with confidence scores, a language code or redacted text — that code can store directly. A prompted model returns generated text whose wording and layout can vary between runs, so it suits flexible, multi-step instructions better than repeatable pipelines.

What happens if one document in a Language SDK batch is invalid?

The other documents still succeed. The Language SDK returns one result per input document in the same order, and the rejected document appears in its position as a DocumentError with is_error set to True, so code should check is_error on each result before reading its fields.

How do you send audio to a GPT-4o audio model in Foundry?

Read the recording, base64-encode it and include it in the user message as an input_audio content part with the data and its format (for example wav). For a spoken reply, also set modalities to text and audio and pass an audio argument with a voice and output format; the reply's words are in message.audio.transcript and the sound in message.audio.data.

What is the difference between recognize_once and continuous recognition in the Speech SDK?

recognize_once returns after a single utterance, which suits short commands or questions. Continuous recognition keeps listening until it is stopped and delivers each finalized phrase through the recognized event, which suits live captions and long recordings; the app waits for session_stopped before calling stop_continuous_recognition.

Why does my Speech SDK call return Canceled with CancellationReason.Error?

That result means the request failed. The most common cause is a key or endpoint in SpeechConfig that is wrong, missing or has been rotated, though other service-side failures such as a bad request or throttling report the same reason. cancellation_details.error_details describes the specific failure.

How do you make Azure text to speech pause or speak more slowly?

Write SSML: a voice element keeps the chosen neural voice, a prosody element changes rate or pitch, and a break element inserts a pause. Send the SSML with speak_ssml_async(), because speak_text_async() treats its input as plain text.

What is Azure Speech Voice Live?

Voice Live is a fully managed speech-to-speech API for real-time voice agents. It combines speech recognition, a generative model and speech synthesis in one service and handles natural turn-taking such as interruptions, so you do not wire separate speech-to-text, model and text-to-speech components together.

Source

This lesson covers the "Implement AI solutions by using Microsoft Foundry" domain of the official AI-901 exam guide. Vendors revise their guides — check the source for the current version.

Sign up free to mark lessons complete, bookmark topics and track your exam readiness.

Spotted a mistake in this lesson?