Azure
AzureRealtimeModel connects to Azure OpenAI’s GA
realtime protocol with the server-side Pydantic AI agent loop. Start with the
realtime quickstart or text-to-audio example.
Azure OpenAI realtime uses the OpenAI realtime stack, so install pydantic-ai-slim with the
openai-realtime optional group:
pip install "pydantic-ai-slim[openai-realtime]"
uv add "pydantic-ai-slim[openai-realtime]"
Set AZURE_OPENAI_ENDPOINT and AZURE_OPENAI_API_KEY as for the
Azure AI Foundry provider. Use the azure: prefix followed
by your Azure deployment name:
from pydantic_ai import Agent
agent = Agent(instructions='You are a helpful voice assistant.')
async def main():
async with agent.realtime('azure:my-realtime-deployment').session() as session:
await session.send('Say hello.')
async for part in session.stream_transcripts():
print(f'{part.speaker}: {part.transcript}')
#> assistant: Hello from the realtime assistant.
if part.speaker == 'assistant':
break # keep listening in a real call; we stop after one reply
(This example is complete, it can be run “as is” — you’ll need to add asyncio.run(main()) to run main)
For explicit configuration, use
AzureProvider.for_realtime(). It accepts
a bare resource endpoint or its /openai/v1 form. The GA realtime protocol uses /openai/v1/realtime
and does not take an api_version. Requests authenticate with the resource API key by default, or
with a Microsoft Entra ID token when a credential is passed (see
Browser WebRTC and Microsoft Entra ID).
Pass the Azure deployment name, which is chosen when the model is deployed and need not match the underlying model ID. Available realtime models and regions are documented in the Azure OpenAI realtime documentation.
Azure uses
OpenAIRealtimeModelSettings — the
realtime counterpart of model run settings — including the
shared settings plus:
openai_voicefor the provider voice;openai_input_noise_reductionandopenai_output_speed;openai_turn_detectionfor server or semantic VAD (see turn detection);openai_truncationfor session context management.
See OpenAI settings for the common settings shape. Azure realtime does not
expose temperature through Pydantic AI.
Azure resolves the input_transcription_model setting against
deployments in your resource. The default 'auto' selects gpt-realtime-whisper; a resource
without a matching deployment emits a DeploymentNotFound transcription error on every turn.
Deploy a realtime-capable transcription model such as gpt-realtime-whisper or
gpt-4o-transcribe, then set input_transcription_model to that deployment name. A classic
whisper deployment is not accepted. Set the field to None to disable transcription and use
audio_retention='input_audio' if the spoken turn must remain
available as audio.
Azure OpenAI supports the same browser WebRTC flow as OpenAI — the audio flows browser ↔ Azure directly
while your backend runs a control-plane sideband. See Connecting a frontend
for the topology, and use
AgentRealtime.answer_webrtc_offer /
AgentRealtime.create_client_secret exactly as on OpenAI.
Azure relays the offer with webrtcfilter=on, which limits the events forwarded to the browser to a
safe subset so the session instructions stay on the server’s control connection.
Azure requests authenticate with the resource’s API key by default. To use Microsoft Entra ID
instead — so no API key is involved, e.g. when the resource is locked to managed identity — pass a
credential (any azure.identity
credential, e.g. DefaultAzureCredential). It authenticates every request to the resource — the
realtime WebSocket session and the WebRTC signaling — with a bearer token for the Azure OpenAI data
plane (scope https://ai.azure.com/.default), which requires the Cognitive Services User role on
the resource:
from azure.identity import DefaultAzureCredential
from pydantic_ai.providers.azure import AzureProvider
from pydantic_ai.realtime.azure import AzureRealtimeModel
model = AzureRealtimeModel(
'gpt-realtime',
# `entra_authenticated=True` so no resource key is required — a resource locked to managed
# identity has none. Omit `provider=` entirely to take the endpoint from `AZURE_OPENAI_ENDPOINT`.
provider=AzureProvider.for_realtime(
azure_endpoint='https://my-resource.openai.azure.com', entra_authenticated=True
),
credential=DefaultAzureCredential(),
)
# The realtime session, `answer_webrtc_offer`, and `create_client_secret` now authenticate with an Entra
# bearer token; the browser only ever receives the short-lived ephemeral secret, never it or the API key.
| Feature | Support | Notes |
|---|---|---|
| Audio format | Full feature support | Mono PCM16, 24 kHz input and output |
| Text output | Full feature support | Select with output_modality='text' |
| Image input | Full feature support | Images provide context for the next turn |
| Manual turns | Full feature support | turn_detection=False plus commit/create verbs |
| Interruption/truncation | Full feature support | interrupt(played_ms=...) records the heard cutoff |
| Input transcription | Limited parameter support | Requires a compatible transcription deployment in the Azure resource |
| Native tools | Unsupported | Configure local fallbacks for web capabilities |
| Usage | Full feature support | Token, audio, and cache breakdowns |
| Reconnection | Full feature support | Pydantic AI replays completed local history; in-flight media is lost |
See Audio, images, and transcripts, Turns and interruptions, Tools, and Connection lifecycle for the provider-agnostic workflows.
- A failed input transcription leaves the user turn represented as
retained audio when available, or as a content-less
SpeechPartotherwise. - This page covers Azure OpenAI Realtime only; Azure AI Voice Live support is coming in #6642.