Skip to content

Timeouts

Bounding how long one step inside a run may take, and ending a run from inside a tool, are answered by separate mechanisms with separate failure modes. This page maps them. To stop a run that is already in flight, see Cancelling a Run.

Bounding how long a step takes

Each knob below bounds a different unit of work. None of them bounds the wall-clock duration of a whole run.

What you want to boundHow to set itWhat happens on expiry
A single model request attempt — a provider SDK client’s retries re-arm it for every attempttimeout on ModelSettingsThe provider client raises; the run fails unless a FallbackModel or a transport retry handles it
A function tool callAgent(tool_timeout=...), or timeout= on an individual tool — see Tool TimeoutThe model receives a retry prompt 'Timed out after N seconds.', consuming that tool’s retry budget. A def tool is not actually stopped: the deadline is enforced around the await, so the worker thread runs to completion
A hook functiontimeout= on the @hooks.on.* decoratorHookTimeoutError, which is an AgentRunError and aborts the run. Like a def tool, a def hook is not actually stopped: the worker thread runs to completion
Connecting to an MCP serverMCPToolset(init_timeout=...), default 5 secondsThe connection and initialize handshake fail
A single MCP requestMCPToolset(read_timeout=...), default 300 secondsThe request fails; under the default tool_error_behavior='retry' the model sees it as a retryable tool error
Opening a realtime sessionhandshake_timeout on RealtimeModelSettings, default 30 seconds — OpenAI, Azure OpenAI, xAI, and GeminiOpening the session raises RealtimeError. On a reconnect it consumes a ReconnectPolicy attempt instead
Total work done by a runUsageLimits — requests, tool calls, tokens, or cost — see Usage LimitsUsageLimitExceeded
Wall-clock duration of a whole runNothing built in — wrap agent.run() in asyncio.timeout (Python 3.11+) or anyio.fail_after(), or cancel a CancellationToken from a timerThe run is cancelled

Two of these need qualifying:

  • ModelSettings['timeout'] is applied per model class, not universally. The model classes that forward it to their provider client are listed under ModelSettings.timeout; the ones built on OpenAI’s inherit the forwarding from OpenAIChatModel / OpenAIResponsesModel. Other model classes ignore the setting, and the timeout on the HTTP client they were built with applies instead. When Pydantic AI creates that client itself, it defaults to a 600-second total timeout with a 5-second connect timeout. Google and Mistral additionally reject an httpx.Timeout object and accept only a number of seconds.

    To bound a request on a model class that ignores the setting, configure the timeout where that provider actually takes one. Most providers accept your own http_client, but several don’t: XaiProvider takes a client-level timeout (or a preconfigured xai_client), BedrockProvider takes aws_read_timeout and aws_connect_timeout (or a preconfigured bedrock_client), and HuggingFaceProvider rejects http_client outright in favor of hf_client.

    On a client Pydantic AI created, including one from create_async_httpx2_client(), a request timeout given in seconds can shorten but never lengthen the client’s connect timeout (5 seconds by default) and pool timeout (600 seconds by default). This includes the 600 seconds google-genai sends with every Gemini request. To connect for longer, pass an httpx.Timeout whose connect differs from its other phases, or your own http_client.

  • Tool timeouts are enforced by FunctionToolset only, and each toolset carries its own. Agent(tool_timeout=...) sets the default for tools you register on the agent — it does not reach into a FunctionToolset you constructed yourself and passed via toolsets=[...]. Give that toolset its own FunctionToolset(timeout=...), or set timeout= on the individual tools. Tools coming from an MCP server, an external toolset, or a custom AbstractToolset read neither; bound those with the server-side or transport-level timeout instead.

If you enforce a deadline inside a tool body yourself, catch the TimeoutError and re-raise it as ModelRetry or ToolFailed rather than letting it escape. What happens to a bare TimeoutError depends on whether that tool has a timeout of its own:

  • No timeout on the tool or its toolset. It is an ordinary exception and propagates out of the agent run — unless a capability implements on_tool_execute_error, which can turn it into a replacement tool result or a ModelRetry.
  • A timeout is configured. The call runs inside anyio.fail_after(timeout), which signals expiry with TimeoutError too, so a TimeoutError you raised yourself is indistinguishable from the deadline expiring and becomes the same 'Timed out after N seconds.' retry prompt — reporting a deadline that may never have passed.

Re-raising in the tool is the more local choice; the hook is for applying one policy across every tool.

Ending a run from inside a tool

What a tool raises decides whether the run continues, and what the model gets to see:

RaiseRun continues?The model sees
ModelRetryYesA retry prompt asking it to correct the call — consumes that tool’s retry budget
ToolFailedYesA failed tool result to adapt to — does not consume the retry budget
ApprovalRequired / CallDeferredEnds the run with a DeferredToolRequests output, unless a HandleDeferredToolCalls handler resolves the call inlineNothing yet — see Deferred Tools
Any other exceptionNoBy default nothing — it propagates out of agent.run(). A capability implementing on_tool_execute_error sees it first and can return a replacement tool result or raise ModelRetry, letting the run continue

The deferred row reads differently inside a realtime session, which has no way to pause: a live conversation can’t wait for an out-of-band result. A HandleDeferredToolCalls handler still gets the chance to resolve the call inline, but where a run would end with a DeferredToolRequests output, a session instead answers the model with an explanation that the tool can’t complete during the session, and keeps going. See Deferred and approval-required tools.

A tool can also end the run without raising, by calling RunContext.cancel() — the run ends with RunCancelled and the tool’s return value is discarded. See Cancelling the Run from a Tool.

There is no exception that ends a run early with a successful output. To let a tool finish the run with a value, make that value the run’s output: give the agent an output tool the model can call, or an output function that produces the result.