pydantic_ai.profiles
Bases: TypedDict
Describes how requests to and responses from specific models or families of models need to be constructed and processed to get the best results, independent of the model and provider classes used.
All fields are optional; absent keys mean “use the documented default” (defaults are documented per field below and applied at access sites).
Subclasses (OpenAIModelProfile, AnthropicModelProfile, …) add provider-specific keys; cross-class merging via dict-spread is supported.
Whether the model supports tools. Default: True.
Type: bool
Whether the model supports text output. Default: True.
Type: bool
Whether the model natively supports tool return schemas. Default: False.
When True, the model’s API accepts a structured return schema alongside each tool definition. When False, return schemas are injected as JSON text into tool descriptions as a fallback.
Type: bool
Whether the model supports JSON schema output. Default: False.
This is also referred to as ‘native’ support for structured output.
Relates to the NativeOutput output type.
Type: bool
Whether the model supports a dedicated mode to enforce JSON output, without necessarily sending a schema. Default: False.
E.g. OpenAI’s JSON mode
Relates to the PromptedOutput output type.
Type: bool
Whether the model supports image output. Default: False.
Type: bool
Whether the model supports audio in user messages. Default: False.
Used when converting SpeechParts from realtime session history in
Model.prepare_messages: if True, retained audio is sent to the model as BinaryContent;
otherwise the transcript text is used.
No shipping profile sets this to True yet, so retained realtime audio is currently always
forwarded as transcript text on handoff; enabling it needs per-model-family verification that the
provider accepts audio in user messages.
Type: bool
Whether the provider’s API accepts SystemPromptParts inline at any position. Default: False.
When False, non-leading SystemPromptParts are wrapped as UserPromptParts with
<system>...</system> content in Model.prepare_messages. Leading ones still hoist to the
provider’s top-level system parameter.
APIs that only accept an inline system prompt in certain positions (e.g. Anthropic requires it
to follow a user turn) still set this to True; it’s on their model adapters to make the
positions the API rejects legal. Preserving the part’s authority is worth more than preserving
the exact position it was authored at — an instruction only governs the generation that follows
it, and that’s the same generation either way — so prefer adjusting placement over falling back
to the <system>...</system> rendering, which the model reads as user-authored. Anthropic slides
the entry past intervening user turns and gives it a minimal user turn to follow when nothing
legal precedes it.
Provider.model_profile is resolved from the model name alone, so when support also turns on
something it can’t see — which SDK client the provider was built with, say — the adapter narrows
this in its own Model.profile override, as Anthropic does for Microsoft Foundry. Narrowing the
flag rather than special-casing the adapter’s own rendering keeps Model.prepare_messages the
only place that knows the <system>...</system> fallback.
Type: bool
The default structured output mode to use for the model. Default: 'tool'.
Type: StructuredOutputMode
The instructions template to use for prompted structured output. The {schema} placeholder will be replaced with the JSON schema for the output. Default: DEFAULT_PROMPTED_OUTPUT_TEMPLATE.
Type: str
Whether to add prompted output template in native structured output mode. Default: False.
Type: bool
The transformer to use to make JSON schemas for tools and structured output compatible with the model. Default: None.
Type: type[JsonSchemaTransformer] | None
How long the provider keeps a cached prompt prefix when the request doesn’t ask for a specific retention. Default: None (unknown).
Measured from the last request that used the prefix. Only documented values are populated. When a
provider documents a range, the higher end is used: consumers of a 'cold' outlook are about to pay
for a full prefix re-write, so a false 'cold' sacrifices a live cache hit while a false 'warm'
merely defers maintenance. Because retention is provider infrastructure, providers populate this
field; model-family profile functions must never set it. Providers without an honest documented
expectation boundary leave it None.
Retention requested through model settings, such as anthropic_cache='1h' or
openai_prompt_cache_retention='24h', is resolved by
Model.resolve_cache_retention; this field is
what applies when the settings request nothing.
Consumed by prompt_cache_outlook to classify, from a
message history alone, whether the next request is likely to hit a warm cache. A 'cold' outlook is
a free moment to run history-mutating maintenance (compaction, pruning, repair): the next request
pays a full prefix re-write either way, so the marginal cache cost of the mutation is ~zero.
Type: timedelta | None
Whether the model supports thinking/reasoning configuration. Default: False.
When False, the unified thinking setting in ModelSettings is silently ignored.
Type: bool
Whether the model always uses thinking/reasoning (e.g., OpenAI o-series, DeepSeek R1). Default: False.
When True, thinking=False is silently ignored since the model cannot disable thinking.
Implies supports_thinking=True.
Type: bool
Whether the model thinks when the request doesn’t configure thinking. Default: False.
True for models that think unless told not to, such as Claude Opus 5, DeepSeek V4 and the OpenAI o-series. Pydantic AI
uses it to tell whether a request without a thinking setting will think, for example to decide whether a
tool call can be forced. Unlike thinking_always_enabled, it doesn’t mean thinking can’t be turned off.
Type: bool
Whether the model accepts a forced tool choice: tool_choice='required' or a specific tool. Default: True.
Some models reject forcing on every request, such as Claude Opus 5.5, Claude Fable 5.1 and Claude Mythos 5.1,
as do some OpenAI-compatible providers, such as Moonshot AI. When False, a forced tool choice that Pydantic AI
resolved itself (such as an output tool’s) falls back to 'auto', with the tools filtered to the requested
ones where the API can’t restrict the choice. An explicit forcing
tool_choice raises a UserError.
Type: bool
Whether the model accepts a forced tool choice while it thinks. Default: True.
DeepSeek’s V4 models, for example, only accept forcing while thinking is off. When False and the request
thinks, a forced tool choice is handled as if supports_forced_tool_choice were False. Whether the request
thinks accounts for thinking_enabled_by_default.
Type: bool
Whether the model answers a forced tool choice without thinking. Default: False.
Claude models accept a forced tool choice alongside adaptive thinking, but return no thinking for that
request. When True and the request thinks, Pydantic AI doesn’t force a tool choice it resolved itself (such
as an output tool’s): it falls back to 'auto', and a structured output_type defaults to
Native Output where the model supports it. An explicit forcing
tool_choice is still sent.
Type: bool
The tags used to indicate thinking parts in the model’s output. Default: DEFAULT_THINKING_TAGS.
Whether to ignore leading whitespace when streaming a response. Default: False.
This is a workaround for models that emit `<think> </think>
or an empty text part ahead of tool calls (e.g. Ollama + Qwen3), which we don't want to end up treating as a final result when usingrun_streamwithstra validoutput_type`.
This is currently only used by OpenAIChatModel, HuggingFaceModel, GroqModel, and BedrockConverseModel.
Type: bool
The set of native tool types that this model/profile supports. Default: SUPPORTED_NATIVE_TOOLS (all).
Type: frozenset[type[AbstractNativeTool]]
The maximum number of tokens the model can handle in a single request, input and output combined. Default: None (unknown).
When no profile layer sets this, Model.profile fills it in from
genai-prices data if the model is known there.
Set it explicitly for custom or local models, e.g. profile={'context_window': 128_000}.
When the provider permits a tools entry whose schema is withheld. Default: None.
'standalone' permits the deferral flag on its own. 'with_tool_search' permits it only when a
tool-search tool is present in the same request. None means hidden tools can only be withheld
from the wire. Unsupported deferral is handled on a best-effort basis by withholding the tool.
Type: ToolDeferralMode | None
How the model natively expresses tools added mid-conversation. Default: None.
'by_reference' reveals a tool already declared in the request’s tool definitions (Anthropic
tool_addition blocks referencing a defer_loading entry); 'with_definitions' carries the full
newly available definitions in the reveal (OpenAI Responses additional_tools items). None means
no native channel: Model.prepare_messages projects the change into messages. Additions only —
tool removal (#6985) is not modeled yet and will get its own field.
Type: ToolAdditionMode | None
Deprecated: use tool_addition_mode instead.
Translated (with a deprecation warning) whenever profiles are merged; an explicit
tool_addition_mode in the same profile wins.
Type: ToolAdditionMode | None
Deprecated: use tool_deferral_mode instead.
True translates to tool_deferral_mode='with_tool_search' (with a deprecation warning)
whenever profiles are merged. False carried no signal on its own — deferral capability came
from native tool-search support — so it is dropped; an explicit tool_deferral_mode in the
same profile wins.
Type: bool
Bases: ABC
Walks a JSON schema, applying transformations to it at each level.
The transformer is called during a model’s prepare_request() step to build the JSON schema before it is sent to the model provider.
Note: We may eventually want to rework tools to build the JSON schema from the type directly, using a subclass of pydantic.json_schema.GenerateJsonSchema, rather than making use of this machinery.
The strict parameter forces the conversion of the original JSON schema (self.schema) of a ToolDefinition or OutputObjectDefinition to a format supported by the model provider.
The “strict mode” offered by model providers ensures that the model’s output adheres closely to the defined schema. However, not all model providers offer it, and their support for various schema features may differ. For example, a model provider’s required schema may not support certain validation constraints like minLength or pattern.
Default: strict
Whether the schema is compatible with strict mode.
This value is used to set ToolDefinition.strict or OutputObjectDefinition.strict when their values are None.
Default: True
@abstractmethod
def transform(schema: JsonSchema) -> JsonSchema
Make changes to the schema.
JsonSchema
Bases: JsonSchemaTransformer
Transforms the JSON Schema to inline $defs.
Object keywords (properties, additionalProperties, patternProperties) are only walked when type is
'object', and array keywords (items, prefixItems) only when it is 'array'. On a schema with no type, or
with a type list such as ['object', 'null'], they are left as written, so a $ref inside them is not inlined
and can point at a definition the output no longer contains. If the schema is only meant to accept objects (or
arrays), set type to 'object' (or 'array') to have it inlined; for a nullable one, put it in an anyOf
with {'type': 'null'} instead of using a type list.
def merge_profile(
base: ModelProfile | None,
*overrides: ModelProfile | None,
) -> ModelProfile
Merge profiles via dict-spread. Later arguments override earlier ones; None is treated as empty.
This is the canonical way to layer profiles in providers and tests; replaces the old ModelProfile.update() method.
Deprecated key spellings are translated per input before spreading, so a legacy key in an
override still overrides the base.
def prompt_cache_outlook(
messages: Sequence[ModelMessage],
*,
profile: ModelProfile | None = None,
retention: timedelta | None = None,
now: datetime | None = None,
) -> PromptCacheOutlook
Predict whether the provider’s prompt cache is still warm for the next request on this history.
This is a pure function of the message history and an expectation boundary — it holds no state and makes no
requests, so a history processor, capability, or plain application code can call it with just a
message history to decide whether the next turn is a cheap moment for history-mutating maintenance
(compaction, pruning, repair). When the outlook is 'cold' the next request pays a full prefix
re-write anyway, so the marginal cache cost of mutating history right now is ~zero.
The prediction compares the most recent ModelResponse.timestamp
in messages against now: an idle gap within the retention window is 'warm', a larger gap is 'cold'.
Responses are the anchor because they mark the provider’s last confirmed use of the cache — a request that
has no response after it (like the just-appended request a history processor
sees, which hasn’t been sent yet) never touched the cache, so its timestamp must not reset the idle clock.
Cache points in the history extend whichever boundary applies to their largest TTL, assuming they were honored by the provider that served the requests.
PromptCacheOutlook — 'warm', 'cold', or 'unknown' (see PromptCacheOutlook).
messages : Sequence[ModelMessage]
The message history the next request would be built on, oldest first.
profile : ModelProfile | None Default: None
The model profile whose default_cache_retention
is used as the expectation boundary when retention is None.
retention : timedelta | None Default: None
The retention requested for these requests, replacing the profile’s default. With a model
and its settings in hand, pass
model.resolve_cache_retention(model_settings):
it returns None when the settings request nothing, so the profile’s default still applies.
The reference time to measure idleness against. Defaults to the current UTC time; inject a fixed value for deterministic tests.
Acceptable shapes for the profile= argument on a Model.
- A
ModelProfiledict — a partial profile, merged on top of the provider’s resolved default. - A
Callable[[ModelProfile], ModelProfile]— receives the provider’s resolved default (withDEFAULT_PROFILEalready merged in) and returns the final profile (full control: replace, derive, ignore the default).
Provider classes still expose Provider.model_profile(model_name) (Callable[[str], ModelProfile | None]) — that’s a separate concept used internally by Model.profile to resolve the provider’s default for a given model name.
Type: TypeAlias Default: ModelProfile | Callable[['ModelProfile'], 'ModelProfile']
Predicted state of the provider’s prompt cache for the next request built on a message history.
'warm': the last request happened within the provider’s documented expectation boundary, so the cached prefix is likely still available and the next request should hit it.'cold': the last request happened longer ago than the retention window, so the prefix has likely been evicted and the next request will pay full input price regardless — a free moment to mutate history.'unknown': there’s no retention figure for the model, or the history has no usable timestamp, so no prediction can be made. Treat like'warm'for scheduling (never mutate on a guess).
Type: TypeAlias Default: Literal['warm', 'cold', 'unknown']
Fully populated default ModelProfile. Used as the base layer when resolving a model’s effective profile.
Type: ModelProfile Default: {'supports_tools': True, 'supports_text_output': True, 'supports_tool_return_schema': False, 'supports_json_schema_output': False, 'supports_json_object_output': False, 'supports_image_output': False, 'supports_audio_input': False, 'default_structured_output_mode': 'tool', 'prompted_output_template': DEFAULT_PROMPTED_OUTPUT_TEMPLATE, 'native_output_requires_schema_in_instructions': False, 'json_schema_transformer': None, 'default_cache_retention': None, 'supports_thinking': False, 'thinking_always_enabled': False, 'thinking_enabled_by_default': False, 'supports_forced_tool_choice': True, 'supports_forced_tool_choice_with_thinking': True, 'forced_tool_choice_disables_thinking': False, 'thinking_tags': DEFAULT_THINKING_TAGS, 'ignore_streamed_leading_whitespace': False, 'supported_native_tools': SUPPORTED_NATIVE_TOOLS, 'context_window': None, 'tool_deferral_mode': None, 'tool_addition_mode': None}
Default instructions template for prompted structured output. The {schema} placeholder is replaced with the JSON schema for the output.
Default: dedent("\n Always respond with a JSON object that's compatible with this schema:\n\n {schema}\n\n Don't include any text or Markdown fencing before or after.\n ")
Default (start_tag, end_tag) pair for parsing thinking content out of text responses.
Type: tuple[str, str] Default: ('<think>', '</think>')
Bases: ModelProfile
Profile for models used with OpenAIChatModel.
ALL FIELDS MUST BE openai_ PREFIXED SO YOU CAN MERGE THEM WITH OTHER MODELS.
Non-standard field name used by some providers for model thinking content in Chat Completions API responses. Default: None.
Plenty of providers use custom field names for thinking content. Ollama and newer versions of vLLM use reasoning,
while DeepSeek, older vLLM and some others use reasoning_content.
Notice that the thinking field configured here is currently limited to str type content.
If openai_chat_send_back_thinking_parts is set to 'field', this field must be set to a non-None value.
Whether the model includes thinking content in requests. Default: 'auto'.
This can be:
'auto'(default): Automatically detects how to send thinking content. If thinking was received in a custom field (tracked viaThinkingPart.idandThinkingPart.provider_name), it’s sent back in that same field. Otherwise, it’s sent using tags. Only thereasoningandreasoning_contentfields are checked by default when receiving responses. If your provider uses a different field name, you must explicitly setopenai_chat_thinking_fieldto that field name.'tags': The thinking content is included in the maincontentfield, enclosed within thinking tags as specified inthinking_tagsprofile option.'field': The thinking content is included in a separate field specified byopenai_chat_thinking_field.False: No thinking content is sent in the request.
Defaults to 'auto' to ensure thinking is sent back in the format expected by the model/provider.
Type: Literal[‘auto’, ‘tags’, ‘field’, False]
This can be set by a provider or user if the OpenAI-”compatible” API doesn’t support strict tool definitions. Default: True.
Type: bool
A list of model settings that are not supported by this model. Default: ().
Deprecated: use supports_forced_tool_choice instead.
Translated (with a deprecation warning) whenever profiles are merged.
Type: bool
Deprecated: use supports_forced_tool_choice_with_thinking instead.
Translated (with a deprecation warning) whenever profiles are merged.
Type: bool
The role to use for the system prompt message. If not provided, defaults to 'system'.
Type: OpenAISystemPromptRole | None
Whether the Chat Completions API accepts more than one system-role message at the start of the conversation. Default: True.
OpenAI itself and most compatible providers accept multiple system messages, so this defaults to True.
Set to False for strict OpenAI-compatible backends (e.g. some LiteLLM/vLLM deployments) that require
exactly one initial system message; consecutive system messages at the start will be merged into one
(joined with two newlines) before being sent.
Type: bool
Whether a streamed Chat Completions response must include a non-null finish_reason. Default: False.
When enabled, reaching clean EOF before any chunk supplies a finish_reason raises
ModelAPIError, instead of treating the response as a 'stop'.
This defaults to False because OpenAI-compatible APIs do not consistently guarantee the field.
Type: bool
Whether the model supports web search in Chat Completions API. Default: False.
Type: bool
The encoding to use for audio input in Chat Completions requests. Default: 'base64'.
'base64': Raw base64 encoded string. (Default, used by OpenAI)'uri': Data URI (e.g.data:audio/wav;base64,...).
Type: Literal[‘base64’, ‘uri’]
Whether the Chat API supports file URLs directly in the file_data field. Default: False.
OpenAI’s native Chat API only supports base64-encoded data, but some providers like OpenRouter support passing URLs directly.
Type: bool
Whether the model supports including encrypted reasoning content in the response. Default: False.
Type: bool
Whether the model supports reasoning (o-series, GPT-5+). Default: False.
When True, sampling parameters may need to be dropped depending on reasoning_effort setting.
Type: bool
Deprecated: use thinking_enabled_by_default instead.
Translated (with a deprecation warning) whenever profiles are merged.
Type: bool
Whether the model accepts reasoning_effort='none' and allows sampling parameters (temperature, top_p, etc.)
while reasoning is off. Default: False.
The GPT-5.1+ mainline models support turning reasoning off via effort='none', and sampling params are
accepted in that mode. When reasoning is enabled (low/medium/high/xhigh), sampling params are not supported.
Whether the model reasons by default is tracked separately by thinking_enabled_by_default.
Type: bool
Whether the model accepts reasoning_effort='minimal'. Default: True.
Disabled for GPT-5.6 models, whose documented reasoning efforts exclude minimal. When disabled,
unified thinking='minimal' falls back to reasoning_effort='low'.
See https://developers.openai.com/api/docs/guides/latest-model.
Explicit openai_reasoning_effort='minimal' settings are still passed through unchanged.
Type: bool
Whether the Responses API supports reasoning.mode ('standard' | 'pro') for this model. Default: False.
Currently only supported by the GPT-5.6 family.
Type: bool
Whether the Responses API accepts reasoning.context='all_turns' for this model. Default: False.
auto and current_turn are accepted by every reasoning model, so they are gated on
openai_supports_reasoning instead; only all_turns requires this flag.
Currently supported by the GPT-5.4, GPT-5.5, and GPT-5.6 families.
Type: bool
Whether the Responses API requires the status field on function tool calls to be None. Default: False.
This is required by vLLM Responses API versions before https://github.com/vllm-project/vllm/pull/26706. See https://github.com/pydantic/pydantic-ai/issues/3245 for more details.
Type: bool
Whether the Responses API accepts text.format of type json_schema for this model. Default: False.
Only needed when the Responses API is more capable than Chat Completions for the same model, as with
DeepSeek, whose Chat Completions endpoint rejects response_format of type json_schema with
This response_format type is unavailable now while its Responses endpoint honors the schema. When set,
OpenAIResponsesModel enables supports_json_schema_output on its resolved profile, so
NativeOutput becomes available on the Responses API alone.
Type: bool
Whether Responses API tool call IDs are only unique within one response. Default: False.
When enabled, response IDs are incorporated into tool call IDs as responses are ingested so
normalized message history keeps the history-wide uniqueness required by Pydantic AI. The qualified
response_id:tool_call_id form is restored to the original provider tool call ID when history is
replayed.
Type: bool
Whether the Responses API accepts function calls interleaved with other assistant items in one
assistant turn. Default: True.
DeepSeek’s Responses endpoint merges each
function call into the assistant message next to it, so an assistant item between two calls
splits them into separate messages that each carry an unanswered call
(#7430). When this is False, such a turn
has its calls moved to the end before the request goes out; the message history you hold is
unchanged.
Reordering is best-effort: a turn carrying a native or compaction item, or a call still waiting on its result, is sent in its original order.
Type: bool
Whether the Responses API supports the phase field on assistant messages. Default: False.
phase labels an assistant message as intermediate commentary or the final_answer. When the model
supports it, OpenAI recommends preserving and sending it back unchanged on every assistant message in
follow-up requests; dropping it can cause preambles to be interpreted as final answers and degrade
behavior in long-running or tool-heavy flows.
Supported by gpt-5.3-codex, gpt-5.4 and later mainline models. The official OpenAI Responses API
silently ignores the field on older models, but defaults to False so we don’t risk sending an
unrecognized field to OpenAI-compatible APIs (vLLM, Bifrost, …) that haven’t been verified to accept it.
Type: bool
Whether the Chat Completions API supports document content parts (type='file'). Default: True.
Some OpenAI-compatible providers (e.g. Azure) do not support document input via the Chat Completions API.
Type: bool
Whether the Chat Completions API accepts the max_completion_tokens field for the max_tokens setting. Default: True.
OpenAI itself (including the o-series reasoning models) uses max_completion_tokens, the field that caps
visible output plus reasoning tokens, so this defaults to True. Many OpenAI-compatible providers (e.g.
OpenRouter) only accept the older max_tokens field; set this to False for those so the max_tokens
setting is sent as max_tokens instead.
Type: bool
Whether the model supports OpenAI explicit prompt cache breakpoints. Default: False.
When enabled, CachePoint markers are translated into
prompt_cache_breakpoint fields on the preceding content block, on both the Chat Completions and
Responses APIs. When disabled, CachePoint markers are filtered out.
Type: bool
Whether the Responses endpoint serves streaming responses only. Default: False.
When True, nominally non-streaming requests are sent with stream=True and aggregated from the
terminal response.completed event, so callers see an ordinary non-streaming ModelResponse.
Set for subscription-auth endpoints (e.g. OpenAI Codex) that reject stream=False.
Type: bool
Whether the Responses endpoint requires store=false on every request. Default: False.
When True, store=false is sent even when no openai_store setting is given (the field cannot be
omitted), and an explicit openai_store=True setting is silently overridden, consistent with how
other backend-rejected settings are dropped.
Type: bool
Whether the provider exposes server-side input-token counting (responses/input_tokens). Default: True.
When False, count_tokens() raises a UserError instead of calling a missing endpoint.
Type: bool
Bases: JsonSchemaTransformer
Recursively handle the schema to make it compatible with OpenAI strict mode.
See https://platform.openai.com/docs/guides/function-calling?api-mode=responses#strict-mode for more details, but this basically just requires:
additionalPropertiesmust be set to false for each object in the parameters- all fields in properties must be marked as required
def validate_openai_profile(profile: ModelProfile) -> None
Validate an OpenAI-compatible profile after resolution. Called from OpenAIChatModel.__init__.
def openai_model_profile(model_name: str) -> ModelProfile
Get the model profile for an OpenAI model.
def is_openai_live_model(model_name: str) -> bool
Whether a model name belongs to OpenAI’s GPT-Live API rather than its Realtime API.
The two are different protocols on the same provider, so the model name is what picks between
them — see OpenAILiveModel.
def openai_live_model_profile(model_name: str) -> RealtimeModelProfile
Get the realtime model profile for an OpenAI GPT-Live model.
Live is far more constrained than the Realtime API: it owns turn-taking entirely (no manual turns, no server-side interruption or truncation), speaks rather than writes, takes an image only for its delegated backend to respond to, and seeds from text alone. It also has no end-of-turn frame, so Pydantic AI infers the boundary.
RealtimeModelProfile
def openai_realtime_model_profile(model_name: str) -> RealtimeModelProfile
Get the realtime model profile for an OpenAI realtime model.
RealtimeModelProfile
Maps unified thinking values to OpenAI reasoning_effort strings.
Type: dict[ThinkingLevel, str] Default: {True: 'medium', False: 'none', 'minimal': 'minimal', 'low': 'low', 'medium': 'medium', 'high': 'high', 'xhigh': 'xhigh'}
Sampling parameter names that are incompatible with reasoning.
These parameters are not supported when reasoning is enabled (reasoning_effort != ‘none’). See https://platform.openai.com/docs/guides/reasoning for details.
Default: ('temperature', 'top_p', 'presence_penalty', 'frequency_penalty', 'logit_bias', 'openai_logprobs', 'openai_top_logprobs')
def openai_codex_model_profile(model_name: str) -> ModelProfile
Get the model profile for OpenAI Codex subscription-auth models.
The Codex backend speaks the Responses API with a narrower dialect than the standard OpenAI
endpoint: it serves streaming responses only, requires store=false, rejects sampling/tuning
request fields (verified live on PR #6433),
and does not expose server-side input-token counting.
Bases: ModelProfile
Profile for models used with AnthropicModel.
ALL FIELDS MUST BE anthropic_ PREFIXED SO YOU CAN MERGE THEM WITH OTHER MODELS.
Whether the model supports fast inference speed (anthropic_speed='fast'). Default: False.
Currently Claude Opus 4.6, 4.7, 4.8, 5, and 5.5 support fast mode. See the Anthropic docs for the latest list.
Type: bool
Whether the model supports adaptive thinking (Sonnet 4.6+, Opus 4.6+). Default: False.
When True, unified thinking translates to {'type': 'adaptive'}.
When False, it translates to {'type': 'enabled', 'budget_tokens': N}.
Adaptive thinking, unlike extended thinking, accepts a forced tool_choice, so this also decides whether
an explicit forcing tool_choice raises alongside unified thinking.
Type: bool
Whether the model supports the effort parameter in output_config (Opus 4.5+, Sonnet 4.6+). Default: False.
When True and the unified thinking level is a string (e.g. ‘high’), it is also
mapped to output_config.effort.
Type: bool
Whether the model supports Anthropic-managed dynamic filtering for web search/fetch. Default: False.
When enabled, Pydantic AI selects the web_search_20260209 / web_fetch_20260209 tool versions,
which let Claude filter web results via code execution before they enter context.
Type: bool
Whether the model supports the xhigh effort value in output_config. Default: False.
Claude Opus 4.7, 4.8, and 5 accept xhigh; older Anthropic models should use max instead.
Type: bool
Whether the model rejects budget-based thinking settings. Default: False.
Claude Opus 4.7, 4.8, and 5 require adaptive thinking and return a 400 for
{'type': 'enabled', 'budget_tokens': ...}.
Type: bool
Whether the model rejects sampling settings like temperature and top_p. Default: False.
Claude Opus 4.7, 4.8, 5, and 5.5 require these settings to be omitted from request payloads.
Type: bool
Whether the model rejects xhigh/max effort while thinking is explicitly disabled. Default: False.
Claude Opus 5 caps effort at high when anthropic_thinking={'type': 'disabled'} and returns a
400 for xhigh or max; Claude Opus 4.8 accepts the same combination. Claude Opus 5.5 rejects
disabled thinking at every effort level, so the flag doesn’t apply to it.
Type: bool
The Anthropic code execution tool version used when anthropic_code_execution_tool_version='auto'. Default: '20250825'.
Type: AnthropicCodeExecutionToolVersion
The Anthropic code execution tool versions supported by the model. Default: ('20250825',).
Type: tuple[AnthropicCodeExecutionToolVersion, …]
Whether the model supports output_config.task_budget. Default: False.
Anthropic currently documents task budgets as a Claude Opus 4.7 / 4.8 / 5 / 5.5 beta feature.
Type: bool
Deprecated: use supports_forced_tool_choice instead.
Translated (with a deprecation warning) whenever profiles are merged.
Type: bool
The most output tokens the model can generate in one response, thinking included. Default: None (unknown).
AnthropicModel sends it as max_tokens when the request doesn’t set one, so responses are only cut off at the
model’s limit, like on APIs where the output limit is optional.
Whether the model rejects a request whose input plus max_tokens exceeds its context window. Default: False.
Claude models older than Claude Sonnet 4.5 answer such a request with a 400, where later models accept it and stop
at the context window. When True, a request that doesn’t set max_tokens gets the lower default of 4096, so a
conversation close to the context window still fits. It’s also set for Claude 3 and 3.5, whose maximum output
(4,096 or 8,192 tokens) is below the higher default.
Type: bool
Whether the model binds each thinking block to the conversation prefix that produced it. Default: False.
Claude Fable 5.1, Claude Opus 5.5, and Claude Sonnet 5.5 reject a replayed thinking block once the system prompt text changes or a
non-deferred tool joins the tools array — both of which Pydantic AI causes by design, through
dynamic @agent.instructions and conditional toolsets. When True, Pydantic AI preserves the
account’s default behavior on the first request; if Anthropic rejects a stale block, it retries
once with thinking.block_binding.prefix_mismatch_behavior='drop_block' and warns after the
retry succeeds.
Type: bool
def resolve_anthropic_effort(
level: ThinkingEffort,
*,
supports_xhigh: bool,
) -> AnthropicEffort
Resolve a unified thinking effort level to the Anthropic output_config.effort value.
Shared between the direct Anthropic path and any provider that translates to the
Anthropic output_config wire shape (e.g. Bedrock Converse for Anthropic models).
Keeps ANTHROPIC_THINKING_EFFORT_MAP as the single source of truth for the
base mapping, while letting the xhigh passthrough decision live in one place.
AnthropicEffort
def anthropic_model_profile(model_name: str) -> ModelProfile | None
Get the model profile for an Anthropic model.
The unified sampling settings gated by anthropic_disallows_sampling_settings.
Models whose profile sets that flag reject these settings outright with a 400, so every provider
serving them has to drop and warn rather than forward. It lives here rather than beside either
model because models/bedrock.py cannot import from models/anthropic.py without pulling the
anthropic SDK into the bedrock extra.
Default: ('temperature', 'top_p', 'top_k')
Concrete Anthropic code execution tool version to send for CodeExecutionTool.
Type: TypeAlias Default: Literal['20250825', '20260120']
Maps unified thinking values to Anthropic budget_tokens for non-adaptive models.
Type: dict[ThinkingLevel, int] Default: {True: 10000, 'minimal': 1024, 'low': 2048, 'medium': 10000, 'high': 16384, 'xhigh': 32768}
Effort values Anthropic accepts at output_config.effort.
Type: TypeAlias Default: Literal['low', 'medium', 'high', 'xhigh', 'max']
Maps unified thinking effort levels to Anthropic output_config.effort.
xhigh maps to 'max' by default; callers that target a model with
anthropic_supports_xhigh_effort should pass supports_xhigh=True to
resolve_anthropic_effort
to preserve xhigh instead of downshifting.
Type: dict[ThinkingEffort, AnthropicEffort] Default: {'minimal': 'low', 'low': 'low', 'medium': 'medium', 'high': 'high', 'xhigh': 'max'}
Bases: ModelProfile
Profile for models used with GoogleModel.
ALL FIELDS MUST BE google_ PREFIXED SO YOU CAN MERGE THEM WITH OTHER MODELS.
Whether the model supports combining function declarations with native tools and response_schema. Default: False.
Gemini 3+ supports all tool combinations:
- function_declarations + native_tools
- output_tools (function declarations) + native_tools
- response_schema (NativeOutput) + function_declarations See https://ai.google.dev/gemini-api/docs/tool-combination
Type: bool
Whether the model accepts the include_server_side_tool_invocations tool-config field. Default: False.
When enabled, Gemini emits explicit tool_call/tool_response parts for server-side
native tools (Google Search, URL Context, File Search) that we round-trip through
NativeToolCallPart /
NativeToolReturnPart. Pre-Gemini-3 models
reject the field with 'Tool call context circulation is not enabled'.
This is a Gemini Developer API (ML Dev) only parameter: the google-genai SDK’s Vertex
converter raises ValueError when the field is set, so GoogleModel skips it for
Google Cloud (Vertex) even on Gemini 3+ models.
Distinct from google_supports_tool_combination
even though both currently flip on for Gemini 3+ — the former gates the SDK request
field, the latter gates which combinations of native / function / output tools are
allowed in the same request.
Type: bool
MIME types supported in native FunctionResponseDict.parts. Default: ().
See https://ai.google.dev/gemini-api/docs/function-calling#multimodal-function-responses
Whether the model uses thinking_level (enum: LOW/MEDIUM/HIGH) instead of thinking_budget (int). Default: True.
Gemini 3+ models use thinking_level; older models (e.g. Gemini 2.5) use thinking_budget.
Type: bool
Whether the model accepts thinking_level='MINIMAL'. Default: True.
Derived from google_thinking_levels
when that is set. Sparse profiles without the level set fall back to this flag: when disabled,
unified thinking='minimal' and thinking=False resolve to thinking_level='LOW'.
See https://ai.google.dev/gemini-api/docs/thinking.
Type: bool
Thinking levels the model supports, from Google’s per-model thinking table. Default: unset.
Unset means the full GOOGLE_THINKING_LEVELS
scale is assumed. Unified thinking efforts snap to the nearest supported level.
See https://ai.google.dev/gemini-api/docs/thinking.
Type: frozenset[GoogleThinkingLevel]
Whether the model supports Gemini’s VALIDATED function-calling mode. Default: False.
VALIDATED is Gemini’s equivalent of the cross-provider strict tool flag (like OpenAI/Anthropic
strict tool calling): it behaves like AUTO but the API enforces that the model adheres to the
declared function schema. Issue reports also observe that it mitigates the function-name hallucination
some Gemini models exhibit (an observed effect, not a documented guarantee). When the flag is set,
GoogleModel upgrades AUTO to VALIDATED by default (every schema is VALIDATED-compatible — no
rewrites), and a caller opts out per tool with ToolDefinition.strict=False. Because Gemini’s mode is
request-wide, any function or output tool with strict=False keeps the whole request on AUTO.
See https://ai.google.dev/gemini-api/docs/function-calling#function_calling_config.
Type: bool
Whether Google Search grounding is billed once per grounded prompt rather than per search query. Default: False.
Gemini 2.5 and older bill a request once, however many queries it ran, and only when it returned a web source;
Gemini 3+ bills each unique search query. This decides the web_searches count on
RequestUsage.
See https://ai.google.dev/gemini-api/docs/google-search#pricing.
Type: bool
Bases: JsonSchemaTransformer
Transforms the JSON Schema from Pydantic to be suitable for Gemini.
Gemini supports a subset of OpenAPI v3.0.3.
Bases: GoogleJsonSchemaTransformer
Transforms the JSON Schema from Pydantic into the OpenAPI v3.0.3 subset Gemini’s Schema accepts.
A function declaration carries its parameters as either parametersJsonSchema (full JSON Schema,
which GoogleModel sends) or parameters (an
OpenAPI v3.0.3 subset) —
the two are mutually exclusive. The Live API only implements parameters, so
GoogleRealtimeModel needs this narrower form,
where a union is anyOf, an enum is a list of strings, and there are no $refs to resolve.
def google_model_profile(model_name: str) -> ModelProfile | None
Get the model profile for a Google model.
def google_realtime_model_profile(model_name: str) -> RealtimeModelProfile
Get the realtime model profile for a Gemini Live model.
RealtimeModelProfile
Native Gemini thinking_level values.
Type: TypeAlias Default: Literal['MINIMAL', 'LOW', 'MEDIUM', 'HIGH']
The full thinking-level scale, cheapest first. The resolver’s order map derives from this.
Type: tuple[GoogleThinkingLevel, …] Default: ('MINIMAL', 'LOW', 'MEDIUM', 'HIGH')
The full thinking-level scale as a set.
Type: frozenset[GoogleThinkingLevel] Default: frozenset(GOOGLE_THINKING_LEVEL_SCALE)
def meta_model_profile(model_name: str) -> ModelProfile | None
Get the model profile for a Meta model.
def amazon_model_profile(model_name: str) -> ModelProfile | None
Get the model profile for an Amazon model.
def deepseek_model_profile(model_name: str) -> ModelProfile | None
Get the model profile for a DeepSeek model.
Bases: ModelProfile
Profile for Grok models (used with XaiProvider and various OpenAI-compatible providers).
ALL FIELDS MUST BE grok_ PREFIXED SO YOU CAN MERGE THEM WITH OTHER MODELS.
Whether the model supports builtin tools (web_search, x_search, code_execution, mcp). Default: False.
Type: bool
Deprecated: use supports_forced_tool_choice instead.
Translated (with a deprecation warning) whenever profiles are merged.
Type: bool
Native reasoning_effort values supported by the Grok model. Default: empty (frozenset()).
Type: frozenset[GrokReasoningEffort]
def grok_model_profile(model_name: str) -> ModelProfile | None
Get the model profile for a Grok model.
def grok_realtime_model_profile(model_name: str) -> RealtimeModelProfile
Get the realtime model profile for an xAI Grok Voice model.
RealtimeModelProfile
Native xAI reasoning_effort values.
Type: TypeAlias Default: Literal['none', 'low', 'medium', 'high']
def mistral_model_profile(model_name: str) -> ModelProfile | None
Get the model profile for a Mistral model.
def qwen_model_profile(model_name: str) -> ModelProfile | None
Get the model profile for a Qwen model.
Bases: ModelProfile
Profile for models used with GroqModel.
ALL FIELDS MUST BE groq_ PREFIXED SO YOU CAN MERGE THEM WITH OTHER MODELS.
Whether the model always has the web search built-in tool available. Default: False.
Type: bool
Whether thinking=False truly disables reasoning via reasoning_effort='none'. Default: False.
Only the qwen3 family supports this; other Groq reasoning models can at most suppress reasoning
output via reasoning_format='hidden' while still reasoning internally.
Type: bool
Whether the model accepts graded reasoning_effort values (low/medium/high). Default: False.
Only the gpt-oss family supports this; unified thinking levels map to those values via
GROQ_GPT_OSS_REASONING_EFFORT_MAP.
The qwen3 family instead only accepts none/default (see groq_supports_reasoning_disable).
Type: bool
def groq_model_profile(model_name: str) -> ModelProfile
Get the model profile for a Groq model.
Maps unified thinking effort levels to the graded reasoning_effort values the gpt-oss family accepts.
gpt-oss only accepts low/medium/high (not none/default), so minimal folds into low and
xhigh into high. thinking=True (bare enable) maps to medium in GroqModel, mirroring the neutral
default other providers use. See the Groq docs.
Type: dict[ThinkingEffort, Literal[‘low’, ‘medium’, ‘high’]] Default: {'minimal': 'low', 'low': 'low', 'medium': 'medium', 'high': 'high', 'xhigh': 'high'}
Bases: ModelProfile
Profile for Z.AI (Zhipu AI) GLM models.
Whether the model accepts a per-request reasoning_effort level (GLM-5.2 and GLM-5.3).
Type: bool
Substitutions applied to the unified thinking effort level before it is sent as reasoning_effort.
Levels not in the mapping are forwarded unchanged. GLM-5.2 accepts all unified levels, so it needs no
mapping; GLM-5.3 only accepts low/high/max (per the Z.AI docs and the error message returned when
disabling thinking on it), so its other unified levels are mapped.
def zai_model_profile(model_name: str) -> ModelProfile | None
The model profile for ZAI (Zhipu AI) GLM models, matched by Z.AI’s native glm-* ids.
Marks thinking-capable models (glm-5, glm-4.7, glm-4.6, glm-4.5) via supports_thinking=True.
This includes the glm-4.6v and glm-4.5v vision models, which also support thinking mode per the
Z.AI docs. GLM-5.2 and GLM-5.3 additionally accept a per-request reasoning effort level, flagged via
zai_supports_reasoning_effort=True. GLM-5.3 always reasons and cannot disable thinking, flagged via
thinking_always_enabled=True.
The provider-specific request/response shape (e.g. the reasoning_content field used by Z.AI’s API)
is configured in ZaiProvider.model_profile() rather than here. Providers that serve GLM models under
a different id scheme (e.g. Cerebras’s zai-glm-*, which doesn’t match the glm-* prefixes above)
configure thinking support in their own model_profile().
Bases: ModelProfile
Profile for a decision model: what the model can be asked.
These are facts about the model behind the URL, not the class that talks to it, so they are set by the provider
for the model name, or by profile=. A key left out falls back to the class’s
max_choice_options and
max_score_levels.
ALL FIELDS MUST BE decision_ PREFIXED SO YOU CAN MERGE THEM WITH OTHER MODELS.
The most options the model accepts in one pick-one question, or None for no limit.
A pick-one field with more options, or more routes than this on the route question, is a
UserError before a request is sent.
The most levels the model accepts in one rubric, or None for no limit.
Whole numbers from 0 with more levels than this are not a rubric, so a field of them is asked as a pick-one
instead, and counts against decision_max_choice_options.
def decision_model_profile(model_name: str) -> ModelProfile
Get the model profile for a decision model.
A decision model answers typed questions about a state; it does not generate text, call tools, or read anything
but text. Tool-mode structured output is how a decision model fills an output_type, and it rides on
supports_tools, so that stays on. A system prompt anywhere in the history is part of what the model judges, so
it needs no wrapping. Every other capability flag is off, and what no flag covers, such as a file in a prompt,
the model refuses itself.