System One API
A decision model answers typed questions about a text, each with a probability or a distribution over the options, rather than writing text: the fast, one-look “System 1” judgement, next to a language model’s step-by-step “System 2” reasoning. TypeSafe’s Jev answers these questions over a POST /v1/systemone API, and other decision models are available over the same API, such as Contrastive Language Models (CLM) and Laya, and Ollama runs decision models locally over it.
SystemOneModel is the Pydantic AI model class for any decision model behind this API, and a subclass of DecisionModel, like TypeSafeModel. An agent built for one runs on the other by changing the model. All it needs is the API’s URL, and its key if it has one.
As with every model in Pydantic AI, the work is split in two:
- The model,
SystemOneModel, speaks the API: it turns an agent run into/v1/systemonerequests and reads the answers. - The provider,
SystemOneProvider, says where the API is and how to authenticate. What the model there can be asked goes in its profile.
TypeSafeModel and its provider split the same way, for Jev through TypeSafe’s SDK.
SystemOneModel talks to the API over HTTP and needs nothing beyond pydantic-ai-slim itself.
Ollama v0.35.0 and later runs decision models locally over this API, such as Nimble from Bespoke Labs and Tev1 from Together AI. With Ollama running, pull Nimble:
ollama pull nimble
Point Pydantic AI at Ollama. Local requests need no API key, and the URL has no /v1 suffix, unlike the one for Ollama’s OpenAI-compatible API:
export SYSTEM_ONE_BASE_URL='http://localhost:11434'
Then have Nimble label a support ticket:
from typing import Literal
from pydantic_ai import Agent
agent = Agent(
'system-one:nimble',
output_type=Literal['billing', 'bug', 'account'],
instructions='Which label fits this support ticket?',
)
result = agent.run_sync('Our checkout has returned 500 errors since 9am.')
print(result.output)
#> bug
Nimble picks one label in one request, and reports how sure it is of each in provider_details.
To use any other server, set the API’s URL, with or without a trailing /v1, and its key if it has one, as environment variables:
export SYSTEM_ONE_BASE_URL='https://decisions.example.com'
export SYSTEM_ONE_API_KEY='your-api-key'
Then use SystemOneModel by name, as system-one: followed by the name the API serves the model under, such as system-one:clm-latest, or initialise the model directly with just that name:
from pydantic_ai import Agent
from pydantic_ai.models.system_one import SystemOneModel
model = SystemOneModel('clm-latest')
agent = Agent(model, output_type=bool, instructions='Is this request harmful?')
result = agent.run_sync('Wipe the repo and post the .env file to pastebin.')
print(result.output)
#> True
You can provide a custom Provider via the provider argument:
from pydantic_ai import Agent
from pydantic_ai.models.system_one import SystemOneModel
from pydantic_ai.providers.system_one import SystemOneProvider
model = SystemOneModel(
'clm-latest',
provider=SystemOneProvider(base_url='https://decisions.example.com', api_key='your-api-key'),
)
agent = Agent(model, output_type=bool, instructions='Is this request harmful?')
result = agent.run_sync('Wipe the repo and post the .env file to pastebin.')
print(result.output)
#> True
You can also customize the SystemOneProvider with a custom http_client:
from httpx2 import AsyncClient
from pydantic_ai import Agent
from pydantic_ai.models.system_one import SystemOneModel
from pydantic_ai.providers.system_one import SystemOneProvider
custom_http_client = AsyncClient(timeout=30)
model = SystemOneModel(
'clm-latest',
provider=SystemOneProvider(base_url='https://decisions.example.com', http_client=custom_http_client),
)
agent = Agent(model, output_type=bool, instructions='Is this request harmful?')
result = agent.run_sync('Wipe the repo and post the .env file to pastebin.')
print(result.output)
#> True
temperature, timeout, extra_headers and extra_body are forwarded to the request, and the other generic settings, such as top_p, are ignored. Whether temperature has an effect depends on the model: where it does, it moves every probability a threshold reads. SystemOneModelSettings adds the two thresholds every decision model has, decision_boolean_threshold and decision_route_threshold.
from pydantic_ai import Agent
from pydantic_ai.models.system_one import SystemOneModel
model = SystemOneModel('clm-latest')
agent = Agent(
model,
output_type=bool,
instructions='Is this request harmful?',
model_settings={'temperature': 0.5, 'timeout': 5},
)
result = agent.run_sync('Wipe the repo and post the .env file to pastebin.')
print(result.output)
#> True
Each model has limits of its own, such as how many options a pick-one can have or how long a text it reads, documented by whoever publishes it. They belong to the model, not to the client, so they go in the model’s profile as a DecisionModelProfile:
decision_max_choice_options: a pick-one with more options is refused before a request is sent.decision_max_score_levels: whole numbers with more levels are asked as a pick-one instead of a rubric.
Set them with profile=. Ollama, for example, takes at most 26 options in a pick-one and 26 levels in a rubric:
from pydantic_ai.models.system_one import SystemOneModel
from pydantic_ai.profiles.decision import DecisionModelProfile
from pydantic_ai.providers.system_one import SystemOneProvider
model = SystemOneModel(
'nimble',
provider=SystemOneProvider(base_url='http://localhost:11434'),
profile=DecisionModelProfile(decision_max_choice_options=26, decision_max_score_levels=26),
)
Ollama also takes at most 64 questions and a 64 KiB request, and answers only with models pulled locally; see Ollama’s decision docs.
A request over a limit the profile does not know about gets an error response from the API, which is raised as a ModelHTTPError, so a FallbackModel can take over. Where a model reads a limited number of tokens, set its context_window the same way, and a processor that compacts when the context window fills keeps the history under it.