Skip to content

Set reliability targets for services and large language model providers

Set a measurable reliability goal for a service or large language model (LLM) provider, then watch its error budget and burn rate before the target is missed.

A service level objective (SLO) is a target for how reliable a service should be, such as “99.9% of checkout requests succeed over the last 30 days.” Logfire calls each SLO a reliability target. An SLO is an internal engineering target; a service level agreement (SLA) is an external contract.

You’ll find Service levels under Reliability in the project sidebar. The page lists every target in the project. Search by target name, service, provider, or description; filter by status; or group by service, status, time window, source, or not at all. Service detail pages show the same targets on their Reliability page.

Choose what good looks like

A target combines:

  1. The total events or values to measure.
  2. The subset that counts as bad.
  3. The percentage that should be good over a rolling time window.

The resulting good-event ratio is the service level indicator (SLI). Enter any target percentage below 100%, or choose a preset from 99% to 99.99%. The default is 99.9%. Window presets are 1, 7, 28, 30, and 90 days.

Two derived values show whether the goal is at risk:

  • Error budget is the amount of unreliability the target allows. A 99.9% target allows 0.1% of included events to be bad.
  • Burn rate is how quickly the service is spending that budget. At 1x, one complete budget would be consumed over the target’s rolling window.

Start from a template

From a service detail page, select Reliability, then New target and choose:

  • Availability for records that did not error.
  • Latency for records below a duration threshold.
  • AI / LLM calls for calls without a provider-side failure.
  • AI quality (eval score) for records that reach an evaluation threshold. An evaluation (eval) is a repeatable test of model or agent behavior. This template uses gen_ai.evaluation.score.value, which Pydantic Evals can populate.
  • Custom SQL, using Structured Query Language conditions, when the preset conditions do not fit.

To manage targets as code, use the experimental logfire_slo resource of the Terraform and Pulumi providers. See Infrastructure as Code.

Choose records or metrics

Records measure rows such as requests, remote procedure calls (RPCs), background jobs, and LLM calls. Define the total population and bad subset with SQL boolean conditions. The setup wizard previews matching records before you save.

Metrics measure a count or other additive value, a gauge fraction, a cumulative counter, or the share of histogram observations on the good side of a threshold. For a histogram, choose Below is good for values such as request latency, or Above is good when higher values are better. Their burn-rate history appears after you save the target.

You can restrict either source to selected deployment environments. Set the target percentage and window, then choose the starting notification channels for its three generated alerts. You can change the channels of each alert later.

Let Logfire watch the burn rate

Each target generates three alerts, one for each burn-rate tier:

TierWindowsThresholdSeverity
Fast burn1 hour and 5 minutes14.4xpage
Medium burn6 hours and 30 minutes6xpage
Slow burn3 days and 6 hours1xticket

These tiers follow the multiwindow, multi-burn-rate method in the Google Site Reliability Engineering (SRE) Workbook. Checking two windows keeps a brief spike from firing an alert meant for a sustained regression.

A low target can make a tier unable to fire. The highest possible burn rate is reached when every event is bad: at a 90% target, that is 10x, which is below the fast burn threshold of 14.4x. The fast burn tier can fire only at a target of about 93.06% or higher, and the medium burn tier only at about 83.34% or higher.

Logfire keeps the alert for a tier that cannot fire, but does not evaluate it. The target detail page shows the tier as Disabled with Target too low to fire, and the alert page says that the reliability target is too low for the alert to fire. A disabled tier does not count as an active alert. If you raise the target enough, the same alert runs again, with its notification channels.

Route alerts from the target detail page. See Alerts for notification-channel setup.

Track the budget and investigate failures

The target detail page shows Current service level, Target, and Error budget remaining. Reliability history has two views: Error budget shows the remaining budget over time, and Burn rate shows the hourly burn rate. Both views start with the full rolling window of the target. To look at a shorter span, choose 24 hours, or 7 days when the window is longer than 7 days.

To investigate failures, drag across the chart to select a time range. Then select Open in SQL Workbench to query that range in the SQL Workbench: the failing records for a target that measures records, or the metric rows for a target that measures metrics. For a target that measures records, Open in Live view shows the same failing records in Live view. Without a selection, these links use the range that the chart shows. If a burn-rate alert has fired, its row has an Investigate link. Select it to open the chart at the time range around the most recent firing of the alert.

The How it’s measured panel records the scope and exact conditions.

Monitor an LLM provider

On the Providers tab of LLMs and Providers, select Set reliability targets for an observed provider. Choose an availability or latency target for that provider across every service that calls it. To limit an LLM target to one service, create it from that service’s Reliability page instead.

Provider targets default to 99% over 28 days. Availability targets treat server failures, rate limits, timeouts, and connection errors as bad. Recognized OpenAI and Anthropic client errors, such as invalid requests or credentials, are excluded because they do not measure provider availability.