| Production context | Browser, agents, services, databases, logs, metrics, and infrastructure | AI and application traces, logs, and OpenTelemetry spans |
| Full-stack observability suite | Infrastructure monitoring, service maps, logs, metrics, traces, and AI | No |
| Browser session replay | Browser tracing and session replay alongside the agent trace | No |
| First-class feature flags | OpenFeature/OFREP flags, targeting, and controlled rollout | No |
| Evaluation workflow | Pydantic Evals: the same evaluators online and offline | Experiments, playgrounds, CI/CD, and online scoring |
| Score pricing | No separate score meter; normal record pricing applies | $1.50 per 1,000 scores after 50K/month on Pro |
| Human review | Annotation queues on production runs and evals | Human-review scores and assigned trace review |
| From failure to change | Trace-backed optimizer and managed agent configuration | Prompt, scorer, dataset, and environment workflows |
| Managed agent configuration | Prompts, agent specs, tools, skills, versioning, targeting, and rollout | No |
| Controlled rollout | Immutable versions, labels, targeting, weighted rollout, and feature flags | Prompt, dataset, and parameter environments |
| Investigation workflow | Agent trace investigator, PostgreSQL-compatible SQL, and MCP across full telemetry | Logs, trace views, SQL, and MCP for Braintrust data |
| Deployment options | Cloud, Dedicated, or the same product self-hosted on Kubernetes | Cloud or an Enterprise self-hosted data plane |