Datasets and experiments
Use the Logfire evaluation workspace to answer two questions: what should you test, and did the latest change help?
A dataset is a reusable collection of test cases. Each case has an input and can also have an expected output and metadata. An experiment is one run of your AI system over a dataset. Evaluators judge each result and record assertions or scores that Logfire can aggregate and compare.
- Curate cases. Keep cases in a hosted dataset that your team edits in Logfire, or define them with your eval code.
- Run an experiment. Use pydantic-evals to run the same task and evaluators over every case.
- Read the overview. Check completion, assertions, task errors, evaluator results, duration, tokens, and cost.
- Review case evidence. Inspect the input, output, evaluator result, and trace for failures or surprising scores.
- Compare iterations. Choose a baseline run and confirm whether a candidate improved the cases that matter.
The New dataset flow asks how you want to manage cases:
- Manage in Logfire creates a hosted dataset. Your team can add and edit cases in the web UI, define JSON schemas, import cases from code, and add production traces.
- Manage in code keeps the source of truth with your eval code. Experiment results still appear in Logfire. You can also sync a copy of the cases to Logfire when you want to browse or edit them there.
Use stable dataset names so successive runs appear together. If you plan to create a hosted copy, choose a hosted-compatible name: start with a letter or number and use only letters, numbers, dots, underscores, and hyphens.
The name is also the identity. Hosted names are unique within a project, and Logfire keys the Datasets list on them. A hosted dataset and a code-defined one sharing an exact, hosted-compatible name form a single entry carrying both the hosted cases and the experiment history. The choice above is therefore a starting point rather than a commitment: a dataset that begins in code can gain a hosted copy later under the same name.
- Manage datasets to find, create, import, and edit test cases.
- Run evals in code to execute a task and send results to Logfire.
- Review experiments to diagnose failures and compare a candidate with a baseline.
- Use the Datasets SDK to create, publish, and fetch hosted datasets from Python.
- Configuring Logfire during an eval sends its inputs, outputs, evaluator results, and traces to your project.
- Experiments run from code. The web UI helps you prepare datasets and analyze results; it does not execute the task under test.
- Code remains the source of truth for code-defined cases. Syncing imports new cases into the hosted copy and updates cases with matching names; edits in Logfire do not change your code.
- Every experiment is also trace data. Open a case’s trace when the result alone does not explain what the agent did.
Start with Manage datasets if you need cases. If you already have a dataset and results, go directly to Review experiments.