---
title: AI will misbehave. Will you notice?
description: >-
  Engineers are shipping AI features fast, often without a dedicated AI
  specialist. See how to run production AI well with observability across the
  LLM layer and the whole system, evals before and after you ship, and failures
  that turn into tests.
date: '2026-06-25'
dateDisplay: 'Thu, Jun 25 · 3:30pm UTC (11:30am EDT)'
duration: 45 min
status: past
tags:
  - Logfire
speakers:
  - Chris Samiullah
canonical: 'https://pydantic.dev/events/intro-to-logfire'
---

> Markdown version of [AI will misbehave. Will you notice?](https://pydantic.dev/events/intro-to-logfire) — the canonical HTML page.
>
> Thu, Jun 25 · 3:30pm UTC (11:30am EDT) · 45 min · Speakers: [Chris Samiullah](https://pydantic.dev/authors/chris-samiullah.md)
>
> All events: [/events.md](https://pydantic.dev/events.md) · Site index: [/llms.txt](https://pydantic.dev/llms.txt)

---

# AI will misbehave. Will you notice?

Engineers are shipping AI features fast, often without a dedicated AI specialist. See how to run production AI well with observability across the LLM layer and the whole system, evals before and after you ship, and failures that turn into tests.

## What you'll learn

- Why production AI needs visibility at both the LLM layer and across the full system around it
- How traces, dashboards, and alerts surface cost, latency, and behavior across real traffic
- How offline and online evals measure whether answers are any good, before you ship and once you're live
- How a failure in production becomes a test, so the system improves as it runs

## About

Engineers are shipping AI features fast, often on lean teams without a dedicated AI specialist. The good news is AI is still engineering. The discipline that keeps software reliable is what keeps AI reliable too.

Running production AI well requires knowing what your AI is actually doing once it's live, and using that to keep making it better. Every feature you ship is one more thing that can fail quietly. To catch it before users do, AI observability should cover two levels: the LLM layer and the whole system around it. You need both because when something breaks, a tool that sees only one leaves you guessing about the other.

[Pydantic Logfire](https://pydantic.dev/logfire?utm_source=pydantic.dev&utm_medium=event_page&utm_campaign=intro-to-logfire) is AI observability for production: LLM calls, costs, prompts, tool calls, and eval scores, alongside every other step that produced the response, in one view. It's built on OpenTelemetry, so the model and the rest of your stack, in Python, TypeScript, Rust, or anything OTel-compatible, land in one place.

A 45-min session, with time for your questions at the end.

## Recording

Watch the recording on the event page: https://pydantic.dev/events/intro-to-logfire
