Observability & Tracing

Track every step of every agent run with detailed metrics, monitored continuously by Iris, our evaluation harness.

Agent Path: every route the compliance assistant took across 5,076 traces, from the model to each tool and sub-agent
99.9%
Less time to review 5,076 production traces: 7 minutes with Iris, against about 85 hours by hand
100%
Of production traces included in every analysis
2,110 → 6
Production errors traced to six named root causes in one analysis
Demo

Debug a Failed Run

Open a trace, see the route it took and why it failed, then confirm the fix held.

Per-Action Traces

Every agent action is recorded with its tool calls, retrievals and model calls on one timeline.

Continuous Monitoring

Iris checks new production traces for you, so failures surface without anyone having to ask.

Chat Support

Ask which agents are failing or what drove last week's cost, and get the numbers back.

Detailed Metrics

Cost, tokens and latency by agent, model, prompt version or day.

Analysis

The Number Nobody Had Looked Up

42.8%of runs had at least one error inside them

Every one of them returned an answer, so nobody noticed.

Errors in production agents are normal. Silent ones are the problem. Iris reads every trace, groups the errors by what actually went wrong, such as a denied permission, a knowledge base that was never attached or a sub-agent that failed, and takes you to the exact step where it happened. It tells you what to change, and the next run shows whether the fix held.

Iris's whole-estate analysis: 4,933 traces and 605 sessions, with the overall error rate at 42.8%
One question, covering every trace up to July 2, 2026.
Monitoring

Continuous Monitoring

Iris reviews new production traces on a schedule, groups failures by root cause, and reports each one with a proposed fix.

1

Runs on a Schedule

New production traces are reviewed automatically, without anyone having to ask.

2

Groups Failures by Root Cause

Errors are sorted by what went wrong, per agent and per error type.

3

Proposes the Fix

Each group comes with the likely cause and what to change.

Iris scheduled analysis, errors by agent: 2,110 errors across 5,076 traces, with the root cause for each agentThe same analysis by root cause: six causes, each with a proposed fix

One scheduled analysis, by agent and by root cause

Ask Iris

Conversational Assistance on Your Traces

Ask a question, search by meaning, or break cost down the way you need.

Iris listing the five most expensive traces this month and explaining what drove the cost
Iris finding 14 runs where the assistant cited a rule that has since changed, by meaning rather than keywords
The Traces dashboard: totals for traces, cost, tokens and response time, traces, cost and response time per day, top traces by name and cost by trace name
How it works

From Trace to Fix