A Full Evaluation, Start to Finish
From a request in plain language to a calibrated judge, in eight steps.
Custom Rubrics
Define the criteria that matter for your domain, so a score means what your reviewers mean by it.
Regression Runs
Re-run the same graded cases after a change and see exactly what moved.
Continuous Evaluation
Evaluators score new production traces as they arrive, between test runs.
Judge Calibration
Score a sample blind, chosen at random or by the agent, until the judge agrees with your experts.
What Iris Does
It reads every run your agents produce and names exactly where they broke: a failed tool call, a knowledge source that came back empty, a sub-agent that errored or a context window that dropped the evidence. What matters becomes a graded set the next version has to pass.
- Shows Its WorkEvery tool call and reasoning step is one click away.
- Asks Before It Changes AnythingNothing is written until you say yes.
- Scales Your Reviewers' JudgmentExpert-level scoring on every run, at a fraction of the cost.
Dataset Writer
Writes the test set, every category at once.
Score Analyst
Groups the failures, names the pattern and drafts the fix.
Trace Analyst
Pulls the trace behind any case you question.
Score It Blind. Then Check the Checker
Check the judge against your own experts before you trust its scores.



From Failed Cases to a New Prompt Version
See Why Cases Failed
Iris groups the low scores and names the pattern behind them.
Approve the Proposed Change
It drafts the prompt edit and waits for your yes.
Re-run the Same Cases
The same dataset runs against the new version, so the scores compare directly.
Pattern: three cases scored low, all on well-established topics the agent should know cold.
Hypothesis: the prompt lets it hedge on obscure topics, and it is hedging on familiar ones too.
Change: answer confidently on established policy; keep the hedge for the genuinely obscure. Want me to stage this?

Prompt v2 compared with v3The Evaluation Loop
Create Custom Evaluations
Create a custom evaluator and choose what evidence the judge can read.
Evaluate and Analyze
Run an LLM evaluation and analyze the results with Iris.
Score Live Traffic
Evaluators score new production traces as they arrive.
Calibrate the Judge
Score a sample yourself and improve the evaluator until it agrees with you.