Evaluation Harness
The Evaluation Studio overview: evaluation jobs, templates and scores
Iris's welcome screen, with five ways to start
Iris returning an analysis and a proposed test plan for the compliance assistant
The generated dataset, real cases ready to open and edit
Agent, dataset and template confirmed, then the run triggered
Score distribution, the failure cluster and the proposed change
Ten traces picked across the score range to score blind
The alignment result: mean absolute error 0.053, well calibrated