Live
What a live eval costs money with: the judge and a required budget.
Live evals need agents too.
Fields
judgeModelrequiredThe model that grades each case against its rubric (judge.v1 instructions, one call per case).
budgetBudgetrequiredThe run budget of every agent run and every judge run, subagents included. A run it stops is an error with budget_exhausted.
rubricreadonly string[]default []Criteria added after each case's own.