Threads AI

Live

What a live eval costs money with: the judge and a required budget.

Live evals need agents too.

Fields

judgeModelrequired

The model that grades each case against its rubric (judge.v1 instructions, one call per case).

budgetBudgetrequired

The run budget of every agent run and every judge run, subagents included. A run it stops is an error with budget_exhausted.

rubricreadonly string[]default []

Criteria added after each case's own.

Edit on GitHub

On this page