Test

Evals

Evals are regression tests for your agents. Each evaluation scripts a conversation, runs it against the real agent — its actual prompt, knowledge and tools — and checks the answers, so you know a prompt or knowledge change didn't quietly break something that used to work.

Open Test → Evals and pick the agent or Agent Team to test. Evaluations are saved with that agent or team, so they sync across devices and teammates.

The Evals page

  • Summary — the pass rate of graded tests, with Passed, Failed and Not graded counts, the average time per test and when you last ran them.
  • Evaluations list — filter by Failed, Passed, Not graded or Not run, or search. Each row shows its name, steps and checks; run it on its own, open the last result to read the conversation and why it passed or failed, edit it, or delete it.
  • Run all — runs every evaluation for the selected agent or team.
  • Create evaluation — opens the editor.
  • Generate tests — drafts tests from the agent's own hours, services, FAQs and knowledge. They wait for your review; nothing is added until you pick them. Check each expected answer first — a wrong expectation fails forever.
  • Find issues — the platform writes tricky questions for this agent (made-up facts, prompt leaks, off-topic and unsafe requests, tool use), runs them, and lists real problems with the reason each one is a problem. Pin as test turns a finding into a permanent evaluation using the exact question that broke it.

Building an evaluation

The editor has four parts.

Details — a name and a description. Leave the name blank and the first caller turn is used.

Evaluator — the model that grades AI judge checks. Exact-phrase checks don't use a model.

Conversation — choose Agent or Agent Team and which one. The agent's prompt is shown read-only for reference (the knowledge base, flow and memory are added when it runs). Then build the script from the bar at the bottom:

  • Caller turn — what the caller says. The agent answers it.
  • Agent reply — a fixed reply put into the conversation as if the agent had said it. The model isn't asked that turn, which lets you set up a later moment in a conversation reliably.
  • Check — grades the agent's answer to the caller turn right above it:
  • AI judge — pass if the reply meets your rule, e.g. "offers Tuesday or Wednesday and asks which suits the caller". Catches tone and correctness, not just words.
  • Exact phrase — pass if the reply contains your text. Fast and free.
  • Optionally, tools the agent must call while answering that turn.

Checks can go anywhere in the conversation, not just at the end. Reorder or delete steps with the controls on each one; the editor warns about a check that doesn't follow a caller turn.

Whole-conversation checks (optional) — tools that must be called at some point, and a time limit for the whole test.

Variables

Use {{name}} in a caller turn or in the agent's prompt, then set a test value under Variables in the Test runs panel. Variables found in the prompt and the script are listed for you; you can also add your own.

Running it

Click Test (or press ⌘↵). It runs the draft as it is — you don't need to save first. The Test runs panel shows each run's conversation with every check's verdict under the reply it graded, how long it took, the tools called, and for teams which member answered each turn. Save (⌘S) stores the evaluation; if you switched it to another agent or team, saving moves it there.

Passed, failed, not graded

  • Passed — every check that ran passed.
  • Failed — at least one check genuinely failed: the agent said the wrong thing, missed a required tool, or was too slow.
  • Not graded — the test never got a verdict: no LLM key, budget spent, the judge was unavailable, or the request failed. This is not a failure, and it isn't counted in the pass rate — it means nothing was verified either way. Fix the cause and run it again.

Branches

The same evaluations run before you promote an agent branch to live, against the branch's own settings, so you can see what would change before it goes out.

Next steps