Skip to main content

Why this is a common failure point

Every team tests its agent differently, so “our agent is good” means nothing outside the team. Comparing two models, two voice providers or two versions of your own agent needs the same calls, the same conditions and the same grading every time. Otherwise a better score may just be an easier test.

How Roark tests it

Live Bench is a fixed battery of calls with standardised flows, conditions and metrics. Nothing about it can be configured, so a score means the same thing across agents and against Roark’s published baselines. You pick the agent to put on the line.

Setting it up

The suite runs four batteries:

What it measures

The template seeds these, and you can add or remove metrics in Advanced before running.

What to look for

  • The latency bottleneck. It names the stage that owned most of each slow turn, so you know whether to change the transcriber, the model or the voice.
  • Compare like with like. Live Bench is built for A/B comparisons: run it on both versions of the agent and read the differences, not the absolute scores.

Over the API

Live Bench is not available over the API yet.