Why this is a common failure point
People call when something has gone wrong, so a large share of callers arrive frustrated, anxious or confused. Emotional callers speak differently (faster, louder, in fragments) and behave differently: they interrupt, repeat themselves and change direction. An agent that only ever met calm test callers mishears them, keeps the same cheerful tone with an angry caller, or loses the thread when a confused caller circles back.How Roark tests it
Emotion handling runs the same flow with a caller in each mood you pick, with the same goal and information every time. It tells you whether any mood makes your agent misunderstand the caller or fail the call, and measures how your agent sounded in response, so you can see whether it stayed calm. It is a property sweep: every value runs the same flow with everything else held the same, and the report tells you which values did worse than the rest and which differences are just chance.The values
Mood changes how the caller speaks (tone, pace, pitch) and how they behave. The caller’s goal and information stay the same.
What it measures
The template seeds these, and your flow’s own metrics run alongside them. Only checks can mark a value Worse on; the other measures explain why.What to look for
- Mirroring. An agent that gets louder or faster with an angry caller shows it in the loudness and pace deltas. Steady deltas across moods mean it held its tone.
- Confused and distracted callers repeat themselves and change direction. Comprehension failures there are the most actionable finding this sweep produces.
- The vocal measures explain, the check decides. Only Comprehension Failure is a check, so only it can mark a mood Worse on.
Setting one up
Pick the flow to test, then the values to run it across. Every value starts selected; deselect the ones you don’t need. Each value runs your flow again, so the values you keep are the cost of the run. Compared against sets the baseline: a reference value the report pins first, such as silence for background noise. The template suggests the natural one, and you can change it or choose none. It doesn’t change the verdict: whether a value did worse is always decided against every other value combined. Iterations live in the sweep panel rather than in Advanced. Each value runs every attached test case once per iteration:Calls per value = test cases x iterationsEach value needs at least 5 calls to be compared with the others, and the Iterations field won’t go below the count that reaches it. It starts at 6 calls per value, so one call excluded after the run doesn’t drop a value out of the comparison. More calls let the report see smaller differences. For a sweep of 10 values:
Below 5 calls a value can never be flagged, however badly it does, which is why the minimum is enforced rather than suggested. Over the API, a sweep template defaults
iterationCount to 6 and refuses a plan with fewer than 5 calls per value with a 400.How the results are decided
The sweep’s table leads the run report. Its Compared with the rest column gives each value one result:
How “too large to be chance” is decided:
- Each value is compared with all the other values combined, one check at a time, never with a single other value. Two values that happen to land at opposite ends are not a finding on their own.
- A sweep makes many comparisons at once (every value on every check), and some will look different by luck alone. The report accounts for that: of all the values it flags in a run, at most 1 in 10 is expected to be a fluke.
- A value is flagged only when it does worse. Doing better than the rest is not a problem to fix.