Skip to main content
POST
JavaScript

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <token>, where <token> is your auth token.

Body

application/json

A run plan to configure, or a built-in template to save as one.

name
string
required

Name of the run plan

Minimum string length: 1
Example:

"My Run Plan"

direction
enum<string>
required

Direction of the simulation (INBOUND or OUTBOUND)

Available options:
INBOUND,
OUTBOUND
Example:

"INBOUND"

maxSimulationDurationSeconds
integer
required

Maximum duration in seconds for each simulation

Required range: x >= 1
Example:

300

agentEndpoints
object[]
required

Agent endpoints to include in this run plan

Minimum array length: 1
description
string

Description of the run plan

Example:

"A run plan for testing inbound calls"

iterationCount
integer
default:1

Number of iterations to run for each test case (1-10000)

Required range: 1 <= x <= 10000
Example:

1

maxConcurrentJobs
integer
default:5

Maximum number of concurrent simulation jobs

Required range: x >= 1
Example:

5

silenceTimeoutSeconds
integer
default:30

Timeout in seconds for silence detection

Required range: x >= 1
Example:

30

maxNoResponseRetries
integer
default:0

How many more times to run a test case when the agent under test never responds: it never speaks on a call or never replies in a chat (0-10). 0 turns retries off. Failed checks and failures on Roark’s side are never retried.

Each retry is a separate attempt, billed like any other, so a plan retrying N times can place up to N + 1 calls per test case. Every silent attempt stays on the run with its own call; the run settles once each test case has a final attempt, and the agent never spoke verdict is judged on each test case’s last attempt.

Required range: 0 <= x <= 10
Example:

2

noResponseRetryBackoffSeconds
integer
default:90

Seconds a retry waits before it dials (30-600). Only used when maxNoResponseRetries is above 0.

Required range: 30 <= x <= 600
Example:

90

endCallPhrases
string[]

Phrases that trigger end of call. Empty array disables the feature.

Example:
endCallReasons
string[]

Semantic conditions that trigger end of call. The LLM evaluates the conversation against these conditions. Empty array disables the feature.

Example:
executionMode
enum<string>
default:PARALLEL

Execution mode (PARALLEL or SEQUENTIAL)

Available options:
PARALLEL,
SEQUENTIAL_SAME_RUN_PLAN,
SEQUENTIAL_PROJECT
Example:

"PARALLEL"

scenarios
object[]
deprecated

Deprecated: use flows instead. Scenarios to include in this run plan. The same scenario ID can appear multiple times with different variables.

Minimum array length: 1
flows
object[]

Customer flows to include in this run plan. The same flow can appear more than once with a different persona override, different variables, or different overrides: attaching it once per value of one property is how you compare that property without a template.

Minimum array length: 1
personas
object[]

Personas to include in this run plan. Required with scenarios; ignored with flows, where each variant carries its own persona.

Minimum array length: 1
metrics
object[]

Metric definitions to include in this run plan. Reference each by id (UUID) or slug.

Optional when the attached flows carry the grading: metrics a flow declares itself (with includeFlowMetrics), or the Agent Expectations and Keypad Entry metrics a run adds for flows with expectations or expected keypad entries (with includeAutomaticMetrics). A plan with nothing to grade is rejected with a 400.

Minimum array length: 1
includeFlowMetrics
boolean
default:true

Also collect each attached flow's own metrics, on top of the metrics named here.

Default true, which is what you want when you brought your own flows and their graders. Set false for a run whose metric list is meant to be exhaustive: a template like Load Testing or Voicemail deliberately grades a narrow set, and inheriting every flow metric on top multiplies analysis cost across the volume without adding signal.

GET /v1/simulation/template returns the value each template expects.

Example:

true

includeAutomaticMetrics
boolean
default:true

Let the run add metrics by itself off the attached flows, on top of the metrics named here.

Two attach this way today: Agent Expectations wherever an attached flow has agent expectations written on it, and Keypad Entry wherever one has steps where the agent is expected to press keys. Both grade something authored on the flow that nothing else measures, which is why it is on by default.

Set false when the metrics list is meant to be exhaustive: a plan testing only whether the caller can complete the flow may not want the agent graded on its expectations as well. False also pins the plan against any automatic metric Roark adds later.

Example:

true

comparisonProperty
enum<string> | null

The property this run plan investigates: the one thing its arms differ by.

Set it and the run report compares the arms on that property, so a run answers "what did background noise cost" rather than just "what did each arm score". Every value is a field already recorded on each call, so the report can label an arm CRYING_BABY rather than repeating a flow variant's title.

Omit it and the report still compares when it can: it detects which property varies across the arms. Setting it is what tells the written summary what you were trying to find out, which detection cannot infer.

Available options:
ACCENT,
AGE,
BACKGROUND_NOISE,
BACKGROUND_NOISE_VOLUME,
BASE_EMOTION,
CONFIRMATION_STYLE,
GENDER,
INTENT_CLARITY,
LANGUAGE,
INTERRUPTION,
MEMORY_RELIABILITY,
RESPONSE_TIMING,
SPEECH_CLARITY,
SPEECH_PACE
Example:

"BACKGROUND_NOISE"

comparisonBaseline
string | null

The reference value of comparisonProperty, for example NONE for BACKGROUND_NOISE or NORMAL for SPEECH_PACE: shown first in the results. Must be a value that property can take. Whether a value did significantly worse does not depend on it: that is decided against every other value combined (see sweepAttribution).

Stored rather than assumed, so the report can say "compared against US accent" instead of implying Roark decided which value is normal. Omit it and the property's own norm is used, as the dashboard prefills it, or none when your comparisonValues leave the norm out. GENDER has no norm, so choose the one you are testing against.

Example:

"NONE"

comparisonValues
(string | object)[]

The arms to run, for a plan that sweeps comparisonProperty. This is what the plan costs: the flow is attached once per arm, so ten arms is ten times the calls of one.

Attach each flow once, as you would without a comparison: the plan builds the arms, running the happy path or edge cases you selected under every arm. Built arms need at least 5 calls per arm (iterationCount times the test cases per arm), or the plan is refused with 400. Flows that all carry overrides on comparisonProperty already are the arms and are kept as you wrote them; a mix of flows with and without one is refused.

Each entry is one arm. A bare value runs it plain: "CITY". An object runs the value with something pinned on that arm only, such as a noise level per bed: { "value": "OFFICE", "backgroundNoiseVolume": 0.6 } plays OFFICE at 60% while the other beds keep the default. List a value more than once with different pins to run it as several arms: DRIVING at 0.7 and DRIVING at 1 are two arms, reported as Driving (70% noise) and Driving (100% noise), and "DRIVING" beside them keeps the plain arm too. The sweep still varies one property; what an arm pins is part of "everything else" for that arm only, so the report still compares the arms on comparisonProperty.

Omit it to run every value the property has, plain, which for ACCENT is more than twenty. A comparisonBaseline outside the values listed is rejected, because it would anchor every difference to an arm the run never made. A value the property cannot take, a pin the sweep cannot account for, or the same arm listed twice is rejected with 400.

Not stored as a field: the arms are the values. Reading the plan back returns them as its flow attachments, each with its pins as overrides.

Minimum array length: 1

A value of comparisonProperty, run as its plain arm.

Minimum string length: 1
Example:
enrichWithLiveConversation
boolean
default:false

Merge the customer's own recording of the real call into each simulation, so metrics can be scored against the live leg as well as the simulated one. This is the API equivalent of the dashboard's live-enrichment toggle.

With this on, the run provisions a phone number and holds each call open for up to 15 minutes waiting for a matching call to be posted to POST /v1/call. A call matches on the provisioned number (roarkPhoneNumber on the job) with a start time inside the simulation window. If nothing arrives, the simulation still completes and any LIVE-sourced metric produces no value.

Required by any metric whose requiresLiveConversation is true: without it that metric is silently skipped.

Example:

false

autoRun
boolean
default:false
deprecated

Deprecated: use POST /v1/simulation/run, which starts a run and accepts runtime variables as well. This flag runs the plan with only the values pinned on it.

Example:

false

Response

The created run plan

data
object
required

Response when creating a run plan, optionally including a triggered job