Create a run plan
Creates a new simulation run plan.
To run a simulation, use POST /v1/simulation/run instead: it starts a run from a plan or from an inline configuration, and takes runtime variables. Create a plan here when you want a reusable, named one to run later.
Send template instead of a full configuration to save one of the built-in templates as a
plan. It takes the same fields as the template variant of POST /v1/simulation/run, builds the
same plan, and never starts it.
To compare one property, attach the flow once and send comparisonProperty with the
comparisonValues to run: the plan attaches the flow once per value.
Authorizations
Bearer authentication header of the form Bearer <token>, where <token> is your auth token.
Body
- New plan
- From template
A run plan to configure, or a built-in template to save as one.
Name of the run plan
1"My Run Plan"
Direction of the simulation (INBOUND or OUTBOUND)
INBOUND, OUTBOUND "INBOUND"
Maximum duration in seconds for each simulation
x >= 1300
Agent endpoints to include in this run plan
1Description of the run plan
"A run plan for testing inbound calls"
Number of iterations to run for each test case (1-10000)
1 <= x <= 100001
Maximum number of concurrent simulation jobs
x >= 15
Timeout in seconds for silence detection
x >= 130
How many more times to run a test case when the agent under test never responds: it never speaks on a call or never replies in a chat (0-10). 0 turns retries off. Failed checks and failures on Roark’s side are never retried.
Each retry is a separate attempt, billed like any other, so a plan retrying N times can place up to N + 1 calls per test case. Every silent attempt stays on the run with its own call; the run settles once each test case has a final attempt, and the agent never spoke verdict is judged on each test case’s last attempt.
0 <= x <= 102
Seconds a retry waits before it dials (30-600). Only used when maxNoResponseRetries is above 0.
30 <= x <= 60090
Phrases that trigger end of call. Empty array disables the feature.
Semantic conditions that trigger end of call. The LLM evaluates the conversation against these conditions. Empty array disables the feature.
Execution mode (PARALLEL or SEQUENTIAL)
PARALLEL, SEQUENTIAL_SAME_RUN_PLAN, SEQUENTIAL_PROJECT "PARALLEL"
Deprecated: use flows instead. Scenarios to include in this run plan. The same scenario ID can appear multiple times with different variables.
1Customer flows to include in this run plan. The same flow can appear more than once with a different persona override, different variables, or different overrides: attaching it once per value of one property is how you compare that property without a template.
1Personas to include in this run plan. Required with scenarios; ignored with flows, where each variant carries its own persona.
1Metric definitions to include in this run plan. Reference each by id (UUID) or slug.
Optional when the attached flows carry the grading: metrics a flow declares itself
(with includeFlowMetrics), or the Agent Expectations and Keypad Entry metrics a run
adds for flows with expectations or expected keypad entries (with
includeAutomaticMetrics). A plan with nothing to grade is rejected with a 400.
1Also collect each attached flow's own metrics, on top of the metrics named here.
Default true, which is what you want when you brought your own flows and their graders. Set false for a run whose metric list is meant to be exhaustive: a template like Load Testing or Voicemail deliberately grades a narrow set, and inheriting every flow metric on top multiplies analysis cost across the volume without adding signal.
GET /v1/simulation/template returns the value each template expects.
true
Let the run add metrics by itself off the attached flows, on top of the metrics named here.
Two attach this way today: Agent Expectations wherever an attached flow has agent expectations written on it, and Keypad Entry wherever one has steps where the agent is expected to press keys. Both grade something authored on the flow that nothing else measures, which is why it is on by default.
Set false when the metrics list is meant to be exhaustive: a plan testing only whether the caller
can complete the flow may not want the agent graded on its expectations as well. False also pins
the plan against any automatic metric Roark adds later.
true
The property this run plan investigates: the one thing its arms differ by.
Set it and the run report compares the arms on that property, so a run answers "what did
background noise cost" rather than just "what did each arm score". Every value is a field
already recorded on each call, so the report can label an arm CRYING_BABY rather than
repeating a flow variant's title.
Omit it and the report still compares when it can: it detects which property varies across the arms. Setting it is what tells the written summary what you were trying to find out, which detection cannot infer.
ACCENT, AGE, BACKGROUND_NOISE, BACKGROUND_NOISE_VOLUME, BASE_EMOTION, CONFIRMATION_STYLE, GENDER, INTENT_CLARITY, LANGUAGE, INTERRUPTION, MEMORY_RELIABILITY, RESPONSE_TIMING, SPEECH_CLARITY, SPEECH_PACE "BACKGROUND_NOISE"
The reference value of comparisonProperty, for example NONE for BACKGROUND_NOISE or
NORMAL for SPEECH_PACE: shown first in the results. Must be a value that property can take.
Whether a value did significantly worse does not depend on it: that is decided against every
other value combined (see sweepAttribution).
Stored rather than assumed, so the report can say "compared against US accent" instead of
implying Roark decided which value is normal. Omit it and the property's own norm is used,
as the dashboard prefills it, or none when your comparisonValues leave the norm out.
GENDER has no norm, so choose the one you are testing against.
"NONE"
The arms to run, for a plan that sweeps comparisonProperty. This is what the plan costs: the
flow is attached once per arm, so ten arms is ten times the calls of one.
Attach each flow once, as you would without a comparison: the plan builds the arms, running
the happy path or edge cases you selected under every arm. Built arms need at least 5 calls
per arm (iterationCount times the test cases per arm), or the plan is refused with 400.
Flows that all carry overrides on comparisonProperty already are the arms and are kept as
you wrote them; a mix of flows with and without one is refused.
Each entry is one arm. A bare value runs it plain: "CITY". An object runs the value with
something pinned on that arm only, such as a noise level per bed: { "value": "OFFICE", "backgroundNoiseVolume": 0.6 } plays OFFICE at 60% while the other beds keep the default.
List a value more than once with different pins to run it as several arms: DRIVING at 0.7 and
DRIVING at 1 are two arms, reported as Driving (70% noise) and Driving (100% noise), and
"DRIVING" beside them keeps the plain arm too. The sweep still varies one property; what an
arm pins is part of "everything else" for that arm only, so the report still compares the arms
on comparisonProperty.
Omit it to run every value the property has, plain, which for ACCENT is more than twenty.
A comparisonBaseline outside the values listed is rejected, because it would anchor every
difference to an arm the run never made. A value the property cannot take, a pin the sweep
cannot account for, or the same arm listed twice is rejected with 400.
Not stored as a field: the arms are the values. Reading the plan back returns them as its
flow attachments, each with its pins as overrides.
1A value of comparisonProperty, run as its plain arm.
1Merge the customer's own recording of the real call into each simulation, so metrics can be scored against the live leg as well as the simulated one. This is the API equivalent of the dashboard's live-enrichment toggle.
With this on, the run provisions a phone number and holds each call open for up to 15 minutes
waiting for a matching call to be posted to POST /v1/call. A call matches on the provisioned
number (roarkPhoneNumber on the job) with a start time inside the simulation window. If
nothing arrives, the simulation still completes and any LIVE-sourced metric produces no value.
Required by any metric whose requiresLiveConversation is true: without it that metric is
silently skipped.
false
Deprecated: use POST /v1/simulation/run, which starts a run and accepts runtime variables as well. This flag runs the plan with only the values pinned on it.
false
Response
The created run plan
Response when creating a run plan, optionally including a triggered job