Appearance
Test runs
Available now In a vendor workspace, the portal's Test runs page playsa test suite into your workspace and checks what your app does with it. Every case gets a verdict, with what was expected, what your app did, and what the check rests on.
The suite is som-1.0.0+lib-0.2.2: SOM 1.0 and the skill library 0.2.2. Its messages are the published SOM 1.0 examples (the hurricane run, the orphan clip), the same fixtures the skill dev kit will play. All of it is synthetic data.
A result reads "passes som-bus suite som-1.0.0+lib-0.2.2". It is a sandbox result, not a certification, and it says nothing about suites or apps it didn't run.
Who can run tests
| Viewer | Developer | Admin | Owner | |
|---|---|---|---|---|
| See runs and their verdicts | ✓ | ✓ | ✓ | ✓ |
| Start a run, submit a consumer's end state | ✓ | ✓ | ✓ |
- Vendor workspaces only. A run injects messages; your vendors run theirs in their own workspaces and can share the results with you. A publisher grades its own in-house apps in its workbench, a vendor-type workspace of its own: see Your workbench.
- Only your organisation's apps. A run is refused while another organisation's app has a connection in the workspace (a revoked consumer connection counts, since restoring it reattaches its queue), so test messages never reach anyone else.
- One run playing at a time per workspace, and at most 50 runs in any 24 hours.
- Every start and every result is recorded in the workspace's audit log. Runs are kept 90 days.
What you need first
- A consumer connection that reads your workspace, and a way to read it: a client secret for the HTTPS pull API, or an AWS reader role you register on the connection. Create it under Consumer connections.
- For the skill harness, a producer connection that may publish
skill.warning.raised, with thesystem_idyour executor stamps (Credentials).
How a run works
- You choose consumer scenarios or the skill harness, and which cases.
- The bus plays each case's messages into your workspace, in order, as the producer
som-bus-test-runs. They skip the gateway on purpose: some scenarios send what the gateway would refuse, such as a late, older snapshot. Each message carries two extra attributes,sandbox_scenario(the case id) andtest_run(the run id), which an IAM reader of the queue sees. The HTTPS pull API doesn't return them: tell a run's messages apart by their story ids (next step). - Story ids are made unique per run and case:
<story_id>.<run>.<n>, for examplehurricane-2026-0911.rk7m2q4xhab3n.1. Every run starts clean, even for an app that keeps state per story (and it should). - The verdicts are recorded, and the run is complete. The workspace's owners get an email with the pass and fail counts, unless the workspace chose otherwise: see Notifications.
Consumer scenarios
The bus can't see inside your consumer, so after the run has played, you report the end state your consumer holds. The Test runs page shows the run's story ids and a template. Submit it within 24 hours:
json
{
"stories": { "hurricane-2026-0911.rk7m2q4xhab3n.1": 7 },
"processed": ["0199a1c4-7a2e-7b31-8c55-4d2f9e6a1b07", "…"]
}stories: for each story, the highestsequence_numberyour consumer holds.processed: everymessage_idyour consumer acted on, once per time it acted. A message you acted on twice appears twice.
| Case | What it plays | A conforming consumer |
|---|---|---|
consumer.replay | The seven hurricane snapshots | Holds sequence 7 and acted on all seven |
consumer.duplicate | Snapshot 2 delivered twice, same message_id | Acts on it once |
consumer.late_snapshot | Snapshots 1–3, then snapshot 2 again as a new message | Keeps sequence 3 |
consumer.unknown | An unknown envelope extension, a message type SOM 1.0 doesn't define, the next snapshot | Ignores what it doesn't know and holds sequence 2 |
consumer.som_1_1 | Two orphan stories, one from a 1.0.0 and one from a 1.1.0 producer | Holds both; doesn't branch on som_version |
Why: Ordering and delivery and Forward compatibility.
Skill harness
Your executor must have this configured instance loaded: house-breaking-indicative-category of smart-stories/raise-flag-on-match 0.2.2, watching lifecycle.phase for BREAKING, severity flag. To see a passing run before you build your own, run RND's reference executor, which has this instance loaded, against your workspace. Each warning is matched to its trigger by causation_id, and the harness waits 30 seconds after each trigger (plus a few seconds for the bus to record the warning).
| Case | What it plays | Expected |
|---|---|---|
skill.hurricane | The hurricane run: DEVELOPING, then BREAKING from snapshot 2 | One warning, on snapshot 2 |
skill.killed | A BREAKING snapshot, then the story KILLED | One warning, on the first snapshot |
skill.telling | A BREAKING snapshot, then telling.started | One warning, on the snapshot |
skill.redelivery | Snapshot 2 delivered again with the same message_id | Still exactly one warning |
skill.orphan | An ORPHAN shell (no lifecycle) | No warning |
skill.unknown_extension | The orphan shell with an unknown envelope extension | No warning |
skill.som_1_1 | The orphan shell from a 1.1.0 producer | No warning |
skill.ingest | The orphan shell, then asset.ingested | No warning. Pending spec |
A warning the gateway refuses never reaches the bus, so the harness sees it as missing. Look up the refusal in Activity; the Playground shows the same verdict for a single message.
Pending spec
skill.ingest needs asset.ingested, a message type with no SOM 1.0 schema. It runs and is shown, greyed out with the reason, but never fails a run. A run's result doesn't count it.
Where a check comes from
Every check names what it rests on:
| Rests on | Meaning | Verdict when it doesn't hold |
|---|---|---|
| SOM | The SOM 1.0 schema or its conformance rules | Fail |
| Library | The skill library 0.2.2: the skill file and its conventions | Fail |
| som-bus | This bus's delivery model: at least once, per correlation_id order | Fail |
| P-nn | An RND position where SOM and the library are silent. Ours, not the standard | Note, never a failure |
The one exception is P-15, the 30-second limit, which fails a case. The positions cited here:
| Position | Says |
|---|---|
| P-06 | A trigger that isn't a snapshot is a wake-up, evaluated against the latest snapshot |
| P-08 | Raise when a declaration opens or changes; nothing while it is unchanged |
| P-09 | skill_warning_ref is <rule_id>/<scope>/<n>, the same for every warning of one declaration |
| P-12 | Nothing is raised once a story is KILLED, SPIKED or ARCHIVED |
| P-14 | A configured instance never triggers on its own warning |
| P-15 | Publish within 5 seconds; the harness fails a case at 30 |
Checks
| Check | Kind | Holds when |
|---|---|---|
warning.raised | skill | The expected warning arrived, with the expected fields |
warning.missing | skill | Fails: no expected warning within 30 seconds |
warning.fields | skill | Fails: a field differs from what the configured instance raises (skill_id, skill_version, rule_id, severity, non_overridable, affected_fields, blocks, story_id, scope) |
warning.causation_id | skill | Every warning's causation_id is a message the case played |
warning.correlation_id | skill | A warning carries its trigger's correlation_id |
warning.story_id | skill | A warning's story_id is its trigger's story |
warning.duplicate | skill | One trigger, one warning. A retry must reuse the stored bytes, so the bus absorbs it |
warning.unexpected | skill | No warning where none is expected |
warning.timeout | skill | Fails: a warning arrived after 30 seconds (P-15) |
warning.ref | skill | skill_warning_ref follows P-09 (a note otherwise) |
warning.none | skill | No warning, as expected |
consumer.sequence | consumer | The story is held at the expected sequence_number |
consumer.processed | consumer | The consumer acted on the messages it must |
consumer.processed_once | consumer | No message_id was acted on twice |
cross.telling_pairs | run | telling.ended follows telling.started; exposure_start never changes |
cross.warning_after_terminal | run | No warning on a terminal story (P-12) |
cross.warning_repeated | run | A declaration raises again only when the warning changes (P-08) |
cross.warning_self_trigger | run | No warning caused by the same rule's warning (P-14) |
Cross-message checks (cross.*) run over everything the run played and everything your apps published about its stories. A check with nothing to look at shows n/a.
What a run keeps
Ids, the compared values and the verdicts. Never your app's payloads: a warning's detail is compared, not stored. The messages a run plays stay in your workspace's story timeline for 7 days, like any other traffic.
Not yet Coming
- Submitting a consumer's end state from your app with its own credentials, instead of from the portal.
- Registering your configured skill instances on the bus, checked against the suite's schema.
- More cases: the other library skills,
usage[]againstlink.*events, ORPHAN shells retiring. - Comparing two runs.
To share what passed, make a results statement from your complete runs.