Retries
Guard failures and business-logic errors are not retried. The handler retry base is
configurable via
policy.yaml handler_retry_base_seconds.
Dead letters
After retries are exhausted, or on a non-retryable error, the event lands indead_letters
with a failure_type:
retry_exhausted: retried the maximum number of times, all failed.handler_error: a non-retryable validation, business-logic, or guard-escalated error.target_resolution_failed: a pin-routed event could not resolve a recipient.chain_depth_exceeded: the causal chain hit its limit (see below).
Chain depth
Each event carries achain_depth counter. When an emit would exceed 50, the emitted
event is intercepted and dead-lettered with failure_type: chain_depth_exceeded. The
triggering handler still succeeds: chain-depth overflow is an emission interception, not a
handler failure, so the handler’s other side effects still commit.
Terminal-state rejection
An event targeting an entity in a terminal state is rejected before any handler runs (outcometerminal_reject). This is unconditional.
Stalled-run escalation
A run that sits inoperational_state=stalled for longer than a policy threshold has the
runtime emit platform.run_stalled so an operator can react. The default threshold is
5 minutes (300 seconds); the event is platform-emitted, so subscribe to it directly with
no declaration in your events.yaml.
policy.yaml:
policy.yaml
platform.run_stalled
is emitted per (run_id, blocking_layer, blocking_reason, last_progress_at) tuple. A new
event fires only when canonical progress advances or the diagnosis tuple changes while the
run remains stalled.
The platform’s own
platform.run_stalled emit does not count as progress, so it cannot
accidentally suppress later escalations.swarm run status <run-id> surfaces
on demand. The escalation just makes it active rather than poll-only.
Incidents and logs
swarm incidents aggregates failures by error code and component. swarm logs [--follow]
queries or streams runtime logs. Both read through the runtime API.
How failures are named
Every runtime failure is recorded as a structured record with four parts: what class of failure it was, a more specific reason code, whether it can be retried, and what to do about it. These are produced by the runtime itself, never guessed from error text. What this buys you operationally:- Deterministic: the same recorded inputs always produce the same failure class on replay.
- Closed namespace: product contracts cannot mint
platform.*classes; authoring mistakes surface as verify findings instead, so a runtime failure always means “valid contract, real runtime fact.” - Cause carried: dead letters and incidents persist the validated envelope with its
cause chain —
swarm event viewon a dead letter andswarm incidentsgroup by the same class/selector taxonomy, so “what failed” has one vocabulary from log line to aggregate.

