Skip to main content
The engine handles failures at each layer with explicit, bounded behavior.

Retries

Guard failures and business-logic errors are not retried. The handler retry base is configurable via policy.yaml handler_retry_base_seconds.

Dead letters

After retries are exhausted, or on a non-retryable error, the event lands in dead_letters with a failure_type:
  • retry_exhausted: retried the maximum number of times, all failed.
  • handler_error: a non-retryable validation, business-logic, or guard-escalated error.
  • target_resolution_failed: a pin-routed event could not resolve a recipient.
  • chain_depth_exceeded: the causal chain hit its limit (see below).
A dead-letter record carries the original event and payload, route identities, the error message, retry count, chain depth, and handler node.

Chain depth

Each event carries a chain_depth counter. When an emit would exceed 50, the emitted event is intercepted and dead-lettered with failure_type: chain_depth_exceeded. The triggering handler still succeeds: chain-depth overflow is an emission interception, not a handler failure, so the handler’s other side effects still commit.

Terminal-state rejection

An event targeting an entity in a terminal state is rejected before any handler runs (outcome terminal_reject). This is unconditional.

Stalled-run escalation

A run that sits in operational_state=stalled for longer than a policy threshold has the runtime emit platform.run_stalled so an operator can react. The default threshold is 5 minutes (300 seconds); the event is platform-emitted, so subscribe to it directly with no declaration in your events.yaml.
Tune the threshold and toggle the surface in policy.yaml:
policy.yaml
You get one stall alert per stall, not a stream of them: at most one platform.run_stalled is emitted per (run_id, blocking_layer, blocking_reason, last_progress_at) tuple. A new event fires only when canonical progress advances or the diagnosis tuple changes while the run remains stalled.
The platform’s own platform.run_stalled emit does not count as progress, so it cannot accidentally suppress later escalations.
The diagnosis underlying the escalation is the same one swarm run status <run-id> surfaces on demand. The escalation just makes it active rather than poll-only.

Incidents and logs

swarm incidents aggregates failures by error code and component. swarm logs [--follow] queries or streams runtime logs. Both read through the runtime API.

How failures are named

Every runtime failure is recorded as a structured record with four parts: what class of failure it was, a more specific reason code, whether it can be retried, and what to do about it. These are produced by the runtime itself, never guessed from error text. What this buys you operationally:
  • Deterministic: the same recorded inputs always produce the same failure class on replay.
  • Closed namespace: product contracts cannot mint platform.* classes; authoring mistakes surface as verify findings instead, so a runtime failure always means “valid contract, real runtime fact.”
  • Cause carried: dead letters and incidents persist the validated envelope with its cause chain — swarm event view on a dead letter and swarm incidents group by the same class/selector taxonomy, so “what failed” has one vocabulary from log line to aggregate.
Routing failures are their own taxonomy slice worth knowing. Each publish-time dead letter names the exact identities involved, so a routing gap is a one-line diagnosis rather than digging through the event store: