Skip to main content

swarm run vs swarm serve

Two ways to start a runtime, picked by how long the work lives:
  • swarm run start is a shortcut: boot a runtime in-process, publish one trigger event, stream the trace, exit when the run goes quiet — nothing left running and no events waiting to be delivered (we call this quiescing). It works for any flow that completes from a single trigger, deterministic or agent-driven. The Quickstart uses it, CI checks use it, and event publish-style offline iteration uses it. The process going away when the run finishes is the point.
  • swarm serve is the long-running runtime you operate. It owns the API, health, and MCP listener and stays up across many runs and many external interactions. Use it when the flow needs the runtime to outlive a single trigger: timers waiting on wall-clock, human decisions arriving later, multiple externally-driven events, anything you’ll interact with over time.
The most common newcomer confusion: trying swarm run start for a flow that needs you to interact with it mid-flight (decide a mailbox item, publish a follow-up event, watch a timer fire later), then being surprised the process exits before you can. For those, use swarm serve and drive it from a second terminal with swarm event publish, swarm mailbox, and the rest of the CLI. Both need a workspace data directory: pass --data <dir> (or set workspace.data_source in swarm.yaml) — boot fails without one. The rest of this page is about swarm serve.

swarm serve

swarm serve is the long-running runtime. It owns the API, health, and MCP listener. By default it binds 127.0.0.1:8081, serving health, readiness, /v1/rpc, /v1/ws, and the MCP routes on one listener.
The loopback default means the runtime is only reachable from the same machine. To accept connections from elsewhere, bind a public address on purpose (for example --health-addr 0.0.0.0:8081); the loopback default is there so you do not expose the API by accident. A few flags matter for operations:
  • --shutdown-grace <duration> (default 30s) is how long shutdown waits for in-flight work to finish before it cancels what is left and exits.
  • --no-require-bundle-match turns off a safety check. Normally, if a run is still active and was started from a different contract bundle than the one you are now booting, swarm serve refuses to start, so you do not quietly process an old run with new contracts. This flag lifts that check for recovery or development.
  • --abandon-active-runs is destructive: at startup it cancels every run that was still running or paused and moves their in-flight work to the dead-letter queue (failed work parked for inspection). Use it to get a clean slate after a crash, not against runs you care about. (It cancels work; it does not delete data or stop containers. That is swarm control nuke.)
  • --dev bundles the loose settings for local work: it turns on --abandon-active-runs and --no-require-bundle-match, makes boot verbose, and deletes the per-entity workspace containers when the runtime shuts down. Handy locally, but not something to point at state you want to keep.
Editing contracts and restarting on the same store can currently leave publish failing with connect route snapshot generation is stale (tracked upstream). Until fixed, treat a contract edit as a fresh store: stop the runtime and delete the project’s directory under ~/.swarm/stores/projects/ before restarting.

Boot verification

At boot the engine loads the bundle, runs it through the static analyzer, initializes the Postgres state stores, then starts system nodes and agents. If any error-severity check fails, boot aborts; there is no partial startup. A runtime that is up has passed the analyzer.

Run lifecycle

A run is one execution of your flow. It moves through a few states you will see in swarm run list and swarm run status:
  • running: work is in progress.
  • completed: the run finished cleanly. This happens only when every entity has reached a terminal state and no events are still waiting to be delivered. An open timer counts as pending work, so a run with a timer still ticking stays running until that timer fires or is cancelled.
  • failed: something broke at the platform level (the database went away, the contract bundle was corrupt).
  • paused: held by swarm control pause (below).
  • stalled: still running but not making progress. swarm run status <run-id> reports this and names the blocking layer and reason.
Handler and agent failures don’t fail the run. A failed handler becomes a dead letter and the run keeps going — so a run can reach completed with dead letters inside it. Watch dead letters separately; run status won’t surface logic errors.
One thing about querying: entity state is scoped to the current run. swarm entity list and the entity query tools see the entities in the run you are looking at, not across past runs, so any “have I seen this before?” check spanning runs has to be built explicitly.

Health

Two unauthenticated process probes are always available: GET /healthz (the process is alive) and GET /readyz (the runtime is ready: database up, contracts loaded). For an authenticated operator summary with runtime identity, use swarm health (the health.check method).

Pausing ingress

swarm control pause --all pauses runtime ingress — the runtime stops handing new events and calls to your flow (the runtime.pause method); swarm control continue --all resumes it. Use it for maintenance windows. While paused:
  • New events are still accepted and saved; they wait in the queue and are delivered once when you resume.
  • In-flight work finishes its current step and stops there.
  • Read-only operator methods keep working; live tool and MCP calls are refused until you resume.
  • Lifecycle signals: pausing emits platform.paused, resuming emits platform.resumed.
Note what --all does and does not do. pause --all gates new ingress; it does not reach into each running run and freeze it in place. To act on one run, name it: swarm control pause <run-id>. A third control, swarm control stop, ends runs rather than pausing them. stop --all works differently again: it stops each active run one at a time rather than flipping a single ingress switch, so it asks for confirmation.

Surfaces

The runtime exposes one API (/v1/rpc for request/response, /v1/ws for subscriptions), the MCP gateway, and the swarm CLI as a client. See the API reference.