swarm run vs swarm serve
Two ways to start a runtime, picked by how long the work lives:swarm run startis a shortcut: boot a runtime in-process, publish one trigger event, stream the trace, exit when the run goes quiet — nothing left running and no events waiting to be delivered (we call this quiescing). It works for any flow that completes from a single trigger, deterministic or agent-driven. The Quickstart uses it, CI checks use it, andevent publish-style offline iteration uses it. The process going away when the run finishes is the point.swarm serveis the long-running runtime you operate. It owns the API, health, and MCP listener and stays up across many runs and many external interactions. Use it when the flow needs the runtime to outlive a single trigger: timers waiting on wall-clock, human decisions arriving later, multiple externally-driven events, anything you’ll interact with over time.
swarm run start for a flow that needs you to
interact with it mid-flight (decide a mailbox item, publish a follow-up event, watch a timer
fire later), then being surprised the process exits before you can. For those, use
swarm serve and drive it from a second terminal with swarm event publish, swarm mailbox,
and the rest of the CLI.
Both need a workspace data directory: pass --data <dir> (or set
workspace.data_source in swarm.yaml) — boot fails without one.
The rest of this page is about swarm serve.
swarm serve
swarm serve is the long-running runtime. It owns the API, health, and MCP listener. By
default it binds 127.0.0.1:8081, serving health, readiness, /v1/rpc, /v1/ws, and the
MCP routes on one listener.
--health-addr 0.0.0.0:8081); the loopback default is there so you do not expose the API by
accident.
A few flags matter for operations:
--shutdown-grace <duration>(default30s) is how long shutdown waits for in-flight work to finish before it cancels what is left and exits.--no-require-bundle-matchturns off a safety check. Normally, if a run is still active and was started from a different contract bundle than the one you are now booting,swarm serverefuses to start, so you do not quietly process an old run with new contracts. This flag lifts that check for recovery or development.--abandon-active-runsis destructive: at startup it cancels every run that was stillrunningorpausedand moves their in-flight work to the dead-letter queue (failed work parked for inspection). Use it to get a clean slate after a crash, not against runs you care about. (It cancels work; it does not delete data or stop containers. That isswarm control nuke.)--devbundles the loose settings for local work: it turns on--abandon-active-runsand--no-require-bundle-match, makes boot verbose, and deletes the per-entity workspace containers when the runtime shuts down. Handy locally, but not something to point at state you want to keep.
Boot verification
At boot the engine loads the bundle, runs it through the static analyzer, initializes the Postgres state stores, then starts system nodes and agents. If anyerror-severity check
fails, boot aborts; there is no partial startup. A runtime that is up has passed the
analyzer.
Run lifecycle
A run is one execution of your flow. It moves through a few states you will see inswarm run list and swarm run status:
- running: work is in progress.
- completed: the run finished cleanly. This happens only when every entity has reached a
terminal state and no events are still waiting to be delivered. An open timer counts as
pending work, so a run with a timer still ticking stays
runninguntil that timer fires or is cancelled. - failed: something broke at the platform level (the database went away, the contract bundle was corrupt).
- paused: held by
swarm control pause(below). - stalled: still
runningbut not making progress.swarm run status <run-id>reports this and names the blocking layer and reason.
Handler and agent failures don’t fail the run. A failed handler becomes a
dead letter and the run keeps going — so a run can reach
completed
with dead letters inside it. Watch dead letters separately; run status won’t surface logic
errors.swarm entity list and the
entity query tools see the entities in the run you are looking at, not across past runs, so any
“have I seen this before?” check spanning runs has to be built explicitly.
Health
Two unauthenticated process probes are always available:GET /healthz (the process is
alive) and GET /readyz (the runtime is ready: database up, contracts loaded). For an
authenticated operator summary with runtime identity, use swarm health (the health.check
method).
Pausing ingress
swarm control pause --all pauses runtime ingress — the runtime stops handing new events and
calls to your flow (the runtime.pause method); swarm control continue --all resumes it. Use it for maintenance windows.
While paused:
- New events are still accepted and saved; they wait in the queue and are delivered once when you resume.
- In-flight work finishes its current step and stops there.
- Read-only operator methods keep working; live tool and MCP calls are refused until you resume.
- Lifecycle signals: pausing emits
platform.paused, resuming emitsplatform.resumed.
--all does and does not do. pause --all gates new ingress; it does not reach into
each running run and freeze it in place. To act on one run, name it: swarm control pause <run-id>. A third control, swarm control stop, ends runs rather than pausing them.
stop --all works differently again: it stops each active run one at a time rather than
flipping a single ingress switch, so it asks for confirmation.
Surfaces
The runtime exposes one API (/v1/rpc for request/response, /v1/ws for subscriptions), the
MCP gateway, and the swarm CLI as a client. See the
API reference.
