The mental model
Before the files, three ideas. Everything in Swarm is built from them:- Events are messages. Nothing happens except in response to an event.
ticket.created,ticket.classified, and so on are events, each carrying a typed payload. - A system node is deterministic code. No LLM. It subscribes to events and, for each one, runs a fixed pipeline: check a condition, write some data, advance the stage, emit the next event. Same input, same result, every time.
- Agents are LLM workers. They subscribe to events, reason, and emit events. An agent never changes state directly; it emits an event with its result, and a system node decides what to write. That is what keeps the system auditable.
The files you will write
A flow is a directory of small YAML files, each with one job:
Create them one at a time.
Every contract below passes
swarm verify. The bundle uses only built-in scalar types and a
single self-contained flow, so there is nothing else to set up.1
Name the flow (package.yaml)
Every flow starts with a manifest. It names the flow and declares which Swarm versions it
runs on.
package.yaml
flows is where larger systems list child flows; ours is a single flow, so it is empty.
See the flow package for the complete file list.2
Declare the lifecycle (schema.yaml)
schema.yaml is the flow’s public surface. It declares three things: the stages a
ticket moves through, the pins (which events the flow accepts from outside), and the
agent roles the flow needs.schema.yaml
initial: true marks where new
tickets start and terminal: true marks the absorbing end. The analyzer rejects a stage
nothing ever advances to, so the map is exactly the four the flow uses.Pins are the flow’s public interface. An input event pin is the front door:
ticket.created is the only event the outside world can send in.Required agents are the roles the flow depends on. Each says what it listens to and
what it emits; we will provide matching agents in agents.yaml. Boot fails if a declared
role has no matching agent. See the state machine reference.3
Declare the data (entities.yaml)
An entity is the thing moving through the staged lifecycle, here a single ticket. It is a
record with typed fields plus its current stage. Declare the fields handlers will fill in.Most fields use the short form
entities.yaml
name: type. Three things to note:resolutionuses the longer form because it is written by the resolver and then read from external operator views such asswarm entity view, not by another internal handler._unused_reader_reasontells the analyzer that the missing internal reader is intentional.escalation_countuses the longer form because it needs aninitialvalue: every declared field must have aninitialor a handler that writes it, or the analyzer reports it as uncovered.- You do not declare the ticket’s stage here; the platform tracks it for you (exposed
to expressions as
_entity.current_state).
4
Declare the events (events.yaml)
Events are how every part of the flow talks to every other part. Each event declares the
fields its payload carries. Routing is not declared here; it comes entirely from who
subscribes to what (the next two files).Fields go directly under the event name (there is no
events.yaml
payload: wrapper). swarm.source: external tells the analyzer that ticket.created comes from outside, so it does not
expect something inside the flow to produce it. Whatever emits an event must fill every
field it declares; there are no defaults and no automatic copy from the triggering event.Notice each event carries only what its handler needs: ticket.created brings the raw text
from outside, but the internal events pass just the structured verdict (category,
priority) and the result. The platform identifies the ticket by its entity, so there is no
need to thread a ticket_id field through every event. See
Events and routing.5
Write the orchestrator (nodes.yaml)
This is the engine room. A system node is deterministic code that subscribes to events
and runs one handler per event. A handler runs a fixed pipeline and commits it in a
single transaction: optionally guard (check a condition), write fields, advance
the stage, and emit the next event.Read it handler by handler:
nodes.yaml
ticket.createdis the entry point. A stateful flow’s input-pin handler must say how it gets its entity;create_entity: truemints a fresh ticket. It then advances tonew.ticket.classifiedcopiescategoryandpriorityfrom the event onto the ticket (data_accumulation), advances toassigned, and emitsticket.assignedfor the resolver. Thesource_field/target_fieldform copies a payload field to an entity field.ticket.escalatedruns a guard first:entity.escalation_count < policy.max_escalations. If the ticket has bounced too many times the guard fails andon_fail: rejectstops it; otherwise it increments the counter (a computed write,expression) and routes back toassignedto try again.ticket.resolvedsaves the resolution and emitsticket.resolution_confirmed. Here the shortwrites: [resolution, resolved_by]form copies same-named payload fields.ticket.resolution_confirmedadvances toclosed, a terminal state.
ticket.classified uses the object form of emit, with a
fields map, instead of the bare emit: ticket.assigned. Each value is a small expression,
here entity.category and entity.priority, reading the ticket fields the same handler just
wrote. (The bare string form is only for events that declare no payload fields.)Exactly one system node may handle a given event, which keeps state changes unambiguous.
For every field a handler can use, see the handler reference.6
Add the agents (agents.yaml)
Agents are the LLM workers. Each subscribes to events, reasons, and emits events. Note what
they do not do: they never write the ticket’s fields directly. The resolver puts its
answer in the Each agent’s
ticket.resolved payload, and the orchestrator’s handler writes it. Agents
emit; system nodes decide what to persist.agents.yaml
emit_events automatically gives it a tool to emit those events (for example
emit_ticket_classified). memory: false means a fresh conversation per event with no
memory between tickets.One persistent conversation per ticket requires a flow that has one instance per
ticket — this single flow does not. See Composing flows.
7
Set policy (policy.yaml)
Policy holds configuration values. Guards read them as
policy.X, and prompts use them as
{{X}}. The escalation guard above read policy.max_escalations.policy.yaml
8
Write the prompts (prompts/)
Each agent gets a markdown prompt that tells it what to do and what to emit. The resolver decides from
{{variable}}
placeholders are filled from policy and instance values at run time.prompts/classifier-agent.md
prompts/resolver-agent.md
category because that is what ticket.assigned carries. An
agent acts on the event payload it receives, so the payload is the contract for what each
agent can see: the classifier reads the raw subject and body from ticket.created, while
the resolver works from the structured verdict.9
Verify
How the pieces connect
Trace one ticket through the diagram at the top:ticket.createdarrives from outside. The orchestrator mints the ticket and advances it tonew; the classifier (subscribed to the same event) reads it and emitsticket.classified.- The orchestrator handles
ticket.classified: it savescategoryandpriority, advances toassigned, and emitsticket.assigned. - The resolver handles
ticket.assignedand emits eitherticket.resolved(the orchestrator saves the resolution, advances toresolved, and emitsticket.resolution_confirmed, which closes the ticket) orticket.escalated(the guard checks the count, increments it, and sends the ticket back toassigned).
What the contracts buy you
A single LLM with a script is fine for one task. It breaks down when you make the model be the system — tracking state, following rules exactly, staying coherent across a long job. Swarm keeps the judgment in the LLM and the rigid parts in deterministic nodes.- Reliability: the model cannot fake the outcome. An agent can say it resolved the
ticket, but saying so does not make it so. Agents only emit events; a deterministic system
node decides what is written and whether the ticket advances. It reaches
resolvedbecause a handler advanced it afterticket.resolvedarrived, never because the model asserted it. Rules like the escalation cap are checks the platform runs every time, not instructions you hope the model remembers. And because each transition commits atomically, the ticket is always in exactly one declared state, never half-updated. - Compose many agents, for as long as it takes. Each agent is a small unit wired to the others only through typed events and pins, never one shared mega-prompt. You grow a system by adding roles and whole sub-flows (coordinator, managers, workers), not by enlarging a central script. This flow has two agents; the same model coordinates hundreds, and a run can span hours or days, surviving restarts along the way.
- No context explosion, and sharper agents. A single agent, or a flat “everyone in one chat”, drowns as the work grows: the context window fills and quality drops. Each Swarm agent runs in a scoped session that sees only its own events, so the classifier never carries the resolver’s history. Splitting work into small, focused tasks is not just how you scale; it is what makes each individual LLM call more reliable.
- The wiring is checked before it runs.
swarm verifycaught the bugs a script hides until production: a state nothing reaches, a missing payload field, an unfilled role, a dangling event. - Every run is auditable and replayable. The event log, the before/after of every field, and each agent’s turn are recorded and tied to the run, so “why did it do that?” is a query. You can replay a run or fork it against fixed contracts.
- A human step is one line away (
mailbox_write): the decision becomes just another event, with no queue or resume code to build.
When a script is the better call
Be honest about the fit. For a single LLM call, a short conversation, or a workflow you will rewrite next week, a plain script is the right tool and Swarm is overkill. The contracts pay off when many agents must coordinate, the work is long-running, the state has to stay consistent, or it runs at volume. See Why Swarm for the full picture.Run it
The runtime defaults to SQLite at.swarm/dev.db, so no DB service is needed; the agents do
need an LLM runtime configured to answer (see Installation). swarm run start
boots a runtime in process on loopback, publishes the trigger, and streams the trace:
swarm run start uses a built-in dev API token; an explicit token (--api-token-file) is only
needed when exposing the API beyond loopback.
This flow has LLM agents, so it advances past
new only when a runtime that can answer them
is configured: a real provider (swarm serve --backend anthropic with ANTHROPIC_API_KEY) or
the Claude CLI runtime (swarm serve --backend claude_cli). The deterministic system-node steps
run either way; the agent steps wait until a runtime is present.Watch one run
Here is the trace from a real run of this bundle, lightly cleaned (internalplatform.* log
lines removed, IDs shortened) and annotated. It is the whole flow on one screen, the abstract
pieces made concrete:
subscriber=node/__runtime_replay_scope__ you see on most lines is not your
ticket-orchestrator. It is the platform’s internal subscriber that writes every event to the
log (what makes replay possible). Your orchestrator’s work shows up not as a delivery line but as
the next event it emits, which is how you read the chain below.
What each numbered line is doing, mapped back to the four ideas:
ticket.createdis published and persisted. The event is written to the log before anything runs, so the run can be replayed from here. The orchestrator (the one system node) handles it in a single transaction: it mints the ticket and advances it tonew. No separate “state changed” line appears because the transition is the handler, committed atomically.- Same event, delivered to the classifier agent. Routing came only from subscriptions:
both the orchestrator and
classifier-agentsubscribe toticket.created, so both receive it.in_progressmeans the agent’s LLM turn is running. - The classifier emits
ticket.classified. The orchestrator handles it, writescategoryandpriorityonto the ticket, advances toassigned, and emitsticket.assigned. One event in, one transition, one event out. ticket.assignedis delivered to the resolver. It carries thecategorythe orchestrator projected withemit.fields, which is all the resolver needs to decide.- The resolver emits
ticket.resolved. (This was atechnicalticket, so the resolver resolved it; abillingoraccountticket would emitticket.escalatedhere, and the guard you wrote would cap the retries.) The orchestrator saves the resolution and advances toresolved. ticket.resolution_confirmedcloses the loop. Its handler advances the ticket toclosed, a terminal state, and the run goes quiet.
Want to see the other branch? Send a ticket the classifier will label
billing or account
(for example, subject “I was double charged”). The resolver emits ticket.escalated, the
orchestrator’s guard checks escalation_count < max_escalations, increments the count, and
re-emits ticket.assigned, the retry loop, until the cap rejects it. The LLM keeps proposing;
the deterministic guard decides when to stop.What to read next
Core concepts
The model behind flows, events, handlers, and agents.
Writing handlers
Guards, branching, accumulation, and actions in depth.
Handler patterns
Reusable shapes for common orchestration problems.
Testing
The analyzer, test packages, and agent fixtures.

