Skip to main content
This guide builds a small but complete flow from an empty directory. By the end you will have a contract bundle that passes the static analyzer and boots, and, more importantly, you will understand what each file does and why. You will build a support-ticket system: a ticket arrives, an agent classifies it, a deterministic orchestrator routes it to a resolver agent, and it closes. If the resolver cannot solve it, the ticket loops back and tries again until it hits a limit.

The mental model

Before the files, three ideas. Everything in Swarm is built from them:
  1. Events are messages. Nothing happens except in response to an event. ticket.created, ticket.classified, and so on are events, each carrying a typed payload.
  2. A system node is deterministic code. No LLM. It subscribes to events and, for each one, runs a fixed pipeline: check a condition, write some data, advance the stage, emit the next event. Same input, same result, every time.
  3. Agents are LLM workers. They subscribe to events, reason, and emit events. An agent never changes state directly; it emits an event with its result, and a system node decides what to write. That is what keeps the system auditable.
Here is the whole flow as events moving between the orchestrator (deterministic) and the two agents (LLM):

The files you will write

A flow is a directory of small YAML files, each with one job:
That is more declaration than you would write to script an LLM directly. If you are already wondering whether it is worth it, jump to What the contracts buy you, then come back.
Create them one at a time.
Every contract below passes swarm verify. The bundle uses only built-in scalar types and a single self-contained flow, so there is nothing else to set up.
1

Name the flow (package.yaml)

Every flow starts with a manifest. It names the flow and declares which Swarm versions it runs on.
package.yaml
flows is where larger systems list child flows; ours is a single flow, so it is empty. See the flow package for the complete file list.
2

Declare the lifecycle (schema.yaml)

schema.yaml is the flow’s public surface. It declares three things: the stages a ticket moves through, the pins (which events the flow accepts from outside), and the agent roles the flow needs.
schema.yaml
Stages. A ticket is always in exactly one stage; initial: true marks where new tickets start and terminal: true marks the absorbing end. The analyzer rejects a stage nothing ever advances to, so the map is exactly the four the flow uses.Pins are the flow’s public interface. An input event pin is the front door: ticket.created is the only event the outside world can send in.Required agents are the roles the flow depends on. Each says what it listens to and what it emits; we will provide matching agents in agents.yaml. Boot fails if a declared role has no matching agent. See the state machine reference.
3

Declare the data (entities.yaml)

An entity is the thing moving through the staged lifecycle, here a single ticket. It is a record with typed fields plus its current stage. Declare the fields handlers will fill in.
entities.yaml
Most fields use the short form name: type. Three things to note:
  • resolution uses the longer form because it is written by the resolver and then read from external operator views such as swarm entity view, not by another internal handler. _unused_reader_reason tells the analyzer that the missing internal reader is intentional.
  • escalation_count uses the longer form because it needs an initial value: every declared field must have an initial or a handler that writes it, or the analyzer reports it as uncovered.
  • You do not declare the ticket’s stage here; the platform tracks it for you (exposed to expressions as _entity.current_state).
4

Declare the events (events.yaml)

Events are how every part of the flow talks to every other part. Each event declares the fields its payload carries. Routing is not declared here; it comes entirely from who subscribes to what (the next two files).
events.yaml
Fields go directly under the event name (there is no payload: wrapper). swarm.source: external tells the analyzer that ticket.created comes from outside, so it does not expect something inside the flow to produce it. Whatever emits an event must fill every field it declares; there are no defaults and no automatic copy from the triggering event.Notice each event carries only what its handler needs: ticket.created brings the raw text from outside, but the internal events pass just the structured verdict (category, priority) and the result. The platform identifies the ticket by its entity, so there is no need to thread a ticket_id field through every event. See Events and routing.
5

Write the orchestrator (nodes.yaml)

This is the engine room. A system node is deterministic code that subscribes to events and runs one handler per event. A handler runs a fixed pipeline and commits it in a single transaction: optionally guard (check a condition), write fields, advance the stage, and emit the next event.
nodes.yaml
Read it handler by handler:
  • ticket.created is the entry point. A stateful flow’s input-pin handler must say how it gets its entity; create_entity: true mints a fresh ticket. It then advances to new.
  • ticket.classified copies category and priority from the event onto the ticket (data_accumulation), advances to assigned, and emits ticket.assigned for the resolver. The source_field/target_field form copies a payload field to an entity field.
  • ticket.escalated runs a guard first: entity.escalation_count < policy.max_escalations. If the ticket has bounced too many times the guard fails and on_fail: reject stops it; otherwise it increments the counter (a computed write, expression) and routes back to assigned to try again.
  • ticket.resolved saves the resolution and emits ticket.resolution_confirmed. Here the short writes: [resolution, resolved_by] form copies same-named payload fields.
  • ticket.resolution_confirmed advances to closed, a terminal state.
Filling the emitted payload. When a handler emits an event that carries fields, it must populate every one: there are no defaults and nothing is copied automatically from the triggering event. That is why ticket.classified uses the object form of emit, with a fields map, instead of the bare emit: ticket.assigned. Each value is a small expression, here entity.category and entity.priority, reading the ticket fields the same handler just wrote. (The bare string form is only for events that declare no payload fields.)Exactly one system node may handle a given event, which keeps state changes unambiguous. For every field a handler can use, see the handler reference.
6

Add the agents (agents.yaml)

Agents are the LLM workers. Each subscribes to events, reasons, and emits events. Note what they do not do: they never write the ticket’s fields directly. The resolver puts its answer in the ticket.resolved payload, and the orchestrator’s handler writes it. Agents emit; system nodes decide what to persist.
agents.yaml
Each agent’s emit_events automatically gives it a tool to emit those events (for example emit_ticket_classified). memory: false means a fresh conversation per event with no memory between tickets.
One persistent conversation per ticket requires a flow that has one instance per ticket — this single flow does not. See Composing flows.
7

Set policy (policy.yaml)

Policy holds configuration values. Guards read them as policy.X, and prompts use them as {{X}}. The escalation guard above read policy.max_escalations.
policy.yaml
8

Write the prompts (prompts/)

Each agent gets a markdown prompt that tells it what to do and what to emit. {{variable}} placeholders are filled from policy and instance values at run time.
prompts/classifier-agent.md
prompts/resolver-agent.md
The resolver decides from category because that is what ticket.assigned carries. An agent acts on the event payload it receives, so the payload is the contract for what each agent can see: the classifier reads the raw subject and body from ticket.created, while the resolver works from the structured verdict.
9

Verify

Expected output:
The analyzer checks payload coverage, state reachability, agent fulfillment, entity writer coverage, handler fields, and CEL parsing, among others. If it reports an error, the analyzer-checks reference explains each one and how to fix it.

How the pieces connect

Trace one ticket through the diagram at the top:
  1. ticket.created arrives from outside. The orchestrator mints the ticket and advances it to new; the classifier (subscribed to the same event) reads it and emits ticket.classified.
  2. The orchestrator handles ticket.classified: it saves category and priority, advances to assigned, and emits ticket.assigned.
  3. The resolver handles ticket.assigned and emits either ticket.resolved (the orchestrator saves the resolution, advances to resolved, and emits ticket.resolution_confirmed, which closes the ticket) or ticket.escalated (the guard checks the count, increments it, and sends the ticket back to assigned).
Notice the division of labor throughout: agents decide what (classify, resolve, escalate); the orchestrator decides what gets written and what happens next. Every declared state is reachable, and every event has either a handler or a subscriber.

What the contracts buy you

A single LLM with a script is fine for one task. It breaks down when you make the model be the system — tracking state, following rules exactly, staying coherent across a long job. Swarm keeps the judgment in the LLM and the rigid parts in deterministic nodes.
  • Reliability: the model cannot fake the outcome. An agent can say it resolved the ticket, but saying so does not make it so. Agents only emit events; a deterministic system node decides what is written and whether the ticket advances. It reaches resolved because a handler advanced it after ticket.resolved arrived, never because the model asserted it. Rules like the escalation cap are checks the platform runs every time, not instructions you hope the model remembers. And because each transition commits atomically, the ticket is always in exactly one declared state, never half-updated.
  • Compose many agents, for as long as it takes. Each agent is a small unit wired to the others only through typed events and pins, never one shared mega-prompt. You grow a system by adding roles and whole sub-flows (coordinator, managers, workers), not by enlarging a central script. This flow has two agents; the same model coordinates hundreds, and a run can span hours or days, surviving restarts along the way.
  • No context explosion, and sharper agents. A single agent, or a flat “everyone in one chat”, drowns as the work grows: the context window fills and quality drops. Each Swarm agent runs in a scoped session that sees only its own events, so the classifier never carries the resolver’s history. Splitting work into small, focused tasks is not just how you scale; it is what makes each individual LLM call more reliable.
And, almost for free, the things you would otherwise hand-build:
  • The wiring is checked before it runs. swarm verify caught the bugs a script hides until production: a state nothing reaches, a missing payload field, an unfilled role, a dangling event.
  • Every run is auditable and replayable. The event log, the before/after of every field, and each agent’s turn are recorded and tied to the run, so “why did it do that?” is a query. You can replay a run or fork it against fixed contracts.
  • A human step is one line away (mailbox_write): the decision becomes just another event, with no queue or resume code to build.

When a script is the better call

Be honest about the fit. For a single LLM call, a short conversation, or a workflow you will rewrite next week, a plain script is the right tool and Swarm is overkill. The contracts pay off when many agents must coordinate, the work is long-running, the state has to stay consistent, or it runs at volume. See Why Swarm for the full picture.

Run it

The runtime defaults to SQLite at .swarm/dev.db, so no DB service is needed; the agents do need an LLM runtime configured to answer (see Installation). swarm run start boots a runtime in process on loopback, publishes the trigger, and streams the trace:
On loopback, swarm run start uses a built-in dev API token; an explicit token (--api-token-file) is only needed when exposing the API beyond loopback.
This flow has LLM agents, so it advances past new only when a runtime that can answer them is configured: a real provider (swarm serve --backend anthropic with ANTHROPIC_API_KEY) or the Claude CLI runtime (swarm serve --backend claude_cli). The deterministic system-node steps run either way; the agent steps wait until a runtime is present.

Watch one run

Here is the trace from a real run of this bundle, lightly cleaned (internal platform.* log lines removed, IDs shortened) and annotated. It is the whole flow on one screen, the abstract pieces made concrete:
The subscriber=node/__runtime_replay_scope__ you see on most lines is not your ticket-orchestrator. It is the platform’s internal subscriber that writes every event to the log (what makes replay possible). Your orchestrator’s work shows up not as a delivery line but as the next event it emits, which is how you read the chain below. What each numbered line is doing, mapped back to the four ideas:
  1. ticket.created is published and persisted. The event is written to the log before anything runs, so the run can be replayed from here. The orchestrator (the one system node) handles it in a single transaction: it mints the ticket and advances it to new. No separate “state changed” line appears because the transition is the handler, committed atomically.
  2. Same event, delivered to the classifier agent. Routing came only from subscriptions: both the orchestrator and classifier-agent subscribe to ticket.created, so both receive it. in_progress means the agent’s LLM turn is running.
  3. The classifier emits ticket.classified. The orchestrator handles it, writes category and priority onto the ticket, advances to assigned, and emits ticket.assigned. One event in, one transition, one event out.
  4. ticket.assigned is delivered to the resolver. It carries the category the orchestrator projected with emit.fields, which is all the resolver needs to decide.
  5. The resolver emits ticket.resolved. (This was a technical ticket, so the resolver resolved it; a billing or account ticket would emit ticket.escalated here, and the guard you wrote would cap the retries.) The orchestrator saves the resolution and advances to resolved.
  6. ticket.resolution_confirmed closes the loop. Its handler advances the ticket to closed, a terminal state, and the run goes quiet.
Every line is an event moving between the deterministic orchestrator and the LLM agents, exactly the sequence diagram at the top, now with real timestamps behind it. The agents only ever emitted; a system node decided what each event changed and what fired next.
Want to see the other branch? Send a ticket the classifier will label billing or account (for example, subject “I was double charged”). The resolver emits ticket.escalated, the orchestrator’s guard checks escalation_count < max_escalations, increments the count, and re-emits ticket.assigned, the retry loop, until the cap rejects it. The LLM keeps proposing; the deterministic guard decides when to stop.
For LLM agents helping a human author this: see Agentic flow authoring for the ranked artifacts to pull, the iteration loop to run, and anti-patterns to avoid.

Core concepts

The model behind flows, events, handlers, and agents.

Writing handlers

Guards, branching, accumulation, and actions in depth.

Handler patterns

Reusable shapes for common orchestration problems.

Testing

The analyzer, test packages, and agent fixtures.