The problem with LLM-managed orchestration
Most multi-agent frameworks make the LLM be the system: it decides who acts next, tracks the state, and is trusted to follow the rules. LLMs are not reliable at that job, and it shows in production:- The model fakes the state. An agent reports a step done when it is not, or contradicts an earlier decision. Nothing catches it, because the model is also the bookkeeper: there is no independent source of truth.
- Non-deterministic routing. When an LLM decides who acts next, two runs of the same input diverge. You cannot reliably debug, audit, or reproduce a run.
- Context-window exhaustion. Put ten agents in one shared chat and the model loses the thread well before the tenth; quality falls as the task grows.
- No crash recovery. A process dies mid-turn and the in-memory state goes with it.
- No composition. Reusing a pipeline in another project means refactoring code, not importing a package.
A runtime, not just a router
Declaring agent steps in YAML and dispatching them in a fixed order, with no LLM in the routing loop, is the right baseline, and increasingly a standard one. But routing only decides who-acts-next. Running a multi-agent system in production is the larger problem, and it is the layer Swarm adds. Work is a set of entities, each moving through a durable state machine, so the runtime always knows where every one of them is and resumes mid-flight after a crash. Each agent is isolated down to its own workspace, so one cannot reach into another’s work. Token spend is metered per entity and per actor, with automatic throttling before a budget runs away. And humans act through the same event pipeline as everything else. A workflow runner dispatches agents and exits. Swarm is the runtime they would run inside.Strong fit
- Long-running workflows, hours to days. Durable timers, persistence, and crash recovery let a run span days without losing state; replay answers “why did the system do that, and when?” later.
- Coordinating many agents, from a handful to a hundred or more. Fan a batch out and each item gets its own flow instance with its own agent — isolated session, workspace, and lifecycle — while a declared join decides when the batch is complete. Agent count scales on a single deployment without the context-window blowup of a flat chat.
- Approval-gated automation. A risky external effect can require a typed human verdict before it exists: no request, no credential lookup, no provider call until someone approves the exact recorded action. Sign-off is a declared outcome with an audit trail, not a Slack message beside the system.
- Systems assembled from reusable parts. A flow is a self-contained package with typed input and output pins. Build a flow once and import it into other projects, or swap one for a better version, without refactoring.
- Crash recovery as a requirement. Atomic transitions and event replay are the baseline, not a bolt-on.
- Regulated or audited environments. Every event is stored, and so is every field change with its before/after value — so “what changed, when, and because of which event” is a single query.
- High-volume parallel entities. Orders, tickets, claims: hundreds flow through one contract without interfering, each in its own scoped state.
- Cost-controlled deployments. Token usage is tracked per entity and per actor, with thresholds that escalate into throttle and emergency states.
Weak fit
- One-shot LLM calls. Use the provider SDK directly: no runtime needed.
- Conversational chatbots — with a line worth drawing. A chatbot whose value ends at the reply is the provider SDK’s job; the contract overhead does not pay for itself. A durable assistant that remembers per conversation, waits on timers, and asks permission through chat is a different animal — that is a standing flow — a long-lived flow that stays up and reacts to incoming messages — and it is strong fit.
- Exploratory prototyping that changes hourly. The analyzer refusing to boot a
half-finished contract is exactly the wrong friction while you are still sketching. (The
sketch loop is cheaper than it was —
swarm testand the mock backend iterate a flow with zero credentials — but the contract-first discipline remains the price.) - LLM-managed routing at runtime. Swarm refuses this by design. If “the model decides what happens next” is the point of your system, this is the wrong tool.
Swarm vs. alternatives
Swarm sits closer to durable-execution engines (Temporal, Restate, Inngest) than to agent libraries; it applies that rigor to multi-agent LLM systems.The tradeoff, stated plainly
A Swarm flow cannot be re-wired by an LLM at runtime, and the static analyzer will refuse to boot a half-finished contract. In production that refusal is a feature; during exploration it is friction. Sketch loose first if you must, then port to Swarm once the workflow stabilizes. And one honest limit on the word “deterministic”: it stops at the agents. The control loop, routing, state transitions, and the event log are reproducible, but an LLM is still an LLM, so re-run a flow and an agent may reason its way to a different answer. (In tests, even that gap closes: the mock backend scripts the agent turns, so a scenario is deterministic end to end.) What you can always replay and audit in production is what the system did with the model’s answer, and the fact that the answer only changed state by passing through a deterministic handler. Swarm does not make the model deterministic; it makes everything around the model deterministic, which is the part worth auditing, replaying, and recovering. You accept a fuzzy model and wrap it in a system that is not. The cost is conceptual: real contracts to write, and a contract model to learn before complex flows pay off. Local setup is cheap — SQLite ships with the binary and the Quickstart runs with zero env vars. The payoff is everything above; the cost comes first.Ready to try it?
Boot the runtime and trigger your first run.

