Have 30-200 Employees? Make $50k-$500k selling your data for AI Training.

Learn More
Building Deterministic Multi-Agent State Machines in TypeScript
SitePoint Premium
Stay Relevant and Grow Your Career in Tech
  • Premium Results
  • Publish articles on SitePoint
  • Daily curated jobs
  • Learning Paths
  • Discounts to dev tools
Start Free Trial

7 Day Free Trial. Cancel Anytime.

Production multi-agent AI systems that rely on recursive prompt chains and fire-and-forget orchestration produce failure modes that are, by nature, non-deterministic. This article builds a deterministic, checkpoint-backed finite state machine engine in TypeScript that makes multi-agent workflows traceable and recoverable.

Table of Contents

Why Ad-Hoc Agent Chains Break in Production

Production multi-agent AI systems that rely on recursive prompt chains and fire-and-forget orchestration produce failure modes that are, by nature, non-deterministic. The symptoms are well-documented across the industry: silent state corruption when an intermediate agent returns an unexpected schema, impossible-to-reproduce bugs triggered by prompt variation or model temperature drift, and no clear recovery path after a crash mid-pipeline. When a four-agent workflow fails at step three, teams are left asking whether it is safe to retry from the beginning, whether partial results are still valid, and whether the same inputs will even produce the same failure again.

This article builds a deterministic, checkpoint-backed finite state machine engine in TypeScript that makes multi-agent workflows traceable and recoverable. The tech stack is deliberately minimal: TypeScript 5.x running on Node.js 18 LTS or later (recommended) or Bun, Zod for runtime schema validation, better-sqlite3 for synchronous transactional persistence, and architectural concepts borrowed from XState's statechart model. Note that better-sqlite3 requires native Node.js N-API bindings, so you will need build tools (node-gyp, Python 3.x) available in your environment. Every transition is pure. Every state change is atomically persisted. Every failure triggers a classified rollback. The result is a library under 500 lines of engine code with two runtime dependencies (zod and better-sqlite3) that runs in a single process and deploys to serverless environments (AWS Lambda, Fly.io, Bun-based servers). Edge runtimes that prohibit native Node.js bindings (Cloudflare Workers, Vercel Edge) are not supported. Developers get full control over exactly how their agents move through a pipeline.

Core Concepts: State Machines, DAGs, and Multi-Agent Orchestration

What Makes a State Machine "Deterministic" for Agents

A deterministic finite state machine guarantees a single invariant: given the same current state and the same incoming event, the machine always produces the same next state. There is no ambiguity, no probabilistic branching, and no hidden state influencing the outcome. In the context of AI agent orchestration, this means the FSM layer itself never introduces non-determinism. The LLM calls inside agent handlers produce different outputs on every invocation, but the orchestration topology, the transition logic, the checkpoint boundaries, and the recovery paths are all fixed and reproducible. Without this separation, a flaky LLM response changes the control flow graph, and you lose the ability to reason about whether the workflow itself is correct. It allows developers to reason about control flow independently of the stochastic behavior of language models.

A deterministic finite state machine guarantees a single invariant: given the same current state and the same incoming event, the machine always produces the same next state.

Contrast this with non-deterministic prompt chain recursion, where the output of one agent implicitly determines which agent runs next, often through string matching or ad-hoc conditional logic embedded in prompts. Such systems lack a formal transition table, making it impossible to enumerate valid states or prove that the workflow terminates.

Modeling Agent Workflows as Directed Acyclic Graphs

A directed acyclic graph enforces forward progress. Every edge points in one direction, and no path loops back to a previously visited node. This structure makes infinite loops impossible. In the FSM engine, each node in the DAG represents a distinct agent state (such as "summarize" or "classify"), and each directed edge represents a valid transition triggered by a specific event. Transitions between nodes are atomic: the engine either completes the full transition (including context update and checkpoint write) or it does not. There is no intermediate state visible to the rest of the system.

Where XState Ends and a Custom Engine Begins

XState provides a mature statechart implementation with hierarchical states, a visual inspector, and a well-tested ecosystem. For UI-driven state management and general-purpose statecharts, it remains an excellent choice. A custom engine solves three problems XState does not address well in this context. Atomic persistence hooks need to sit directly inside the transition cycle so that every state change is durably recorded before side effects execute; XState's persistence model is pluggable but not transactional by default. Agent-specific rollback semantics, including failure classification and automatic retry with exponential backoff, require control over the interpreter loop that XState's service model does not expose cleanly. And the runtime footprint matters for serverless deployment: a purpose-built engine with zero external state machine dependencies adds only the engine code and two packages to a Lambda bundle, compared to pulling in the full XState runtime.

Designing the State Chart Schema with Zod

Defining States, Events, and Transitions as Types

The core type model consists of four elements: AgentState (a string literal union representing valid states), AgentEvent (a discriminated union of events the machine can receive), Transition (a mapping from a state-event pair to a next state, optionally guarded), and StateMachineDefinition (the complete chart including initial state, terminal states, and the transition table). Zod schemas enforce the shape of this definition at the boundary, whether it arrives from a configuration file, an API payload, or a test fixture. Always construct AgentContext objects via AgentContextSchema.parse({...}) to ensure UUID and structural validation is enforced.

import { z } from "zod";

const TransitionSchema = z.object({
  event: z.string(),
  target: z.string(),
  guard: z.string().optional(),
});

const StateNodeSchema = z.object({
  type: z.enum(["initial", "intermediate", "terminal"]),
  transitions: z.array(TransitionSchema).default([]),
});

const StateMachineDefinitionSchema = z.object({
  id: z.string(),
  initialState: z.string(),
  states: z.record(z.string(), StateNodeSchema),
});

const AgentContextSchema = z.object({
  machineId: z.string().uuid(),
  input: z.unknown(),
  results: z.record(z.string(), z.unknown()).default({}),
  errorLog: z.array(z.string()).default([]),
});

type StateMachineDefinition = z.infer<typeof StateMachineDefinitionSchema>;
type AgentContext = z.infer<typeof AgentContextSchema>;
type StateNode = z.infer<typeof StateNodeSchema>;
type Transition = z.infer<typeof TransitionSchema>;

Compile-time inference via z.infer ensures that TypeScript's type checker stays synchronized with the runtime validation. If a configuration file omits a required field, Zod catches it at parse time with a descriptive error rather than allowing it to propagate as a silent type mismatch.

Validating DAG Integrity at Startup

Before the engine processes a single event, the transition map must be validated as a proper DAG. A topological sort detects cycles, and a reachability check from initialState identifies unreachable states. This validation runs once at startup and prevents an entire class of runtime errors.

function validateDAG(definition: StateMachineDefinition): void {
  const { states, initialState } = definition;
  const stateNames = Object.keys(states);
  const inDegree = new Map<string, number>(stateNames.map((s) => [s, 0]));
  const adjacency = new Map<string, string[]>();

  if (states[initialState]?.type !== "initial") {
    throw new Error(
      `initialState "${initialState}" must have type "initial", got "${states[initialState]?.type}"`
    );
  }

  for (const [name, node] of Object.entries(states)) {
    const targets = node.transitions.map((t) => t.target);
    adjacency.set(name, targets);
    for (const t of targets) {
      if (!inDegree.has(t)) throw new Error(`Transition target "${t}" is not a defined state`);
      inDegree.set(t, (inDegree.get(t) ?? 0) + 1);
    }
  }

  // Kahn's algorithm for cycle detection
  const queue = stateNames.filter((s) => inDegree.get(s) === 0);
  const sorted: string[] = [];
  while (queue.length > 0) {
    const node = queue.shift()!;
    sorted.push(node);
    for (const neighbor of adjacency.get(node) ?? []) {
      inDegree.set(neighbor, (inDegree.get(neighbor) ?? 0) - 1);
      if (inDegree.get(neighbor) === 0) queue.push(neighbor);
    }
  }

  if (sorted.length !== stateNames.length) throw new Error("Cycle detected in state machine definition");

  // Reachability check from initialState
  const reachable = new Set<string>();
  const dfsQueue = [initialState];
  while (dfsQueue.length) {
    const n = dfsQueue.pop()!;
    if (reachable.has(n)) continue;
    reachable.add(n);
    for (const nb of adjacency.get(n) ?? []) dfsQueue.push(nb);
  }
  for (const s of stateNames) {
    if (!reachable.has(s)) {
      // Allow unreachable states only if they are terminal (programmatic sinks
      // such as an error fallback reached via handleFailure, not via transitions).
      if (states[s].type !== "terminal") {
        throw new Error(
          `Non-terminal state "${s}" is unreachable from initialState "${initialState}"`
        );
      }
      console.warn(
        `Terminal state "${s}" is not reachable via transitions from "${initialState}". ` +
        `It may be used as a programmatic sink (e.g., error fallback).`
      );
    }
  }

  const terminals = stateNames.filter((s) => states[s].type === "terminal");
  if (terminals.length === 0) throw new Error("No terminal states defined");
}

The function throws descriptive errors at each failure point rather than returning a boolean. This makes misconfiguration immediately visible during startup, not buried in a log after the first failed transition.

Building the FSM Engine Core

The Transition Function: Pure and Predictable

The heart of the engine is a pure function. It takes the current state, an event, and the current context, then returns the next state and an updated context. It never mutates its inputs.

type MachineSnapshot = { state: string; context: AgentContext };
type GuardMap = Record<string, (context: AgentContext) => boolean>;

function transition(
  definition: StateMachineDefinition,
  snapshot: MachineSnapshot,
  event: string,
  guards: GuardMap = {}
): MachineSnapshot {
  const stateNode = definition.states[snapshot.state];
  if (!stateNode) throw new Error(`Unknown state: "${snapshot.state}"`);

  const candidates = stateNode.transitions.filter((t) => t.event === event);
  if (candidates.length === 0) {
    throw new Error(`No transition for event "${event}" in state "${snapshot.state}"`);
  }

  for (const candidate of candidates) {
    if (candidate.guard) {
      const guardFn = guards[candidate.guard];
      if (!guardFn) throw new Error(`Guard "${candidate.guard}" is not registered`);
      if (!guardFn(snapshot.context)) continue;
    }
    return {
      state: candidate.target,
      context: structuredClone(snapshot.context),
    };
  }

  throw new Error(
    `All guards failed for event "${event}" in state "${snapshot.state}"`
  );
}

Guard conditions allow conditional branching without introducing non-determinism. For a given context value, the same guard always returns the same boolean. The function iterates candidates in definition order, which makes priority explicit and inspectable.

Timeout Utility for Action Handlers

A hung LLM call can stall the interpreter loop indefinitely. The following utility wraps any promise with a configurable timeout, ensuring the engine always makes progress or fails explicitly.

function withTimeout<T>(promise: Promise<T>, ms: number): Promise<T> {
  return Promise.race([
    promise,
    new Promise<never>((_, reject) =>
      setTimeout(
        () => reject(new Error(`timeout: handler exceeded ${ms}ms`)),
        ms
      )
    ),
  ]);
}

The Interpreter Loop

The runtime interpreter consumes events from a queue, calls the pure transition() function, and executes side effects only after the transition succeeds and the checkpoint is written. The send() method enqueues events synchronously; run() must be awaited to process the queue.

type ActionHandler = (context: AgentContext) => Promise<{ event: string; updatedResults: Record<string, unknown> }>;

class Interpreter {
  protected queue: string[] = [];
  protected snapshot: MachineSnapshot;
  protected running = false;
  protected definition: StateMachineDefinition;
  protected actions: Map<string, ActionHandler>;
  protected guards: GuardMap;

  constructor(
    definition: StateMachineDefinition,
    actions: Map<string, ActionHandler>,
    guards: GuardMap,
    protected checkpointStore: CheckpointStore,
    initialContext: AgentContext
  ) {
    this.definition = definition;
    this.actions = actions;
    this.guards = guards;
    const loaded = checkpointStore.load(initialContext.machineId);
    this.snapshot = loaded ?? { state: definition.initialState, context: initialContext };
  }

  send(event: string): void {
    this.queue.push(event);
  }

  async run(): Promise<void> {
    if (!this.running) await this.processQueue();
  }

  async start(initialEvent: string): Promise<void> {
    this.send(initialEvent);
    await this.run();
  }

  protected async processSingleEvent(event: string): Promise<void> {
    const next = transition(this.definition, this.snapshot, event, this.guards);
    this.checkpointStore.save(next, event);
    this.snapshot = next;

    const handler = this.actions.get(next.state);
    if (handler && this.definition.states[next.state].type !== "terminal") {
      const result = await withTimeout(handler(this.snapshot.context), 30_000);
      this.snapshot = {
        ...this.snapshot,
        context: AgentContextSchema.parse({
          ...this.snapshot.context,
          results: { ...this.snapshot.context.results, ...result.updatedResults },
        }),
      };
      this.checkpointStore.save(this.snapshot, `result:${next.state}`);
      this.queue.push(result.event);
    }
  }

  private async processQueue(): Promise<void> {
    this.running = true;
    while (this.queue.length > 0) {
      const event = this.queue.shift()!;
      try {
        await this.processSingleEvent(event);
      } catch (err) {
        await this.handleFailure(err as Error, event);
      }
    }
    this.running = false;
  }

  protected async handleFailure(err: Error, event: string): Promise<void> {
    console.error(`Failure in state "${this.snapshot.state}" on event "${event}":`, err.message);
  }
}

The loop follows a strict sequence: dequeue, validate via transition(), persist the checkpoint, then execute the action handler. This ordering guarantees that the checkpoint always reflects the last successfully entered state, not a state whose side effects have not yet completed. Note that send() only enqueues; callers must await run() (or use start()) to process events, ensuring errors propagate to the caller rather than becoming unhandled rejections. The processSingleEvent method is protected so that subclasses (such as RollbackInterpreter) can invoke it directly during retry, avoiding the this.running guard that would otherwise block re-entry through run().

Registering Agent Action Handlers

Each state node maps to an async handler function that performs the actual agent work, such as calling an LLM API, validating output, or transforming data. Handlers receive typed context and return an event that feeds back into the machine.

type ActionRegistry = Map<string, ActionHandler>;

const agentHandlers: ActionRegistry = new Map([
  ["ingest", async (_ctx) => {
    return { event: "INGESTED", updatedResults: { ingestedAt: Date.now() } };
  }],
  ["summarize", async (_ctx) => {
    // Simulate LLM summarization call
    const summary = `Summary of ${JSON.stringify(_ctx.input).slice(0, 50)}...`;
    return { event: "SUMMARIZED", updatedResults: { summary } };
  }],
  ["classify", async (_ctx) => {
    const classification = "category_a";
    return { event: "CLASSIFIED", updatedResults: { classification } };
  }],
  ["review", async (_ctx) => {
    const approved = true;
    return { event: approved ? "APPROVED" : "REJECTED", updatedResults: { approved } };
  }],
]);

These stubs simulate LLM calls. In production, each handler would wrap a real API call with timeout handling and response validation. The key design constraint: handlers return events, not states. The transition table, not the handler, decides where the machine goes next.

Atomic Checkpointing with SQLite

Why Checkpoints Must Be Atomic

Consider the failure scenario where the engine crashes after completing a state transition but before the action handler finishes executing. Without an atomic checkpoint, the engine lost the record of the new state, so on restart it replays the previous action. If that action is not idempotent (for example, sending an email or charging a credit card), the result is duplication. Alternatively, if the checkpoint writes but the transition metadata does not, the system enters an inconsistent state with no valid recovery path.

The checkpoint contract is strict: the engine transitions, updates the context, and writes the checkpoint in a single SQLite transaction. Either all three succeed or none do.

The checkpoint contract is strict: the engine transitions, updates the context, and writes the checkpoint in a single SQLite transaction. Either all three succeed or none do.

Implementing the Checkpoint Store

better-sqlite3 provides synchronous, transactional writes. This is ideal for single-process agent runners because there is no async gap between "decide to write" and "write committed." Note that better-sqlite3 requires native Node.js bindings and will not run in V8-isolate-based edge runtimes.

import Database from "better-sqlite3";

class CheckpointStore {
  private db: Database.Database;

  constructor(dbPath: string) {
    // Use ':memory:' for tests only; provide a file path in production.
    this.db = new Database(dbPath, { timeout: 5000 });
    this.init();
  }

  private init(): void {
    this.db.pragma("journal_mode = WAL");
    this.db.exec(`
      CREATE TABLE IF NOT EXISTS checkpoints (
        id INTEGER PRIMARY KEY AUTOINCREMENT,
        machine_id TEXT NOT NULL,
        state TEXT NOT NULL,
        context TEXT NOT NULL,
        event_log TEXT NOT NULL DEFAULT '[]',
        timestamp INTEGER NOT NULL
      )
    `);
    this.db.exec(`
      CREATE INDEX IF NOT EXISTS idx_checkpoints_machine_id
      ON checkpoints (machine_id, id)
    `);
  }

  save(snapshot: MachineSnapshot, event: string): void {
    const txn = this.db.transaction(() => {
      // Single query retrieves the event log atomically within the transaction.
      const existing = this.db
        .prepare(
          `SELECT event_log FROM checkpoints
           WHERE machine_id = ? ORDER BY id DESC LIMIT 1`
        )
        .get(snapshot.context.machineId) as { event_log: string } | undefined;

      const eventLog: string[] = existing
        ? [...JSON.parse(existing.event_log), event]
        : [event];

      this.db
        .prepare(
          `INSERT INTO checkpoints (machine_id, state, context, event_log, timestamp)
           VALUES (?, ?, ?, ?, ?)`
        )
        .run(
          snapshot.context.machineId,
          snapshot.state,
          JSON.stringify(snapshot.context),
          JSON.stringify(eventLog),
          Date.now()
        );
    });
    txn();
  }

  load(machineId: string): MachineSnapshot | null {
    const row = this.db
      .prepare(
        `SELECT id, state, context FROM checkpoints
         WHERE machine_id = ? ORDER BY id DESC LIMIT 1`
      )
      .get(machineId) as { id: number; state: string; context: string } | undefined;
    if (!row) return null;
    return {
      state: row.state,
      context: AgentContextSchema.parse(JSON.parse(row.context)),
    };
  }

  rollbackTo(machineId: string, checkpointId: number): MachineSnapshot | null {
    // Warning: this permanently deletes all checkpoints after the given checkpoint id.
    // Uses the AUTOINCREMENT id as the stable, unique boundary — not timestamp,
    // which can produce duplicates at millisecond resolution under fast writes.
    this.db
      .prepare(`DELETE FROM checkpoints WHERE machine_id = ? AND id > ?`)
      .run(machineId, checkpointId);
    return this.load(machineId);
  }

  close(): void {
    this.db.close();
  }
}

The entire save() method, including the event log read and the insert, is wrapped in a single transaction, ensuring atomicity. The id INTEGER PRIMARY KEY AUTOINCREMENT column is used as the deletion boundary in rollbackTo rather than Date.now() timestamps, which can produce duplicates under fast writes. WAL journal mode is enabled at initialization for improved concurrent read performance.

Plugging Checkpoints into the Interpreter

The integration point sits between transition resolution and action handler execution. The processSingleEvent() method (shown in the Interpreter class above) calls this.checkpointStore.save(next, event) immediately after transition() returns and before the handler runs. A second checkpoint write occurs after the handler returns its results, capturing the enriched context. This two-phase checkpointing means a crash during handler execution loses only the handler's output, not the transition itself, and the engine can safely re-enter the state on restart. Note that full atomicity of the two-phase sequence depends on both writes landing in their respective transactions; the transition checkpoint is guaranteed, and re-entry on crash replays only the handler.

Automatic Rollback and Recovery

Detecting and Classifying Failures

Failures fall into two categories. Transient failures include network timeouts, API rate limits, and temporary service unavailability; the engine can retry these. Permanent failures include invalid response schemas, logic errors in guard conditions, and transitions to undefined states; these require human intervention or a dedicated error state. The engine tags each caught error with a classification based on its type and message pattern.

Implementing Rollback Handlers

type RetryPolicy = { maxRetries: number; baseDelayMs: number; backoffMultiplier: number };

const defaultRetryPolicy: RetryPolicy = { maxRetries: 3, baseDelayMs: 500, backoffMultiplier: 2 };

class RollbackInterpreter extends Interpreter {
  private retryPolicy: RetryPolicy;

  constructor(
    definition: StateMachineDefinition, actions: Map<string, ActionHandler>,
    guards: GuardMap, checkpointStore: CheckpointStore,
    initialContext: AgentContext, retryPolicy: RetryPolicy = defaultRetryPolicy
  ) {
    super(definition, actions, guards, checkpointStore, initialContext);
    this.retryPolicy = retryPolicy;
  }

  protected override async handleFailure(err: Error, event: string): Promise<void> {
    const isTransient = this.classifyError(err) === "transient";
    if (isTransient) {
      let lastErr: Error = err;
      for (let attempt = 1; attempt <= this.retryPolicy.maxRetries; attempt++) {
        const delay = this.retryPolicy.baseDelayMs * Math.pow(this.retryPolicy.backoffMultiplier, attempt - 1);
        await new Promise((r) => setTimeout(r, delay));
        console.info(`Retry attempt ${attempt} for event "${event}" after ${delay}ms`);
        try {
          // Reload the last persisted checkpoint before retrying,
          // ensuring we retry against known-good state, not partially mutated in-memory state.
          const checkpoint = this.checkpointStore.load(this.snapshot.context.machineId);
          if (checkpoint) this.snapshot = checkpoint;
          // Directly process this single event without going through run()
          // to avoid the this.running guard blocking re-entry.
          await this.processSingleEvent(event);
          return;
        } catch (retryErr) {
          lastErr = retryErr as Error;
          console.warn(`Retry ${attempt} failed for event "${event}":`, (retryErr as Error).message);
        }
      }
      console.error(`Retries exhausted for event "${event}":`, lastErr.message);
    }
    // Permanent failure or retries exhausted: transition to error state
    this.applyErrorState(err);
  }

  private applyErrorState(err: Error): void {
    const errorStateName = "error";
    if (!this.definition.states[errorStateName]) {
      // Cannot transition to error state — log and rethrow to surface the problem.
      console.error(
        `Permanent failure but no "${errorStateName}" state defined in machine "${this.definition.id}".`,
        err.message
      );
      throw new Error(`Unrecoverable: ${err.message}`);
    }
    this.snapshot = {
      state: errorStateName,
      context: {
        ...this.snapshot.context,
        errorLog: [...this.snapshot.context.errorLog, err.message],
      },
    };
    this.checkpointStore.save(this.snapshot, "ERROR");
  }

  private classifyError(err: Error): "transient" | "permanent" {
    const transientPatterns = [/\bETIMEDOUT\b/, /\bECONNRESET\b/, /\b429\b/, /rate limit/i, /timeout/i];
    return transientPatterns.some((p) => p.test(err.message))
      ? "transient" : "permanent";
  }
}

On transient failure, the engine reloads the last checkpoint (already persisted before the handler ran), restoring the snapshot to known-good state, then directly processes the failed event via processSingleEvent() — bypassing the run() guard that would otherwise block re-entry while the interpreter loop is active. This ensures retries actually execute. On permanent failure or after exhausting retries, the machine transitions to a dedicated error terminal state (after verifying it exists in the definition), preserving the full event log and error message for debugging. The classifyError method uses word-boundary regex patterns rather than substring matching to avoid false positives (e.g., a message like "Step 4291 failed" will not falsely match the 429 rate-limit pattern).

Replaying Event Logs for Debugging

The persisted event log enables full replay. To reproduce a failure, load the initial checkpoint for a given machine_id, feed the recorded events sequentially through the pure transition() function, and compare the output state at each step to the checkpointed state. Because transition() is pure and deterministic, any divergence means either the code changed between the original run and the replay, or the checkpoint data is corrupt. This replay capability requires that the event log is written atomically (as ensured by the transactional save() method) and is one of the most significant operational advantages over ad-hoc orchestration.

Putting It All Together: A Multi-Agent Pipeline

A realistic four-agent pipeline moves through five states: ingest, summarize, classify, review, and done, with an error terminal state reachable from any non-terminal state.

import { randomUUID } from "crypto";

const pipelineDefinition: StateMachineDefinition = {
  id: "doc-pipeline",
  initialState: "ingest",
  states: {
    ingest:     { type: "initial",       transitions: [{ event: "INGESTED", target: "summarize" }] },
    summarize:  { type: "intermediate",  transitions: [{ event: "SUMMARIZED", target: "classify" }] },
    classify:   { type: "intermediate",  transitions: [{ event: "CLASSIFIED", target: "review" }] },
    review:     { type: "intermediate",  transitions: [
      { event: "APPROVED", target: "done" },
      { event: "REJECTED", target: "error" },
    ]},
    done:       { type: "terminal",      transitions: [] },
    error:      { type: "terminal",      transitions: [] },
  },
};

validateDAG(pipelineDefinition);

const store = new CheckpointStore("./pipeline.db");
const context: AgentContext = AgentContextSchema.parse({
  machineId: randomUUID(),
  input: { document: "Quarterly earnings report full text..." },
  results: {},
  errorLog: [],
});

const interpreter = new RollbackInterpreter(
  pipelineDefinition, agentHandlers, {}, store, context
);

// Kick off the pipeline
interpreter.start("INGESTED").then(() => {
  const final = store.load(context.machineId);
  console.log("Final state:", final?.state);
  console.log("Results:", final?.context.results);
  store.close();
});

On a success path, the output shows the machine moving from ingest through summarize, classify, and review to done, with each agent's results accumulated in the context. On a failure path (for example, a simulated timeout in the classify handler), the engine logs the retry attempts, and either recovers to continue the pipeline or transitions to error with the full event history preserved.

The code in this article is self-contained. To run the example, copy all code blocks into src/main.ts in definition order (schemas, validators, timeout utility, engine, checkpoint, rollback, handlers, pipeline), then run with npx tsx src/main.ts (install globally first with npm i -g tsx if needed) or bun run src/main.ts.

Prerequisites

  • Node.js 18 LTS or later (recommended; randomUUID requires 14.17.0 or later)
  • Build tools for native dependencies: node-gyp, Python 3.x (required by better-sqlite3)
  • Packages: npm install typescript zod better-sqlite3 && npm install -D @types/better-sqlite3 @types/node tsx
  • tsconfig.json: "strict": true, "target": "ES2020", "module": "CommonJS"

Testing Determinism: Property-Based and Snapshot Tests

Snapshot Testing Transitions

Vitest or Jest snapshot tests assert that the same (state, event) pair always produces the same next state. For every defined transition in the state chart, a snapshot test calls transition() with fixed inputs and compares the serialized output to a stored snapshot. If a code change alters the transition result, the test fails explicitly, preventing accidental non-determinism from reaching production.

Property-Based Testing with fast-check

The fast-check library (3.x or later) generates random event sequences from the set of valid events and feeds them to the interpreter. The property under test: the machine never enters an undefined state, and every terminal state has an empty transition list. This catches edge cases that handwritten tests miss, such as unusual event orderings that expose missing transitions or guard logic errors.

When to Use This vs. XState, Temporal, or LangGraph

Feature This Engine XState Temporal LangGraph
Persistence model Atomic SQLite checkpoints Pluggable (no built-in transactions) Durable execution (server-side) In-memory or custom
Designed for LLM agent orchestration Yes, handlers return events General-purpose; strongest ecosystem for UI-driven statecharts No, general workflow Yes, LangChain-coupled
Runtime weight 2 runtime deps (zod, better-sqlite3), no framework runtime ~15 npm packages in dependency tree Requires Temporal server cluster (self-hosted) or Temporal Cloud ~10+ packages, inherits LangChain dependency tree
Rollback support Automatic with classification Manual Built-in (activity retry) Manual
TypeScript-first typing Full Zod + inference Full TypeScript SDK available TypeScript support via LangGraph.js (Python SDK has broader feature coverage)

Use XState for statecharts where visualization and hierarchical states matter. Use Temporal for distributed, long-running workflows that span multiple services and require durable execution guarantees beyond a single process. Use LangGraph when tightly integrated with the LangChain ecosystem and willing to accept its runtime assumptions. Use this engine for deterministic multi-agent orchestration where full checkpoint control, a small dependency footprint, and single-process deployment are priorities.

Handlers return events, not states. The transition table, not the handler, decides where the machine goes next.

Next Steps

The architecture gives you deterministic transitions, atomic checkpoints, automatic classified rollback, and full type safety from schema to runtime. The most immediate extension is replacing better-sqlite3 with Turso or libSQL for distributed checkpoint storage; the migration requires swapping synchronous Database calls for the async @libsql/client driver and converting save()/load() to async methods. Beyond that, consider adding parallel branch execution within the DAG for independent agent stages, and building a visual state chart editor that reads the Zod-validated definition. Run the example pipeline and adapt it to your own agent workflows.

SitePoint TeamSitePoint Team

Sharing our passion for building incredible internet things.

© 2000 – 2026 SitePoint Pty. Ltd.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.