Have 30-200 Employees? Make $50k-$500k selling your data for AI Training.

Learn More
Building Production MCP Servers in TypeScript
SitePoint Premium
Stay Relevant and Grow Your Career in Tech
  • Premium Results
  • Publish articles on SitePoint
  • Daily curated jobs
  • Learning Paths
  • Discounts to dev tools
Start Free Trial

7 Day Free Trial. Cancel Anytime.

How to Build a Production MCP Server in TypeScript

  1. Define an explicit lifecycle state machine (Initializing → Ready → Active → Draining → Terminated) with guard clauses that reject invalid transitions.
  2. Scaffold a TypeScript 5.x project with strict, noUncheckedIndexedAccess, and exactOptionalPropertyTypes enabled.
  3. Co-locate Zod schemas with every tool registration using a defineTool wrapper that validates input/output and returns structured MCP errors.
  4. Implement a session pool with TTL-based eviction, max-session limits, and token-indexed reconnect support.
  5. Wire an AbortController per tool invocation to session disconnect and server drain events for deterministic cleanup.
  6. Configure graceful shutdown by handling SIGTERM/SIGINT, draining active sessions, and awaiting in-flight tool completions.
  7. Test with a 50-agent concurrent stress suite in Vitest that verifies zero leaked sessions and structured error responses for malformed input.
  8. Deploy as a long-lived process behind a load balancer with a health-check endpoint that reports the current lifecycle state.

MCP is the protocol AI agents use to call external tools. Most tutorials stop at a registered tool and a working stdio transport. Building production MCP servers in TypeScript demands a different set of concerns: explicit state management across concurrent agent sessions, resilient connection handling that survives network interruptions, and schema validation that prevents silent data corruption from propagating through agent workflows. This article provides a deployable architectural pattern that addresses all three.

Table of Contents

Why Production MCP Servers Are Different from Demos

The Gap Between SDK Quickstarts and Real Workloads

The @modelcontextprotocol/sdk package provides a tool() registration API and transport wiring. What it does not do is manage the lifecycle complexity that surfaces when dozens of concurrent agents connect over hours. Several failure modes appear quickly in production that never show up in single-session development:

When an agent disconnects unexpectedly (network timeout, process crash, user abort), the server retains the session object in memory. Without explicit cleanup, these accumulate indefinitely, consuming heap space and holding references to tool state that no consumer will ever read.

Tools that produce intermediate results, such as document retrieval pipelines or code generation steps, hold large objects on the heap through closures. Without AbortController integration and explicit nullification, a server processing hundreds of tool invocations per hour will show RSS growing on the order of 50-200 MB/hour depending on tool output size. Measure yours with process.memoryUsage().rss.

JSON Schema validation at the protocol level catches structural errors but does not coerce types at parse time, narrow discriminated unions, or infer TypeScript types. An agent sending "42" where a tool expects 42 produces a silent type mismatch that propagates downstream rather than triggering a clean rejection.

Without session token reuse, a reconnecting agent starts a fresh session, losing accumulated context and potentially duplicating side effects from previously completed tool calls.

Architecture Overview: Stateful vs. Stateless MCP Server Lifecycles

When Statelessness Is Enough (and When It Isn't)

Not every MCP server needs the full machinery described here. The decision framework is direct: if every tool invocation is idempotent and requires no context from previous invocations, a stateless design works. A currency conversion tool or a simple API proxy falls into this category. Each request stands alone; the server ignores session identity.

Stateful servers become necessary when tools involve multi-step workflows (a code execution sandbox that accumulates files across invocations) or external resource handles like database connections, file locks, and streaming cursors. Accumulated context that agents expect to persist across a session also forces statefulness. In these cases, the server must track what state each session holds and clean it up deterministically.

The Lifecycle State Machine

Rather than scattering boolean flags like isReady, isShuttingDown, and hasConnections across a server implementation, a production MCP server benefits from an explicit state machine with five states: Initializing, Ready, Active, Draining, and Terminated.

Initializing covers startup: loading configuration, establishing database pools, registering tools. Ready indicates the server can accept connections but has none yet. Active means at least one session is connected. Draining is triggered by a shutdown signal (SIGTERM, SIGINT) and means the server stops accepting new connections while signaling existing sessions to finish and disconnect. Terminated is the final state after all sessions have been released and resources freed.

Guard clauses on each transition method reject invalid state jumps. A server in Terminated cannot move back to Active. A server in Initializing cannot jump directly to Draining. These precondition checks eliminate an entire class of concurrency bugs where race conditions between incoming connections and shutdown signals produce undefined behavior.

Guard clauses on each transition method reject invalid state jumps. A server in Terminated cannot move back to Active. These precondition checks eliminate an entire class of concurrency bugs where race conditions between incoming connections and shutdown signals produce undefined behavior.

import { EventEmitter } from "events";

enum ServerState {
  Initializing = "Initializing",
  Ready = "Ready",
  Active = "Active",
  Draining = "Draining",
  Terminated = "Terminated",
}

type StatePayload =
  | { state: ServerState.Initializing }
  | { state: ServerState.Ready }
  | { state: ServerState.Active; sessionCount: number }
  | { state: ServerState.Draining; remainingSessions: number }
  | { state: ServerState.Terminated; reason: string };

const VALID_TRANSITIONS: Record<ServerState, ServerState[]> = {
  [ServerState.Initializing]: [ServerState.Ready],
  [ServerState.Ready]: [ServerState.Active, ServerState.Draining],
  [ServerState.Active]: [ServerState.Ready, ServerState.Draining],
  [ServerState.Draining]: [ServerState.Terminated],
  [ServerState.Terminated]: [],
};

class ServerLifecycle extends EventEmitter {
  private current: StatePayload = { state: ServerState.Initializing };

  get state(): ServerState {
    return this.current.state;
  }

  transition(next: StatePayload): void {
    const allowed = VALID_TRANSITIONS[this.current.state];
    if (!allowed.includes(next.state)) {
      throw new Error(
        `Invalid transition: ${this.current.state} → ${next.state}`
      );
    }
    const previous = this.current.state;
    this.current = next;
    this.emit("transition", { from: previous, to: next });
  }
}

export { ServerLifecycle, ServerState, StatePayload };

The discriminated union StatePayload ensures that state-specific data (such as sessionCount in the Active state or reason in the Terminated state) is only accessible when the server is actually in that state. TypeScript's narrowing catches misuse at compile time.

Project Scaffolding and Tech Stack Setup

Prerequisites

  • Node.js 18.x or later (Node 20 LTS recommended). The code uses crypto.randomUUID(), which is available globally only in Node 19+. For Node 18 LTS, the explicit import { randomUUID } from 'crypto' shown below is required.
  • npm, yarn, or pnpm as your package manager. Commit your lockfile (package-lock.json, yarn.lock, or pnpm-lock.yaml) to ensure reproducible installs — the caret ranges below are illustrative.
  • TypeScript 5.x (5.8+ recommended).

Initializing the TypeScript 5.x Project

The architecture relies on TypeScript's strictest compiler settings to catch potential issues before runtime. Three settings beyond the standard strict flag are particularly important: noUncheckedIndexedAccess prevents unsafe property access on index-signature and array-index types (critical for the session pool — note that Map.get() already returns T | undefined unconditionally), exactOptionalPropertyTypes distinguishes between undefined and missing properties in tool schemas, and strict itself enables strictNullChecks and strictFunctionTypes that the Zod integration depends on.

Note on module: "Node16": This module mode requires "moduleResolution": "Node16" in your tsconfig, and all relative imports in your source files must use .js extensions (e.g., import { SessionPool } from "./session-pool.js"). The SDK import @modelcontextprotocol/sdk/server/mcp.js already follows this convention.

project-root/
├── src/
│   ├── lifecycle.ts        # ServerLifecycle state machine
│   ├── schema.ts           # defineTool wrapper and Zod utilities
│   ├── session-pool.ts     # SessionPool with TTL eviction
│   ├── tool-execution.ts   # ToolExecution class with AbortController
│   ├── tools/
│   │   └── search.ts       # Example tool definitions
│   └── server.ts           # Entry point and transport setup
├── tests/
│   └── concurrent.test.ts  # Vitest stress tests
├── vitest.config.ts
├── package.json
└── tsconfig.json
// tsconfig.json
{
  "compilerOptions": {
    "target": "ES2022",
    "module": "Node16",
    "moduleResolution": "Node16",
    "strict": true,
    "noUncheckedIndexedAccess": true,
    "exactOptionalPropertyTypes": true,
    "outDir": "./dist"
  },
  "include": ["src"]
}
// package.json (relevant dependencies)
{
  "dependencies": {
    "@modelcontextprotocol/sdk": "^1.12.1",
    "zod": "^3.25.17"
  },
  "devDependencies": {
    "vitest": "^3.2.1",
    "typescript": "^5.8.3"
  }
}

Reproducibility note: The caret ranges above allow minor and patch upgrades. For strict reproducibility, pin exact versions or rely on your committed lockfile.

Zod-Based Tool Schema Validation

Defining Tool Input and Output Contracts with Zod

The MCP SDK uses JSON Schema for tool input definitions at the protocol level, which handles structural validation but lacks several capabilities needed for production safety. JSON Schema does not coerce types at parse time (converting a string "42" to a number 42 based on the schema type). It does not support discriminated unions, a common pattern when a tool accepts multiple input shapes depending on the operation. And it provides no mechanism for TypeScript type inference, meaning developers must manually maintain parallel type definitions that can drift from the schema.

Zod solves all three. By co-locating Zod schemas with tool registrations, the schema becomes the single source of truth for both runtime validation and compile-time TypeScript types. The z.infer utility extracts the type directly from the schema, eliminating type drift entirely. The JSON Schema registered with the MCP protocol layer must also accurately reflect the Zod types — see the type-mapping logic in defineTool below.

McpServer is the high-level server class from @modelcontextprotocol/sdk/server/mcp.js that provides the .tool() registration API.

Wrapping the MCP SDK's tool() with Schema Enforcement

The core pattern is a higher-order function, defineTool, that accepts a Zod input schema, a Zod output schema, and a handler function. It returns a fully typed, validated MCP tool registration. Inside the wrapper, .safeParse() validates the incoming arguments and, on failure, returns a structured MCP-compliant error response with field-level error details rather than crashing the server with an unhandled exception.

This distinction matters: an agent receiving a structured error like { isError: true, content: [{ type: "text", text: "Validation failed: query must be at least 1 character" }] } can retry with corrected input. A server crash gives the agent nothing to work with.

Output validation caveat: The output validation path below returns isError: true to the agent when the server's own output fails validation. This is a server-side defect, not an agent input error — the agent cannot fix it. In production, consider logging the validation failure and throwing an internal error instead of returning isError: true, so the defect surfaces in server monitoring rather than being silently swallowed by the agent.

An agent receiving a structured error like { isError: true, content: [{ type: "text", text: "Validation failed: query must be at least 1 character" }] } can retry with corrected input. A server crash gives the agent nothing to work with.

import { z, ZodSchema, ZodError } from "zod";
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";

function formatZodError(error: ZodError): string {
  return error.issues
    .map((issue) => {
      const path = issue.path.length > 0 ? issue.path.join(".") : "root";
      return `${path}: ${issue.message}`;
    })
    .join("; ");
}

/**
 * Maps a Zod type to its JSON Schema `type` string.
 * For production use, consider the `zod-to-json-schema` package
 * for full support of unions, enums, arrays, and nested objects.
 */
function zodTypeToJsonType(zodType: z.ZodTypeAny): string {
  if (zodType instanceof z.ZodString) return "string";
  if (zodType instanceof z.ZodNumber) return "number";
  if (zodType instanceof z.ZodBoolean) return "boolean";
  if (zodType instanceof z.ZodArray) return "array";
  if (zodType instanceof z.ZodObject) return "object";
  if (zodType instanceof z.ZodDefault) return zodTypeToJsonType(zodType._def.innerType);
  if (zodType instanceof z.ZodOptional) return zodTypeToJsonType(zodType.unwrap());
  if (zodType instanceof z.ZodNullable) return zodTypeToJsonType(zodType.unwrap());
  const typeName = (zodType._def as { typeName?: string }).typeName ?? "unknown";
  console.warn(
    `[defineTool] Unsupported Zod type "${typeName}" in JSON Schema mapping. ` +
    `Use zod-to-json-schema for full type support.`
  );
  return "string";
}

function defineTool<TInput extends ZodSchema, TOutput extends ZodSchema>(
  server: McpServer,
  name: string,
  description: string,
  inputSchema: TInput,
  outputSchema: TOutput,
  handler: (input: z.infer<TInput>) => Promise<z.infer<TOutput>>
): void {
  const jsonSchemaShape: Record<string, { type: string; description?: string }> = {};
  if (inputSchema instanceof z.ZodObject) {
    for (const [key, value] of Object.entries(inputSchema.shape)) {
      jsonSchemaShape[key] = { type: zodTypeToJsonType(value as z.ZodTypeAny) };
    }
  }

  server.tool(name, description, jsonSchemaShape, async (args) => {
    const inputResult = inputSchema.safeParse(args);
    if (!inputResult.success) {
      return {
        isError: true,
        content: [
          { type: "text", text: `Validation failed: ${formatZodError(inputResult.error)}` },
        ],
      };
    }

    let output: z.infer<TOutput>;
    try {
      output = await handler(inputResult.data);
    } catch (err) {
      return {
        isError: true,
        content: [
          { type: "text", text: `Tool execution failed: ${err instanceof Error ? err.message : "internal error"}` },
        ],
      };
    }

    const outputResult = outputSchema.safeParse(output);
    if (!outputResult.success) {
      return {
        isError: true,
        content: [
          { type: "text", text: `Output validation failed: ${formatZodError(outputResult.error)}` },
        ],
      };
    }

    return { content: [{ type: "text", text: JSON.stringify(outputResult.data) }] };
  });
}

// Example tool: searchDocuments
const SearchInput = z.object({
  query: z.string().min(1),
  maxResults: z.number().int().positive().default(10),
});

const SearchOutput = z.object({
  results: z.array(z.object({ title: z.string(), snippet: z.string() })),
  totalCount: z.number(),
});

// Usage: defineTool(server, "searchDocuments", "Search the doc index", SearchInput, SearchOutput, async (input) => { ... });

export { defineTool, SearchInput, SearchOutput };

Session Connection Pooling and Reconnect Handling

Why MCP Sessions Leak Under Real Load

The default behavior of the MCP SDK is to create a session object for each incoming agent connection. When an agent disconnects cleanly, the session can be torn down. But disconnections are frequently unclean: network timeouts, agent process crashes, or user aborts that never send a proper close frame. In these cases, the server retains the session object in memory with all its associated tool state, intermediate results, and event listeners.

This is not hypothetical. A server handling 50 agents where roughly 10% disconnect and reconnect each minute can accumulate hundreds of orphaned sessions over a few hours. Heap snapshots of such servers show session objects that no code path will access again dominating memory, each holding references to closures, buffers, and cached tool results.

Building a Session Pool with TTL and Health Checks

A SessionPool class wraps session management with three mechanisms: a configurable maximum session count that rejects new connections when the pool is full, TTL-based eviction that removes sessions idle beyond a threshold, and periodic health checks that detect dead connections.

The pool integrates with the lifecycle state machine: during the Draining state, the pool signals all active sessions to cancel in-flight work (via their AbortControllers — see integration note below) and waits for them to finish before releasing. During Terminated, the pool rejects all new session acquisitions unconditionally.

Note on token security: In production, session tokens should be cryptographically random and validated for format and length before lookup, to prevent enumeration attacks.

import { randomUUID } from "crypto";

interface PooledSession {
  id: string;
  token: string;
  lastActivity: number;
  data: Map<string, unknown>;
}

class SessionPool {
  private sessions = new Map<string, PooledSession>();
  private tokenIndex = new Map<string, string>(); // token → session id
  private readonly maxSessions: number;
  private readonly ttlMs: number;
  private evictionTimer: ReturnType<typeof setInterval> | null = null;
  private isDraining = false;

  private static readonly TOKEN_MAX_LENGTH = 256;
  private static readonly TOKEN_PATTERN = /^[\w\-]{1,256}$/;

  constructor(maxSessions = 100, ttlMs = 300_000) {
    this.maxSessions = maxSessions;
    this.ttlMs = ttlMs;
    this.evictionTimer = setInterval(() => {
      try {
        this.evictStale();
      } catch (err) {
        console.error({ event: "eviction_error", error: String(err) });
      }
    }, 60_000);
    this.evictionTimer.unref();
  }

  acquire(token: string): PooledSession {
    if (this.isDraining) {
      throw new Error("Session pool is draining; not accepting new sessions");
    }
    if (
      typeof token !== "string" ||
      token.length === 0 ||
      token.length > SessionPool.TOKEN_MAX_LENGTH ||
      !SessionPool.TOKEN_PATTERN.test(token)
    ) {
      throw new Error(
        `Invalid session token: must be 1–${SessionPool.TOKEN_MAX_LENGTH} alphanumeric/dash/underscore characters`
      );
    }
    const existingId = this.tokenIndex.get(token);
    if (existingId !== undefined) {
      const existing = this.sessions.get(existingId);
      if (existing !== undefined) {
        existing.lastActivity = Date.now();
        return existing;
      }
      // Index is stale — clean it up
      this.tokenIndex.delete(token);
    }
    if (this.sessions.size >= this.maxSessions) {
      throw new Error("Session pool exhausted");
    }
    const session: PooledSession = {
      id: randomUUID(),
      token,
      lastActivity: Date.now(),
      data: new Map(),
    };
    this.sessions.set(session.id, session);
    this.tokenIndex.set(token, session.id);
    return session;
  }

  release(sessionId: string): void {
    const session = this.sessions.get(sessionId);
    if (session) {
      session.data.clear();
      this.tokenIndex.delete(session.token);
      this.sessions.delete(sessionId);
    }
  }

  evictStale(): number {
    const now = Date.now();
    const staleIds: string[] = [];
    for (const [id, session] of this.sessions) {
      if (now - session.lastActivity > this.ttlMs) {
        staleIds.push(id);
      }
    }
    for (const id of staleIds) {
      this.release(id);
    }
    return staleIds.length;
  }

  /**
   * Drains the pool by releasing all sessions and stopping the eviction timer.
   *
   * WARNING: This implementation immediately force-releases all sessions.
   * For true graceful drain, you must:
   * 1. Stop accepting new sessions (set a draining flag).
   * 2. Signal in-flight ToolExecution instances to cancel via their AbortControllers.
   * 3. Await a promise that resolves when all running executions reach a terminal state.
   * 4. Only then release sessions and clear the pool.
   * A configurable timeout before force-release is recommended for production use.
   */
  drain(): Promise<void> {
    this.isDraining = true;
    if (this.evictionTimer) {
      clearInterval(this.evictionTimer);
      this.evictionTimer = null;
    }
    const ids = Array.from(this.sessions.keys());
    for (const id of ids) {
      this.release(id);
    }
    return Promise.resolve();
  }

  get size(): number { return this.sessions.size; }
}

export { SessionPool, PooledSession };

Graceful Reconnect Protocol

Session token reuse is the mechanism that enables reconnection. When an agent reconnects using a valid token whose session has not been released or evicted, the acquire() method finds the existing session by token via the tokenIndex and returns it with an updated lastActivity timestamp, allowing the agent to reuse the same token and continue interacting with the server.

Important: Session data is not preserved across releases or TTL evictions. If a session has been released or evicted, acquire() with the same token creates a new, empty session. For true state resumption across disconnects, session data must be persisted to an external store (e.g., Redis) before release and restored on reacquire. The in-process pool shown here supports reconnection only while the session is still alive.

Three edge cases require explicit handling. If the token is valid but the TTL has expired and the session was evicted, the agent receives an error indicating session expiration and must start a new session. If the server is in the Draining state, reconnection attempts are rejected with a status indicating the server is shutting down. If the session was manually released (by an administrator or a cleanup policy), the token lookup returns nothing and the agent starts fresh.

Explicit Tool Lifecycle Management

Tool-Level State Machines for Long-Lived Operations

Some tools complete in milliseconds. Others, such as code execution in a sandboxed environment, multi-step document retrieval, or file generation pipelines, run for seconds or minutes. These long-lived operations need their own lifecycle independent of the server lifecycle, tracking five states: Pending, Running, Completed, Failed, and Cancelled.

The Cancelled state is particularly important. When a session disconnects while a tool is mid-execution, the tool must transition to Cancelled and release any resources it holds: file handles, database cursors, temporary buffers. Without this, a server processing long-running tools under intermittent connectivity will accumulate zombie operations that consume resources indefinitely.

Preventing Memory Leaks in Tool Cleanup

Each tool invocation receives its own AbortController. The controller's signal should be wired to two events: session disconnect and server drain. If either fires, the abort signal propagates into the tool's handler, which can check signal.aborted at each async boundary and bail out cleanly.

Integration note: Wiring the AbortController to session and drain events requires integration code not shown in the base classes here. In a complete implementation, SessionPool would hold a registry of ToolExecution instances per session and call execution.cancel() on release. Similarly, drain() would signal all registered executions before releasing sessions. See the companion repository for a complete example.

For large intermediate results (parsed documents, generated files, fetched datasets), consider wrapping them in WeakRef to allow garbage collection if only the tool execution holds a reference; this pattern is not implemented in the base class shown here and must be added per tool. Explicit nullification of large variables in finally blocks provides a second layer of defense.

enum ToolState {
  Pending = "Pending",
  Running = "Running",
  Completed = "Completed",
  Failed = "Failed",
  Cancelled = "Cancelled",
}

class ToolExecution {
  state: ToolState = ToolState.Pending;
  readonly abortController = new AbortController();
  private cleanupFns: Array<() => void> = [];
  private _error: Error | null = null;

  constructor(private sessionId: string) {}

  get error(): Error | null {
    return this._error;
  }

  onCleanup(fn: () => void): void {
    this.cleanupFns.push(fn);
  }

  start(): void {
    if (this.state !== ToolState.Pending) throw new Error("Cannot start");
    this.state = ToolState.Running;
  }

  complete(): void {
    if (this.state !== ToolState.Running) return;
    this.state = ToolState.Completed;
    this.cleanup();
  }

  fail(error: Error): void {
    if (this.state !== ToolState.Running) return;
    this._error = error;
    this.state = ToolState.Failed;
    this.cleanup();
  }

  cancel(): void {
    if (
      this.state === ToolState.Completed ||
      this.state === ToolState.Failed ||
      this.state === ToolState.Cancelled
    ) {
      return;
    }
    // Pending or Running — both are legal cancel targets
    this.state = ToolState.Cancelled;
    this.abortController.abort();
    this.cleanup();
  }

  private cleanup(): void {
    for (const fn of this.cleanupFns) fn();
    this.cleanupFns = [];
  }
}

export { ToolExecution, ToolState };

Entry Point and Transport Setup

The server.ts file initializes the MCP server, connects the transport, and wires up lifecycle management. Below is a minimal example using the StdioServerTransport:

import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import { ServerLifecycle, ServerState } from "./lifecycle.js";
import { SessionPool } from "./session-pool.js";

const lifecycle = new ServerLifecycle();
const pool = new SessionPool(100, 300_000);

const server = new McpServer({ name: "production-server", version: "1.0.0" });

// Register tools here using defineTool(server, ...)

lifecycle.transition({ state: ServerState.Ready });

let shutdownInProgress = false;

// Graceful shutdown
for (const signal of ["SIGTERM", "SIGINT"] as const) {
  process.on(signal, async () => {
    if (shutdownInProgress) return;
    shutdownInProgress = true;
    lifecycle.transition({ state: ServerState.Draining, remainingSessions: pool.size });
    await pool.drain();
    lifecycle.transition({ state: ServerState.Terminated, reason: signal });
    process.exit(0);
  });
}

try {
  const transport = new StdioServerTransport();
  await server.connect(transport);
  lifecycle.transition({ state: ServerState.Active, sessionCount: 1 });
} catch (err) {
  console.error({ event: "connect_failed", error: String(err) });
  lifecycle.transition({ state: ServerState.Draining, remainingSessions: pool.size });
  await pool.drain();
  lifecycle.transition({ state: ServerState.Terminated, reason: "connect_failed" });
  process.exit(1);
}

Note: For HTTP-based transports (SSE, Streamable HTTP), replace StdioServerTransport with the appropriate transport class from the SDK and add session-per-connection management using the SessionPool.

Testing Concurrent Agent Requests with Vitest

Unit Testing State Transitions

The ServerLifecycle class is deterministic and has no I/O side effects (it extends EventEmitter and emits events on transition), making it an ideal target for unit tests. Valid transitions should succeed and emit the correct events. Invalid transitions (such as Terminated to Active) should throw with a descriptive error message. Event ordering should be verified: a transition from Ready to Active should emit a single transition event with from: "Ready" and to containing the Active payload.

Mock Test Runner: Simulating 50 Concurrent Agents

The stress test creates 50 concurrent mock agent connections, each performing a randomized sequence of tool calls, abrupt disconnections, and reconnection attempts. After all agents have disconnected, assertions verify correctness: the session pool size is zero (no leaked sessions) and no errors were thrown during the run.

Note on the reconnect simulation: The test below calls pool.release() followed by pool.acquire(token), which creates a new empty session — not a true state-resuming reconnect. This validates that the pool correctly handles rapid disconnect/reconnect cycles without leaking sessions. True state resumption requires external session persistence, as discussed above.

Stress-Testing Schema Validation

Fuzz-style tests feed malformed inputs through defineTool registrations: missing required fields, wrong types, extra properties, empty strings where minimum-length constraints exist. Every case must produce a structured error response with isError: true and a human-readable message. The server must never throw an unhandled exception from malformed input.

import { describe, it, expect } from "vitest";
import { ServerLifecycle, ServerState } from "../src/lifecycle.js";
import { SessionPool } from "../src/session-pool.js";
import { ToolExecution, ToolState } from "../src/tool-execution.js";

// ── Unit: evictStale does not skip sessions due to mid-iteration mutation ──
describe("SessionPool.evictStale", () => {
  it("evicts all stale sessions without skipping any", () => {
    const pool = new SessionPool(200, 1000);
    const ids: string[] = [];
    for (let i = 0; i < 10; i++) {
      const s = pool.acquire(`token-${i}`);
      (s as { lastActivity: number }).lastActivity = Date.now() - 5000;
      ids.push(s.id);
    }
    const evicted = pool.evictStale();
    expect(evicted).toBe(10);
    expect(pool.size).toBe(0);
    pool.drain();
  });
});

// ── Unit: drain() blocks new acquire() calls ──
describe("SessionPool.drain", () => {
  it("rejects acquire after drain is called", async () => {
    const pool = new SessionPool(10, 5000);
    pool.acquire("tok-1");
    await pool.drain();
    expect(() => pool.acquire("tok-2")).toThrow("draining");
    expect(pool.size).toBe(0);
  });
});

// ── Unit: ToolExecution stores error on fail ──
describe("ToolExecution.fail", () => {
  it("preserves error for post-mortem diagnosis", () => {
    const ex = new ToolExecution("session-1");
    ex.start();
    const err = new Error("downstream timeout");
    ex.fail(err);
    expect(ex.state).toBe(ToolState.Failed);
    expect(ex.error).toBe(err);
  });
});

// ── Unit: invalid lifecycle transitions throw ──
describe("ServerLifecycle transitions", () => {
  it("throws on Initializing → Active (skipping Ready)", () => {
    const lc = new ServerLifecycle();
    expect(() =>
      lc.transition({ state: ServerState.Active, sessionCount: 0 })
    ).toThrow("Invalid transition");
  });

  it("throws on Terminated → Ready", () => {
    const lc = new ServerLifecycle();
    lc.transition({ state: ServerState.Ready });
    lc.transition({ state: ServerState.Draining, remainingSessions: 0 });
    lc.transition({ state: ServerState.Terminated, reason: "test" });
    expect(() =>
      lc.transition({ state: ServerState.Ready })
    ).toThrow("Invalid transition");
  });

  it("emits transition event with correct from/to", () => {
    const lc = new ServerLifecycle();
    const events: unknown[] = [];
    lc.on("transition", (e) => events.push(e));
    lc.transition({ state: ServerState.Ready });
    expect(events).toHaveLength(1);
    expect((events[0] as { from: string }).from).toBe(ServerState.Initializing);
  });
});

// ── Unit: token validation rejects oversized/malformed tokens ──
describe("SessionPool token validation", () => {
  it("rejects empty token", () => {
    const pool = new SessionPool(10, 5000);
    expect(() => pool.acquire("")).toThrow("Invalid session token");
    pool.drain();
  });

  it("rejects token exceeding max length", () => {
    const pool = new SessionPool(10, 5000);
    expect(() => pool.acquire("a".repeat(300))).toThrow("Invalid session token");
    pool.drain();
  });
});

// ── Integration: 50-agent concurrent stress with deterministic scenario ──
describe("Concurrent agent stress test (deterministic)", () => {
  it("handles 50 agents with no session leaks", async () => {
    const pool = new SessionPool(200, 30_000);
    const errors: Error[] = [];

    const agents = Array.from({ length: 50 }, (_, i) => async () => {
      const token = `agent-token-${i}`;
      try {
        const session = pool.acquire(token);
        await new Promise<void>((r) => setTimeout(r, 10));
        session.data.set("lastTool", `tool-${i}`);
        if (i % 5 === 0) {
          pool.release(session.id);
          const reconnected = pool.acquire(token);
          await new Promise<void>((r) => setTimeout(r, 5));
          pool.release(reconnected.id);
        } else {
          pool.release(session.id);
        }
      } catch (err) {
        errors.push(err as Error);
      }
    });

    await Promise.all(agents.map((fn) => fn()));

    expect(pool.size).toBe(0);
    expect(errors).toHaveLength(0);

    await pool.drain();
  });
});

describe("Concurrent agent stress test (legacy)", () => {
  it("handles 50 agents with no session leaks", async () => {
    const lifecycle = new ServerLifecycle();
    lifecycle.transition({ state: ServerState.Ready });

    const pool = new SessionPool(200, 30_000);
    const errors: Error[] = [];

    lifecycle.transition({ state: ServerState.Active, sessionCount: 0 });

    const agents = Array.from({ length: 50 }, (_, i) => async () => {
      const token = `agent-token-${i}`;
      try {
        const session = pool.acquire(token);

        // Simulate tool calls
        await new Promise((r) => setTimeout(r, Math.random() * 50));
        session.data.set("lastTool", `tool-call-${i}`);

        // Simulate abrupt disconnect for ~30% of agents
        if (Math.random() < 0.3) {
          pool.release(session.id);
          // Simulate new session after disconnect (not true state resumption)
          const reconnected = pool.acquire(token);
          await new Promise((r) => setTimeout(r, Math.random() * 20));
          pool.release(reconnected.id);
        } else {
          pool.release(session.id);
        }
      } catch (err) {
        errors.push(err as Error);
      }
    });

    await Promise.all(agents.map((fn) => fn()));

    expect(pool.size).toBe(0);
    expect(errors).toHaveLength(0);

    // Note: this heap delta assertion is a coarse smoke test, not a reliable
    // leak detector. GC timing makes it non-deterministic across runs and
    // environments. For definitive leak analysis, capture heap snapshots
    // using v8.writeHeapSnapshot() before and after the test and compare
    // retained object counts, or use clinic.js in CI.

    lifecycle.transition({ state: ServerState.Draining, remainingSessions: 0 });
    await pool.drain();
    lifecycle.transition({ state: ServerState.Terminated, reason: "test complete" });
  });

  it("rejects invalid state transitions", () => {
    const lifecycle = new ServerLifecycle();
    expect(() =>
      lifecycle.transition({ state: ServerState.Active, sessionCount: 0 })
    ).toThrowError("Invalid transition");
  });
});

Test runner setup: Vitest requires a vitest.config.ts to handle TypeScript source files. A minimal config: import { defineConfig } from 'vitest/config'; export default defineConfig({ test: { include: ['tests/**/*.test.ts'] } });

Deployment Considerations and Operational Checklist

Running as a Long-Lived Process vs. Serverless

Stateful MCP servers using in-process session storage are incompatible with serverless platforms like AWS Lambda or Google Cloud Functions. The session pool, lifecycle state machine, and TTL eviction timers all assume a persistent process. Cold starts would destroy session state, and the maximum execution duration limits on serverless platforms conflict with long-running tool executions. External session stores (e.g., Redis, DynamoDB) enable serverless deployment with additional complexity, but the architecture shown here targets long-lived processes.

Recommended deployment targets are container orchestrators (Kubernetes, ECS), process managers (PM2, systemd), or dedicated VMs. A health check endpoint that reports the current lifecycle state (Ready, Active, Draining) enables load balancers to route traffic away from servers that are shutting down.

Production Hardening Checklist

  • Implement a lifecycle state machine with guard clauses that reject invalid transitions
  • Co-locate Zod schemas with every tool registration so validation and types share a single source of truth
  • Configure the session pool with TTL eviction and a max session limit sized to your expected agent count
  • Wire an AbortController per tool invocation to session disconnect and server drain events (requires per-deployment integration; see the integration note above)
  • Handle SIGTERM/SIGINT by transitioning to Draining, waiting for in-flight tools to finish, then moving to Terminated
  • Emit structured logs with session ID and tool execution ID as correlation fields
  • Capture a heap snapshot baseline in CI using v8.writeHeapSnapshot() to detect memory regression across releases
  • Run the 50-agent concurrent stress test in CI on every merge. If it passes but production still leaks, your tool handlers are the next place to look.

From Prototype to Infrastructure

Your team gets explicit state machines for server and tool lifecycles, Zod-enforced schema contracts that double as validation and type inference, and session pooling with TTL eviction and reconnect support.

Your team gets explicit state machines for server and tool lifecycles, Zod-enforced schema contracts that double as validation and type inference, and session pooling with TTL eviction and reconnect support. Together, these transform an MCP server from a fragile demo into something you can deploy behind a load balancer. The code above, the Vitest stress tests, and the checklist give you a starting point. Pick the module closest to your current pain point and integrate it first.

SitePoint TeamSitePoint Team

Sharing our passion for building incredible internet things.

© 2000 – 2026 SitePoint Pty. Ltd.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.