Skip to content
English

Coordinated shutdown

system.terminate() stops the actor system, but a production app usually has work to do before that: drain in-flight HTTP requests, tell the cluster you’re leaving, flush a journal, close broker connections. And those steps have an order — leave the cluster before you stop the sharding region; stop the HTTP server before the actors that handle requests.

Coordinated shutdown is the DSL for that. You register tasks against named phases; the framework runs them in dependency order, one phase after the next, each with a timeout cap. Calling run() from any trigger (SIGTERM, K8s PreStop hook, an admin endpoint) executes the whole pipeline once.

import {
ActorSystem,
CoordinatedShutdownId,
Phases,
type Reason,
} from 'actor-ts';
const system = ActorSystem.create('my-app');
const coordinatedShutdown = system.extension(CoordinatedShutdownId);
coordinatedShutdown.addTask(Phases.ServiceUnbind, 'close-http', async (reason) => {
await httpServer.close();
});
coordinatedShutdown.addTask(Phases.ServiceRequestsDone, 'drain-in-flight', async () => {
await waitForInFlightRequests(/* up to 10s */);
});
await system.runUntilTerminated(); // SIGTERM/SIGINT → the pipeline → resolves when it is done

Three things happen when SIGTERM lands:

  1. The runtime calls coordinatedShutdown.run(new ProcessTerminateReason('SIGTERM')).
  2. The phases run in canonical order. Inside each phase, all registered tasks run in parallel; the phase waits for them all (or for their timeouts).
  3. The pipeline finishes with the built-in actor-system-terminate task, which calls system.terminate() for you.

Some of what happens in between is already wired for you; the rest is whatever you added.

Listed in execution order. Wired marks the phases the framework populates by itself — everything else is empty until you register something.

#Phase nameWiredTypical tasks
1before-service-unbind—Last-chance announcements before the server stops accepting connections.
2service-unbindsystem.http(...).bind(...), DevToolsStop the HTTP server / gRPC server / WebSocket listener from accepting new connections.
3service-requests-done—Wait for in-flight requests to finish; abort the rest.
4service-stopbroker actors (MQTT, Kafka, NATS, AMQP, …)Close client connections, release sockets.
5before-cluster-shutdown—Optional pre-cluster-leave hooks.
6cluster-sharding-shutdown-region—Tell the sharding region to hand off entities.
7cluster-leaveCluster.join(...)Issue a Cluster.leave() — gossip leaving status.
8cluster-exiting—Wait for the cluster to acknowledge the leave.
9cluster-exiting-done—Confirm cluster transition is complete.
10cluster-shutdown—Tear down cluster transports.
11before-actor-system-terminate—Last-chance app-level cleanup (flush journals, close brokers).
12actor-system-terminatethe built-in terminatorThe built-in system.terminate() task.

You don’t have to use every phase. Empty phases are no-ops; only phases with registered tasks do anything.

Set actor-ts.coordinated-shutdown.auto-register-tasks = false and the framework registers none of them. The phases stay, the built-in terminator stays, and every resource above is yours to register — for an embedder that owns the lifecycle of what it handed the system. It is one switch rather than one per subsystem, because the reason to reach for it is “I own the lifecycle”, never “unbind the HTTP server but leave the brokers to me”.

The Phases constant exports the canonical names — prefer it over string literals for autocomplete:

import { Phases } from 'actor-ts';
coordinatedShutdown.addTask(Phases.ServiceUnbind, ...); // ✓ typed
coordinatedShutdown.addTask('service-unbind', ...); // ✗ stringly-typed, no auto-complete

For app-specific work that doesn’t fit a canonical phase, declare your own:

coordinatedShutdown.addPhase({
name: 'flush-metrics',
timeoutMs: 3_000,
dependsOn: [Phases.BeforeActorSystemTerminate],
recover: true,
});
coordinatedShutdown.addTask('flush-metrics', 'push-prometheus', async () => {
await metricsRegistry.flush();
});

The dependsOn field is what makes the order DAG-shaped rather than linear — your phase runs after before-actor-system-terminate but before actor-system-terminate (because the latter has the former in its own implicit chain).

Because addPhase requires every dependsOn entry to name an already-registered phase, a cycle can’t be formed in the first place. The order itself is computed by a topological sort when run() executes — that sort would throw Error: cycle in phase dependencies if a cycle ever did arise.

Every task is a function from a Reason to void | Promise<void>:

type ShutdownTask = (reason: Reason) => Promise<void> | void;

The reason lets a task branch on why shutdown was triggered:

coordinatedShutdown.addTask(Phases.BeforeClusterShutdown, 'deregister', async (reason) => {
if (reason instanceof ClusterDowningReason) {
// We were downed — the registry has already written us off.
return;
}
await serviceRegistry.deregister(cluster.selfAddress.toString());
});

Built-in Reason classes:

ClassWhen
ProcessTerminateReason(signal)SIGTERM/SIGINT, from runUntilTerminated() or installProcessHooks().
ActorSystemTerminateReasonPass it yourself when the trigger is an ordinary shutdown.
ClusterLeavingReasonWhat Cluster.bootstrap’s shutdown() passes.
ClusterDowningReasonPass it when you run the pipeline because the cluster forced this node out.
UnknownReasonThe default when run() is called with no argument.

You can subclass Reason for app-specific triggers (AdminEndpointReason, HotReloadReason, etc.).

All tasks in a phase run concurrently — they’re started together and the phase waits for the last one (or its timeout). If you have ordering requirements within a phase (task B must wait for task A), put them in different phases with a dependsOn.

Each phase has a timeoutMs (default 5 s); each task is wrapped in a timeout race. A task that doesn’t finish in time is logged and either:

  • Recovered from (the phase continues, recover: true — the default). The next phase starts.
  • Halts the pipeline (recover: false). Subsequent phases are not run; shutdown stops mid-flight.

Override per phase:

coordinatedShutdown.setPhaseTimeout(Phases.ServiceRequestsDone, 30_000); // 30s drain budget

Or define your own with the wanted timeoutMs / recover:

coordinatedShutdown.addPhase({
name: 'aggressive-cleanup',
timeoutMs: 1_000, // strict cap
dependsOn: [Phases.BeforeActorSystemTerminate],
recover: false, // failure → halt
});

The drain budget sits inside the last phase

Section titled “The drain budget sits inside the last phase”

actor-system-terminate is a task that awaits system.terminate(), and terminate() begins by draining the actors under /user — see Terminating. So two budgets are nested, and the inner one has to be the smaller:

BudgetKeyDefault
The whole actor-system-terminate phaseactor-ts.coordinated-shutdown.default-phase-timeout5 s
The drain inside itactor-ts.system.shutdown-drain-timeout2 s

Raise the drain budget above the phase timeout and the phase is abandoned while the drain is still running — before a single actor has been told to stop. Raise both together, or raise the phase timeout first.

Almost every service wants the same thing — install the handlers, wait, shut down, exit — so that whole shape is one call:

await system.runUntilTerminated();
// Or, to listen for something else as well:
await system.runUntilTerminated(['SIGTERM', 'SIGINT', 'SIGUSR2']);

It installs the handlers, resolves once the system is down and the pipeline has finished, and detaches the handlers on the way out. That last step is not housekeeping: on Deno a signal listener holds the event loop open and has no unref, so a program that shuts down for any other reason — a terminate() from inside, an admin endpoint — would never exit.

A signal the platform cannot deliver is skipped rather than registered. Windows has no SIGTERM under any runtime, and Deno throws on a signal it cannot deliver instead of ignoring it, so asking for one is never a startup failure.

The process stays alive for as long as the promise is pending, and that is a guarantee the call makes rather than one it inherits from the handlers it installed. A signal handler is not by itself a reason for a runtime to keep running: Node unrefs its signal handles, so a service with nothing else on the event loop — no bound port, no open socket, no referenced timer — would otherwise drain its loop the moment it started waiting, and exit before the SIGTERM it had just armed itself for could arrive. So the call holds the loop open itself and lets go in the same step that detaches the handlers. A headless node needs no setInterval of your own to stay up.

The two halves are still available separately when you are embedding the system in a host that owns the process lifetime:

coordinatedShutdown.installProcessHooks(); // → run(new ProcessTerminateReason(signal))
coordinatedShutdown.removeProcessHooks(); // detaches exactly what it installed

Calling either twice is harmless. removeProcessHooks() never touches a listener it did not install, so a second ActorSystem in the same process — or the application’s own SIGTERM handling — is left alone.

In Kubernetes, the pod-shutdown sequence is:

1. K8s sends SIGTERM and ends the pod's grace period clock.
2. K8s also calls the PreStop hook (if configured), running concurrently.
3. After max(graceful-shutdown, grace-period), K8s sends SIGKILL.

The standard recipe:

// On SIGTERM, run coordinated shutdown:
await system.runUntilTerminated();
// PreStop hook script (in your container image):
// #!/bin/sh
// sleep 10 # give upstream LBs time to drain this pod
// exit 0

The sleep in PreStop gives the load balancer time to drop this pod from rotation before the actor system starts shutting down — so in-flight HTTP requests don’t see “I’m draining, go away.”

See Operations — Kubernetes for the full deployment manifest.

coordinatedShutdown.run() is idempotent — calling it multiple times returns the same in-flight promise. Three independent triggers (SIGTERM, an admin endpoint, and a cluster downing) all calling run doesn’t re-run the pipeline. The first call starts it; subsequent calls await the same completion.

This matters because in production you often have multiple shutdown paths:

// SIGTERM path:
await system.runUntilTerminated();
// Admin-endpoint path:
app.post('/shutdown', async (req, res) => {
await coordinatedShutdown.run(new AdminEndpointReason());
res.send('ok');
});
// A cluster downing is not one of them — see below.

Both end up running the same shutdown sequence once.

Being downed by the cluster does not start the pipeline: no reason class runs it for you, and ClusterDowningReason exists for a run() you make yourself. If you want a downed node to shut itself down, subscribe and say so:

cluster.subscribe((event) => {
if (event instanceof MemberDown && event.member.address.equals(cluster.selfAddress)) {
void coordinatedShutdown.run(ClusterDowningReason.instance);
}
});

What runs after coordinatedShutdown.run() completes

Section titled “What runs after coordinatedShutdown.run() completes”

By the time the promise resolves:

  • Every task in every phase has either succeeded or timed out.
  • The built-in actor-system-terminate task has called system.terminate(), which has stopped every actor and closed the dispatcher and scheduler.
  • The log sinks have been flushed and closed, bounded by actor-ts.logger.close-timeout (3 s by default) — so a batching sink’s last records are on disk or on the wire before the process goes. That flush lives inside terminate() rather than in a phase, which means it also covers a program that calls terminate() directly; see multi-sink logging.
  • The process is free to exit (process.exit(0)). Nothing left to do.

A whole production main:

async function main() {
const system = ActorSystem.create('my-app');
await system.http(8080).bind(routes); // registers its own unbind
// Register anything the framework does not own...
await system.runUntilTerminated();
}
main().catch((err) => {
console.error(err);
process.exit(1);
});

When SIGTERM arrives the pipeline runs, the system terminates, runUntilTerminated() resolves, its handlers come off, and the process exits because nothing is keeping the loop alive.

The CoordinatedShutdown API reference covers addTask, addPhase, run, and the full phase constant set.