Skip to content
English

Singleton overview

A cluster singleton is one actor that exists exactly once across the whole cluster. It runs on the leader node; if that node leaves, the next-elected leader re-spawns it. Callers on any node send messages via a proxy that routes to wherever the singleton currently lives.

cluster of 3 nodes

node-1

(leader)

node-2

node-3

Singleton (active)

manager-only

(standby)

manager-only

(standby)

Three actors per node make this work:

  • ClusterSingletonManager — on every node. Watches cluster events; spawns the singleton when this node becomes leader, stops it when it stops being leader.
  • ClusterSingletonProxy — on every node that talks to the singleton. This is what start and ref hand back: a forwarding ActorRef that always points at the current leader’s manager.
  • The singleton actor itself — the user’s actor, only ever instantiated on the leader.

A node running only a proxy can address the singleton but never host it — see Getting a ref without hosting.

The classic use cases:

  • A coordinator — a job scheduler, a saga orchestrator, a rate-limit budget tracker — that must produce one consistent view for the whole cluster.
  • An external-resource owner — the actor that holds a connection to a single external system (a license server, a legacy DB with single-connection licensing).
  • A leader-elected service — your own elected role for some cluster-wide responsibility.

If you’d write if (!alreadyExists) spawn(...), a singleton is probably what you want.

import { Actor, ActorSystem } from 'actor-ts';
import { Cluster, ClusterOptions, SingletonKey } from 'actor-ts/cluster';
class JobScheduler extends Actor<JobCommand> {
static readonly singleton = SingletonKey.of<JobCommand>('job-scheduler');
override onReceive(message: JobCommand): void { /* ... */ }
}
const system = ActorSystem.create('my-app');
const clusterOptions = ClusterOptions.create()
.withHost(host)
.withPort(port)
.withSeeds(seeds);
const cluster = await Cluster.join(system, clusterOptions);
// On every node: one call spawns this node's manager and returns the ref.
// Only the leader's manager actually constructs the JobScheduler.
const scheduler = cluster.singleton.start(JobScheduler);
// Anywhere in the app — same call on every node:
scheduler.tell({ kind: 'schedule', jobId: '42' });

start returns a plain ActorRef<JobCommand>, so a singleton is passed around and stored exactly like any other actor — nothing in a consumer’s signature has to know it is one. Behind the scenes it finds the current leader’s manager via the well-known path /system/cluster/singleton/manager-<typeName> and forwards messages there. When leadership changes, the ref’s target shifts automatically within a gossip round.

SingletonKey.of<Command>('type-name') ties a singleton’s name to its message type. Declaring it as a static readonly singleton on the actor means neither is repeated at the call site, and start / ref can infer the right ActorRef<Command> from the class alone.

An actor that needs constructor dependencies passes a factory as the second argument:

class UserRepository extends PersistentActor<UserRepositoryCommand, Event, State> {
static readonly singleton = SingletonKey.of<UserRepositoryCommand>('user-repository');
constructor(private readonly users: ActorRef<UserCommand>) { super(); }
}
const users = cluster.sharding.start(UserActor);
const repository = cluster.singleton.start(UserRepository, () => new UserRepository(users));

A role belongs on the key, as a second argument:

class Ingress extends Actor<IngressCommand> {
static readonly singleton = SingletonKey.of<IngressCommand>('ingress', 'edge');
}

It is not part of the identity — two keys are equal iff their names match — it rides along so that every node reads the same one. That matters for a node that only calls ref: it has no options object, so a role set only through withRole is invisible to it and its proxy would resolve a different host than the managers do.

For lease, or to override the role per deployment, pass options alongside — or use the full builder form when there is no class to hang a static on:

const singletonOptions = StartSingletonOptions.create<JobCommand>()
.withRole('control-plane') // wins over a role declared on the key
.withBufferSize(5_000); // messages held while no node hosts it
cluster.singleton.start(JobScheduler, singletonOptions);

bufferSize bounds what the proxy holds while the cluster has no host — normally a gossip round, but nothing bounds it in an outage. Past the cap (default 1000) messages go to dead letters with a warning rather than growing the buffer forever.

The buffer drains on the first cluster event that can have produced a host — a leader change, but equally a member coming up, going down, or leaving. That full set matters for a role-restricted singleton: the first member carrying the role can join without the leader changing at all, and before v0.17.0 that left the buffer undrained indefinitely while every later send routed normally.

A member going unreachable is deliberately not in that set, on either side — see Unreachability is deliberately not a trigger.

start puts the node into rotation: it can become the host. A node that only needs to talk to the singleton calls ref instead, which is the counterpart to ClusterSharding.startProxy:

// No manager on this node — messages route to whoever is hosting.
const scheduler = cluster.singleton.ref(JobScheduler);
scheduler.tell({ kind: 'schedule', jobId: '42' });

ref and start return the same memoised ref for a given key, and the local manager is resolved per delivery — so a node that calls ref first and start later keeps using the same ref, which simply begins delivering locally instead of over the wire.

Two more things the returned ref does not do:

  • ref.stop() is a warning no-op. Everywhere else ActorRef.stop() sends a PoisonPill to its target; here that would kill whatever the current leader is hosting. To take this node out of rotation, call cluster.singleton.stop(key).
  • Stopping is asynchronous. The manager releases its lease and its envelope path in postStop, and the actor name stays taken until termination settles, so starting the same singleton again in the same turn throws with an explanation rather than a duplicate-name error.

cluster.singleton also exposes isStarted(key) and managerFor(key) for diagnostics.

When the host node leaves the cluster:

  1. Detection: cluster gossip propagates MemberLeft / MemberRemoved for the leaving node.
  2. Election: the cluster elects a new leader (deterministic based on member sort order).
  3. Hand-over: the new leader’s manager asks every node eligible to host to stand down, and waits for each to confirm its instance has actually terminated.
  4. Spawn: only then does it spawn the singleton.
  5. Routing shift: proxies on every node see the leader change and update their forwarding target.

In flight messages during the transition land in dead letters unless you’ve configured durability — see “State across failover” below. Messages that reach the incoming host while it is waiting on step 3 are held and handed to the new instance.

The transition window is bounded by the failure detector’s timeout (typically a few seconds for unreachable detection). Singletons aren’t a low-latency-failover tool; they trade some unavailability during failover for uniqueness.

How strong that uniqueness is depends on what you configure, and it is worth knowing which one you have:

  • Default (no lease). The hand-over makes “at most one instance” hold whenever the outgoing host can be reached and answers. When it cannot — it is unreachable, or a partition split the cluster — the incoming host waits handOverTimeoutMs (10 s) and then hosts anyway, logging a warning: availability is chosen over an invariant it could not prove. Two instances are possible for the length of that condition.
  • With a lease. Arbitration moves to a third party both sides can reach, so exactly one of them holds it — and the lease is released only once the outgoing instance has genuinely terminated. This is the configuration to use when dual execution would do real damage. See Singleton with lease.

The new instance starts with a clean slate — same as a restarted actor on a single node. For state that should survive:

  • PersistentActor — the singleton persists events; the new instance replays them from the journal. Most production singletons use this. See PersistentActor.
  • DurableState — simpler: snapshot the current state; restore on restart. See DurableState.
  • DistributedData — for state that needs to be readable before the singleton restarts. Most singletons don’t need this; their state is private to the singleton.

Without one of these, every failover is a fresh start. For a short-lived coordinator that just routes incoming work, that’s often fine; for stateful workflows, persist.

There is one case none of the three answers well: state that is expensive to rebuild but cheap to copy — thousands of events to replay, a large read-through cache. Persistence makes the state survive; it does not make the successor fast, because the successor still has to replay.

Warm hand-over closes that gap on a planned move. Implement two methods and the outgoing instance’s state travels to the incoming one on the hand-over that was already happening:

import { Actor } from 'actor-ts';
import { SingletonKey, type WarmHandOverActor } from 'actor-ts/cluster';
class PriceCacheActor extends Actor<PriceCommand> implements WarmHandOverActor {
static readonly singleton = SingletonKey.of<PriceCommand>('price-cache');
private prices = new Map<string, number>();
override async preStart(): Promise<void> {
// Skipped entirely when a predecessor handed its cache over.
if (this.prices.size === 0) await this.loadEveryPrice();
}
serializeForHandOver(): Uint8Array {
return new TextEncoder().encode(JSON.stringify([...this.prices]));
}
restoreFromHandOver(state: Uint8Array): void {
this.prices = new Map(JSON.parse(new TextDecoder().decode(state)));
}
}

The conditional in preStart is the part that matters. restoreFromHandOver runs after the constructor and before preStart, which is the only position from which recovery can still be skipped — a preStart that recovers unconditionally pays the cost anyway and then overwrites what arrived.

Four things are worth knowing before relying on it:

  • It is opted into on the actor, not in the options. There is no flag to set. An actor that does not implement both methods behaves exactly as it did before.
  • It is allowed not to happen. No hooks, a snapshot over the cap, a serializer or a restore that throws, a host that was downed rather than asked, a peer that never answered — every one of those falls back to today’s cold start and logs at warn. Never write a singleton whose correctness depends on the state arriving; a singleton that cannot survive losing its host is a singleton that cannot do its job.
  • serializeForHandOver is called after postStop, so the snapshot is final: no further message will be folded in, and anything still in the mailbox has gone to dead letters rather than into it. An instance that died to a crash or an exhausted supervision budget is never asked — the state that survives such a death is the state that caused it.
  • The snapshot is capped, at 1 MiB by default (withMaxHandOverStateBytes). It travels while the singleton is running nowhere, so a very large one lengthens the outage instead of shortening it. A snapshot is base64 inside a JSON frame and so costs about a third more on the wire; one that would not fit the transport’s maxFrameBytes is refused regardless of the cap, because an oversized frame costs the whole inter-node connection rather than the message.

A PersistentActor singleton needs one extra thought: a warm-started instance holds state that never came from a replay, so its sequence number has to travel in the snapshot too. Otherwise its first persist writes at a position the journal has already used, which fencing correctly reads as a second writer.

If the cluster partitions, two halves might each elect their own leader — and both would spawn the singleton. That’s exactly the case singletons exist to prevent.

Three defenses, in order of complexity:

  1. A downing strategy that picks a winning side during partition (default option, no lease needed). The losing side downs itself; only one half remains active. See Downing strategies.
  2. A lease, passed with the start options:
    const singletonOptions = StartSingletonOptions.create<JobCommand>()
    .withLease(someLeaseImpl); // e.g. K8s lease, or in-memory for tests
    cluster.singleton.start(JobScheduler, singletonOptions);
    The manager must successfully acquire the lease before spawning the singleton. Only one side of a partition can hold the lease, so even with two leaders, only one singleton exists. See Singleton with lease.

The combination of “downing strategy + lease” is paranoid-safe; each alone is usually enough.

A singleton has overhead beyond a normal actor:

  • Every node runs a manager — they’re lightweight (a state machine watching cluster events), but they exist on every node.
  • Every node that calls the singleton runs a proxy — also lightweight, but adds a hop on every tell.
  • Leader-change is the cost of failover — a singleton is unavailable for the few seconds it takes the cluster to converge on a new leader.

If exactness isn’t required (you’d be happy with N replicas), use a cluster router or sharding instead — both scale horizontally with no leader-bottleneck.

The ClusterSingletonManager and ClusterSingletonProxy API references cover the full configuration surface.