跳转到内容
简体中文

Cluster bootstrap

此内容尚不支持你的语言。

A cold start asks every node the same question at the same moment — “is there a cluster yet?” — and discovery answers it differently on each of them. DNS has not fully propagated; a pod is Ready before its IP is in the headless service; the Kubernetes API returns a partial pod list. Act on the first answer and each node forms a cluster out of the subset it happened to see.

Gossip never repairs that. The results are separate clusters with the same name, each with its own leader, its own singletons and its own shard allocations.

Cluster bootstrap is the phase that runs before Cluster.join and makes sure at most one cluster comes out of a simultaneous start.

clusternodediscoveryclusternodediscoveryset changed? restart the marginloop[until unchanged for stableMargin]order by address —lowest is the initial seedinitial seed self-elects,everyone else stays joiningalt[a cluster already exists][nothing answers within the grace]lookup()contact pointsCluster.join(seeds = every other contact point)leader promotes this node to up

1 — Stable observation. Poll the seed provider on pollIntervalMs and require the returned set (plus this node) to be byte-identical for stableMarginMs before acting on it. Any change restarts the margin. A failed lookup is not an empty set — it does not count as an observation at all, so a DNS outage can never be mistaken for “I am alone”.

2 — Deferred self-election. Order the settled set by address. The lowest is the initial seed — but it does not form a cluster on the spot. Every node, winner included, joins with the other contact points as its seeds; only the winner also gets a deadline (selfElectionGraceMs) after which it will form a cluster if nobody has promoted it.

The one-line path — bootstrapCluster runs the phase for you:

import { bootstrapCluster, ClusterBootstrapOptions } from 'actor-ts/cluster';
const bootstrapOptions = ClusterBootstrapOptions.create('my-app')
.withHost(process.env.POD_IP!)
.withPort(2552)
.withDiscovery('kubernetes')
.withStableObservation(true);
const { cluster, shutdown } = await bootstrapCluster(bootstrapOptions);

Pass an object instead of true to override the timings:

const bootstrapOptions = ClusterBootstrapOptions.create('my-app')
.withHost(process.env.POD_IP!)
.withDiscovery('kubernetes')
.withStableObservation({ requiredContactPoints: 3, stableMarginMs: 8_000 });

If you call Cluster.join directly, run the observation and feed both of its outputs into the options — the seed list and the selfElection policy:

import {
Cluster,
ClusterOptions,
NodeAddress,
StableObservation,
StableObservationOptions,
} from 'actor-ts/cluster';
const selfAddress = new NodeAddress('my-app', process.env.POD_IP!, 2552);
const observationOptions = StableObservationOptions.create()
.withSeedProvider(seedProvider)
.withSelfAddress(selfAddress)
.withRequiredContactPoints(3);
const observation = new StableObservation(observationOptions);
const targets = await observation.resolveJoinTargets();
const clusterOptions = ClusterOptions.create()
.withHost(selfAddress.host)
.withPort(selfAddress.port)
.withSeeds(targets.seeds.map((address) => address.toString()))
.withSelfElection(targets.selfElection);
const cluster = await Cluster.join(system, clusterOptions);

resolveJoinTargets() returns the settled set, whether this node won (isInitialSeed), and the selfElection value to pass on. Deriving that value yourself is the one mistake that reintroduces the split brain, so the observation does it for you.

SettingDefaultWhat
stableMarginMs5000How long the contact-point set must stay unchanged.
pollIntervalMs1000How often the seed provider is polled.
maxWaitMs60000Total budget; exceeding it throws.
requiredContactPoints1Fewest contact points a settled observation may contain.
selfElectionGraceMs10000How long the elected node waits before forming a cluster.

Same keys under actor-ts.cluster.bootstrap.*, with the usual precedence — explicit options > HOCON > built-in defaults:

actor-ts.cluster.bootstrap {
stable-margin = 5s
poll-interval = 1s
max-wait = 60s
required-contact-points = 3
self-election-grace = 10s
}

requiredContactPoints is the one worth changing. The margin catches discovery that is slow; only a required count catches discovery that is stably wrong — a resolver that consistently returns two of your three pods will settle happily on the wrong set. The default of 1 exists so single-node development needs no configuration; in production set it to the replica count you expect.

A resolved bootstrapCluster means a ready cluster: this node is a full member (up) and at least minimum-members members are up. When that is not reached within the budget, the bootstrap runs the coordinated-shutdown pipeline and rejects with a ClusterReadyTimeoutError naming self’s status, the up count and the bar — it never resolves for a node that is still joining.

const bootstrapOptions = ClusterBootstrapOptions.create('my-app')
.withHost(process.env.POD_IP!)
.withDiscovery('kubernetes')
.withStableObservation(true)
.withAwaitReady({ minimumMembers: 3, timeoutMs: 30_000 });
const node = await bootstrapCluster(bootstrapOptions);
node.formedNewCluster; // false — joined an existing cluster rather than founding one

awaitReady takes true (the default — wait with the computed budget), false or 0 (skip the wait), a number (the budget in ms), or the options bag above. Unset fields fall through to HOCON and then to the computed budget:

FieldDefaultWhat
minimumMembers1Fewest up members — self included — before the cluster counts as ready. Size it like requiredContactPoints: to the replica count.
timeoutMsderivedawait-ready when set; else self-election-grace + 5 s behind stable observation — the budget every node’s readiness actually hangs on, winner or not — else a flat 5 s.
actor-ts.cluster.bootstrap {
minimum-members = 3
# await-ready = 30s # unset selects the grace-aware computed default
}

The same wait exists standalone — after a hand-wired Cluster.join, or later in the process before a rebalancing-sensitive step:

await cluster.awaitReady({ minimumMembers: 3, timeoutMs: 30_000 });
cluster.isReady(); // synchronous probe of the same predicate
cluster.selfMember(); // this node's own record, tombstone included
cluster.selfElected; // true — this node founded its cluster

Without timeoutMs the promise waits indefinitely, like system.whenTerminated() — pair it with a deadline wherever the cluster may legitimately never form; a node that has left or been removed never becomes ready. Readiness here deliberately means membership: app-registered readiness checks (the /ready endpoint’s aggregate) may depend on initialisation that runs after bootstrap, so waiting on them could deadlock the very call that comes first. /ready stays the load balancer’s view; awaitReady is the in-process one.

To restore the old fire-and-forget shape, pass awaitReady: false and watch membership yourself via cluster.awaitReady().catch(…).

SituationUse
Fixed, known addresses; one designated first nodePlain seed join.
Local development, single nodePlain seed join — the empty seed list already means “I am first”.
Nodes start simultaneously with dynamic addresses (K8s Deployment, autoscaling group)Bootstrap.
Every node is given the same seed listBootstrap — see below.

That last row is easy to miss. Cluster filters this node out of its own seed list, so if every node lists every node then no node is left with the empty list that ordinary self-election needs. Nobody becomes up, nobody becomes leader, and nobody is ever promoted — the cluster deadlocks in joining. The election is what breaks the symmetry.

  • Startup latency. At least stableMarginMs, plus selfElectionGraceMs on a genuine cold start (paid once, by one node). Joining an existing cluster is unaffected — the grace never fires.
  • Discovery load. One lookup() per pollIntervalMs per node until the set settles. The poll cadence is deliberately constant rather than backing off: a growing interval would sample the stable margin at drifting points, and the load it saves — at most a few dozen lookups per node — is not worth weakening the guarantee.