Skip to content
English

Downing strategies

When the cluster partitions, both halves stay running — and both think the other half has failed. Without intervention:

  • Two singletons would exist (one per side).
  • Sharded entities for the same key could spawn on both sides.
  • DistributedData replicas would diverge until reconciliation.

A downing strategy picks a winning side and forcibly downs the losing side. The losing nodes’ actors stop; the winning side continues as the cluster.

network split

During partition (KeepMajority decides)

n1 · n2 · n3

majority → continues

n4 · n5

minority → downs itself

Before partition

n1 · n2 · n3 · n4 · n5

Without a strategy: both halves keep running. When the partition heals, you have two separate clusters with the same name and no automatic merge. Always configure a strategy for production.

StrategyWins on…Trade-off
KeepMajorityThe side with > N/2 members.Simple; ties (exactly N/2-N/2 split) means both sides down.
KeepOldestThe side containing the lowest-addressed member (see the caveat below).Works for even-sized clusters where majority is undefined.
KeepRefereeThe side containing a designated referee node.Predictable but creates a single point of failure (the referee).
StaticQuorumThe side that meets a configured quorum size.Stricter than majority; the quorum-size config is the operator’s choice.
LeaseMajorityMajority + must hold a coordination lease.Paranoid safety; requires a Lease provider (K8s, etc.).

Pick by your cluster’s topology and operational constraints (see the picking section below).

import { Cluster, ClusterOptions, KeepMajority } from 'actor-ts/cluster';
const clusterOptions = ClusterOptions.create()
.withHost(host)
.withPort(port)
.withSeeds(seeds)
.withDowning(new KeepMajority());
const cluster = await Cluster.join(
system,
clusterOptions,
);

Pass the provider as downing. Every cluster event (member unreachable, member reachable, etc.) re-runs the provider with the current view; it returns the set of addresses to forcibly down. Returning an empty set means “no decision yet — wait.”

import { KeepMajority } from 'actor-ts/cluster';
new KeepMajority();

The classic strategy. Counts members; the side with more wins.

Picks:

  • Most general default.
  • No external dependencies.
  • Predictable when N is odd.

Doesn’t:

  • Handle exact N/2-N/2 splits — both sides down themselves.
  • Distinguish between “5 members healthy” and “5 members all on the same machine, the rack burns” — majority by count is blind to physical topology.

Right for clusters with an odd number of nodes and no special-purpose deployments.

import { KeepOldest, KeepOldestOptions } from 'actor-ts/cluster';
const keepOldestOptions = KeepOldestOptions.create().withDownIfAlone(true);
new KeepOldest(keepOldestOptions);

The side containing the “oldest” member wins.

Useful when:

  • You have an even number of nodes, where KeepMajority is ambiguous on 50/50 splits.
  • You can pin the address of the node that should survive — then the strategy is deterministic and needs no quorum arithmetic.

downIfAlone: true means “if I’m the oldest but everyone else is unreachable, I down myself” — prevents a one-node “winner” declaring itself the cluster after a total split.

import { KeepReferee, KeepRefereeOptions } from 'actor-ts/cluster';
const keepRefereeOptions = KeepRefereeOptions.create()
.withRefereeAddress('actor-ts://my-app@10.0.0.1:2552')
.withDownAllIfBelowQuorum(3);
new KeepReferee(
keepRefereeOptions,
);

A designated referee node is the deciding member. The side containing the referee wins; the other side downs itself.

Picks:

  • Most predictable strategy — there’s no ambiguity about which side wins.
  • Works for any cluster size, including even-sized.

Doesn’t:

  • Survive the referee itself going down. If the referee disappears, neither side has it, and the strategy can’t decide. Hence downAllIfBelowQuorum — “if the cluster is below this size, down everyone and let operators rebuild.”

Useful for two-DC clusters with a tie-breaker node in a neutral location (a third DC, a control-plane K8s namespace).

import { StaticQuorum, StaticQuorumOptions } from 'actor-ts/cluster';
const staticQuorumOptions = StaticQuorumOptions.create().withQuorumSize(3);
new StaticQuorum(staticQuorumOptions);

A side wins only if it has at least quorumSize reachable members. Below the quorum, the side downs itself.

Picks:

  • Stricter than majority — protects against minority-survivor scenarios where the minority would otherwise continue.
  • Configurable based on your operator’s confidence threshold.

Doesn’t:

  • Recover automatically. If multiple sub-quorum partitions form, none wins; the operator has to manually rebuild.

Right when you’d rather fail-stop than risk wrong-side survival.

import { LeaseMajority, LeaseMajorityOptions } from 'actor-ts/cluster';
const leaseMajorityOptions = LeaseMajorityOptions.create().withLease(someLeaseImplementation);
new LeaseMajority(
leaseMajorityOptions,
);

Majority plus a coordination lease. Wrap any other strategy: the winning side must additionally acquire a lease (e.g. a K8s Lease resource) before considering itself authoritative.

Picks:

  • Belt-and-braces safety. Two-fold check.
  • Useful when the network is unpredictable (e.g. cloud cross-zone).

Doesn’t:

  • Help if the lease provider itself is partitioned away.

Right for paranoid production scenarios where you’d rather double the cost of split-brain protection.

The interesting case is not “the lease said no” but “the lease said nothing”. Two things leave an acquire() unwanted while it is still on the wire: it has not resolved within acquireTimeoutMs (5 s by default), or the partition view moved on under it — the partition healed, or the unreachable set changed. The second is the likelier of the two, since a membership change inside a 5 s window needs no stall at all. Either way the strategy stops waiting for the attempt, and then has to assume it may still succeed on the wire, leaving the lease claimed by a node that no longer believes it won. The same three rules cover both:

  • The late result is dropped. Every attempt carries an epoch; a result arriving after its epoch was retired cannot write a decision. So a slow acquire() → true can never make this side claim survival after the fact.
  • A late win is released, and nothing new starts until it is. Ownership only begins when the acquire resolves, so the undo waits for it to report back — releasing earlier would be a no-op. While that is outstanding, decide() keeps returning an empty decision rather than starting a second attempt: two attempts overlapping is precisely how the cleanup could delete a lease this side is actively claiming.
  • If the release fails, the strategy stops deciding. The lease is then in an unknown state, so it enters fail-safe and returns an empty decision for this partition view until the partition heals. Deciding nothing is always safe; deciding wrongly is not.

The visible symptom of a lease provider having a bad day is therefore a cluster that waits, not one that splits.

Three questions in order:

  1. Odd or even cluster size?

    • Odd → KeepMajority.
    • Even → KeepOldest or KeepReferee (avoid ties).
  2. Do you have a stable tie-breaker node?

    • Yes → KeepReferee (most predictable).
    • No → KeepMajority or KeepOldest.
  3. Is fail-stop preferable to potential wrong-winner?

    • Yes → StaticQuorum (errs on the side of stopping).
    • No → one of the above.

For a typical 3-node K8s deployment: KeepMajority. For a 2-DC setup with a third-region tie-breaker: KeepReferee. For a 5-node cluster that should never go below 3: StaticQuorum(3).

The provider interface is tiny:

interface DowningProvider {
decide(view: ClusterPartitionView): DowningDecision;
}
type ClusterPartitionView = {
allMembers: ReadonlyArray<Member>;
unreachable: ReadonlySet<string>; // address strings
self: NodeAddress;
};

Implement decide(view) => Set<string> and return the address strings to down. Empty set means “no decision.”

Useful for app-specific rules — e.g. “always keep the node hosting role=primary,” or “if the partition includes the DB-master node, that side wins.”

The DowningProvider API reference covers the strategy interface.