Cluster security
Esta página aún no está disponible en tu idioma.
The cluster transport defaults to plain TCP without authentication — fast, simple, fine for a private network. Not fine for any cluster that crosses an untrusted boundary (the public internet, multi-tenant kubernetes, cross-region links without VPN).
This page covers the production-security setup for the cluster transport.
Two concerns
Section titled “Two concerns”| Concern | How to address |
|---|---|
| Eavesdropping — peer-to-peer traffic readable by anyone on the wire. | TLS on the cluster transport. |
| Unauthorized joins — a malicious node connects + becomes a cluster member. | Mutual TLS (mTLS) — peer certificate verification. |
Enable TLS with mutual authentication (mTLS) for any external-facing cluster — it addresses both.
Enabling TLS
Section titled “Enabling TLS”import { TcpTransport, NodeAddress, Cluster, ClusterOptions } from 'actor-ts/cluster';import fs from 'node:fs';
const transport = new TcpTransport( NodeAddress.parse('my-app@10.0.0.5:2552'), system.log, { cert: fs.readFileSync('./tls/cluster.crt'), key: fs.readFileSync('./tls/cluster.key'), ca: fs.readFileSync('./tls/ca.crt'), rejectUnauthorized: true, // verify peer certs },);
const clusterOptions = ClusterOptions.create() .withHost('10.0.0.5') .withPort(2552) .withSeeds([...]) .withTransport(transport);await Cluster.join(system, clusterOptions);The TLS settings:
cert+key— this node’s certificate + private key. Both carry the material itself, not a path to it: nothing in the transport reads from disk, and none of the three runtimes accepts a filename in these fields either. Load it yourself, as the example above does. Both halves are required together on a listener — see the refusals below.ca— trusted CA bundle. Use to verify peers’ certificates.rejectUnauthorized: true— fail handshakes where the peer’s cert isn’t signed byca.requestClientCert— whether the listener demands a certificate from whoever connects to it. You will not normally set this: it defaults totruewhenevercais present, since a trust bundle on a cluster listener has no other purpose.
With a shared ca, the cluster is mutually authenticated —
every connection requires a peer cert signed by the trusted CA,
in both directions.
The certificate also decides which node
Section titled “The certificate also decides which node”mTLS answers “may this peer be in the cluster”. On its own it
never answered “is this peer the node it claims to be”: the
hello frame carries an address and no credential, so one
CA-signed node could announce itself under another member’s
address. That matters because the gossip-authority rules below
all key off the connection’s peer — they are exactly as strong
as the identity underneath them.
So when a peer presents a certificate, the address it claims in
hello must be one the certificate vouches for. A claim is
accepted when the CN or a SAN covers either:
- the address’s host — the ordinary case, where each node’s certificate carries its own hostname or IP, or
- the full
systemName@host— for deployments that mint a per-node identity and want the tighter binding.
Wildcards are honoured for the host, in the leftmost label only, exactly as TLS hostname verification does.
Nothing changes for a cluster with no peer certificate to read: plain TCP, one-way TLS, and a Deno listener (which cannot report one at all) behave exactly as before. The check strengthens mTLS deployments rather than adding a switch that can be left off.
Three configurations are refused at bind time rather than started in a weaker state than they read as:
- An incomplete server credential —
certwithoutkey,keywithoutcert, or atlsobject carrying neither (acaalone says which peers to trust when dialling; it gives the listener nothing to present). Empty counts as absent, which is what an unset environment variable or a mis-mounted secret looks like by the time it arrives. requestClientCert: truewith noca— there would be nothing to validate peer certificates against.- An mTLS listener on Deno —
Deno.listenTlstakes only a cert and a key, with no way to request or verify a client certificate, so the listener would authenticate nobody.
A Deno node can still join an mTLS cluster: it presents its
own cert / key when dialling, so the listener (on Node.js or
Bun) authenticates it like any other peer. Only hosting the
listener is unavailable — in a mixed deployment, keep the seed
nodes on Node.js or Bun.
One Deno difference worth knowing: rejectUnauthorized has no
equivalent there and is not mapped. Deno always validates the
chain, so reaching a self-signed peer means supplying its signing
CA in ca — which is the shape recommended here anyway.
Certificate management
Section titled “Certificate management”Three approaches:
Self-signed for development
Section titled “Self-signed for development”openssl req -x509 -newkey rsa:4096 -keyout cluster.key -out cluster.crt -days 365 -nodes -subj "/CN=actor-ts"Use the same cert + key on every node. Fine for dev / staging. Don’t use in production.
Private CA for production
Section titled “Private CA for production”# Create a CA once:openssl req -x509 -newkey rsa:4096 -keyout ca.key -out ca.crt -days 3650 -nodes -subj "/CN=actor-ts-ca"
# Per-node certs signed by the CA:openssl req -newkey rsa:4096 -keyout node-1.key -out node-1.csr -nodes -subj "/CN=node-1"openssl x509 -req -in node-1.csr -CA ca.crt -CAkey ca.key -CAcreateserial -out node-1.crt -days 365Each node gets its own cert; everyone trusts the CA. Rotating certs is per-node and doesn’t require touching the CA.
Cert-manager / vault for K8s
Section titled “Cert-manager / vault for K8s”For K8s deployments, use cert-manager with an internal CA or HashiCorp Vault. Certs are mounted as volume secrets; rotation handled by the cert manager.
# Example cert-manager Certificate spec:apiVersion: cert-manager.io/v1kind: Certificatemetadata: name: actor-ts-clusterspec: secretName: actor-ts-cluster-tls issuerRef: name: actor-ts-ca kind: ClusterIssuer commonName: actor-ts dnsNames: - actor-ts-cluster.svc duration: 8760h renewBefore: 720hThe pod mounts the secret as files; the actor reads them.
Firewall the cluster port
Section titled “Firewall the cluster port”Cluster port (2552) — internal-only:- pods can talk to pods on 2552- not exposed via Service / Ingress- LoadBalancer never sees itEven with TLS + auth, expose the cluster port narrowly. A NetworkPolicy in K8s:
apiVersion: networking.k8s.io/v1kind: NetworkPolicymetadata: name: actor-ts-cluster-internal-onlyspec: podSelector: matchLabels: app: actor-ts ingress: - from: - podSelector: matchLabels: app: actor-ts ports: - protocol: TCP port: 2552Only app=actor-ts pods can reach port 2552 on each other.
Threat model
Section titled “Threat model”| Threat | Mitigation |
|---|---|
| Network eavesdropping | TLS |
| Man-in-the-middle | TLS + cert verification |
| Unauthorized cluster join | mTLS (CA-signed peer certs) |
| Insider with stolen cert | Cert rotation + revocation |
| Malformed or hostile wire frame | Shape validation at the decode boundary |
| Frame split into thousands of chunks to multiply decode cost | Linear decode buffer — every byte copied once, on arrival |
| Socket that opens, sends half a frame and goes quiet | Stall deadline on an incomplete frame, plus a cap on inbound connections |
| Sockets opened in bulk that send nothing at all, to fill the inbound cap | Handshake deadline on every accepted socket, so no slot is held open-ended |
| Peer growing a node’s registries without bound | Per-registry caps on pub-sub and the receptionist |
| Address claimed before the node that owns it exists | Tight version-skew cap on every gossiped member version |
| Gossip frame captured off the wire and played back later | Per-sender frame sequence — a frame must out-number the last one accepted from that sender (partial: only once one has been) |
| Spoofed discovery answer steering the bootstrap | mTLS, plus pinnedAddresses on the seed provider |
| Compromised pod inside the cluster | Application-level auth (out of scope here) |
The cluster transport handles transport-level security. Application-level concerns (auth between specific actors, per-tenant isolation) remain your job — the cluster transport is trusted within itself.
Where the seeds come from
Section titled “Where the seeds come from”A seed provider asks an outside party — DNS, the K8s API — which addresses to talk to, and that answer arrives before any certificate has been checked. It cannot let an attacker join the cluster: the address a peer claims must be vouched for by its certificate, so a spoofed seed produces a connection that dies at the handshake.
That guarantee is configuration-dependent, which is the part worth stating plainly. A node only demands a peer certificate if its listener was configured to request one; where TLS is off, or where a half-configured listener never asks, there is nothing to check the claimed address against and a poisoned discovery answer is followed as given.
So pin what the resolver is allowed to say:
const dnsSeedProviderOptions = DnsSeedProviderOptions.create() .withHostname('actor-ts.example.com') .withSystemName('my-app') .withPort(2552) .withPinnedAddresses(['10.0.0.0/8']) .withLog((message) => logger.warn(message));Addresses outside the list never become seeds. Details, including the SRV-mode caveat (targets are hostnames, so pins are suffixes there and the eventual A lookup stays unpinned), are on the DNS seed provider and Kubernetes API seed provider pages.
This is defence in depth: it lowers what a DNS or Endpoints compromise buys, and it is the only discovery-layer control left standing when mTLS is not in place.
What the wire edge rejects
Section titled “What the wire edge rejects”Frames arrive as JSON and used to be cast to the protocol type rather than checked against it, so a peer could hand a node a value the code then read as if the type were true. Every frame is now validated before any handler sees it:
- The frame must be an object carrying a string
kind.null, a bare string and a number are refused — previouslynullalone was an eight-byte remote process kill, and it needed no completed handshake. - Node addresses must have a non-empty
systemNameandhostplus a positive integerport. A port that arrives as the string"2552"is the case worth knowing about: it renders identically in every log line and keys every map the same way, but never compares equal — a node that merged its own address in that shape stopped recognising itself. - A gossiped member’s
statusmust be one of the seven legal values. An unknown one used to reach an exhaustive match that threw after the member had been stored, so the node crashed and re-gossiped the poisoned entry to its peers. - A gossip batch is validated whole. One malformed member refuses the frame, so a bad entry cannot ride in behind a good one.
Two tiers, deliberately: a frame that fails validation is dropped and the connection stays up, because one bad frame should not cost a healthy peer its link. A handler that throws drops the connection — that is the case nobody understands, and it must not escape into the runtime’s socket callback.
Frame kinds registered by extensions (sharding, pub-sub, the receptionist, DistributedData, DevTools) pass this layer and validate their own payloads. So does replicated event sourcing, whose event envelopes ride on pub-sub rather than on a frame kind of their own — see what a replica trusts from its peers.
Validation is only half of it. A frame can be perfectly well-formed
and still speak for a node that did not send it, and that class of
defect is closed the same way everywhere in this codebase: the
identity is taken from the connection, not from the payload. A leave
may only retire the node that sent it; gossip replaces the sending
peer’s contribution and no one else’s; a sharding or singleton
directive is honoured only when it arrived wrapped by the per-path
envelope handler, which is the one place the authenticated address
exists; and a replicated event’s replica is held against the node
that published it. A frame that reaches an actor by generic path
resolution instead — which strips nothing but does carry no sender —
is treated as unauthenticated rather than as trusted.
A CRDT’s own state is the one place that rule does not reach.
DistributedData validates a decoded payload for shape and
plausibility: a counter slot must be a non-negative integer at
or under MAX_COUNTER_SLOT (≈ 2.2e12 — the largest per-slot
value for which a full counter’s value() is still an exact
integer), a register timestamp may be no more than five minutes
ahead of local time, a collection is capped at 4 096 entries, and
a __proto__ key is refused. It does not check that the
sending peer had any business writing the slot it wrote.
Grow-only counters merge by maximum, so any peer able to gossip
can raise any replica’s slot inside those bounds — the owning
node cannot lower it again, the value is written through to the
durable record, and delete(key) is best-effort because a peer
still holding the key re-gossips it. A DistributedData counter is
therefore trustworthy exactly as far as the set of peers you let
onto the wire, which makes the mTLS setup above a prerequisite
for a quota or billing counter rather than a hardening step.
Every rejection is logged at WARN naming the peer and the
offending field, so a dropped frame is diagnosable — a
version mismatch and a hostile peer look different in the log.
What decoding a frame costs
Section titled “What decoding a frame costs”Validation runs on a whole frame, and a whole frame has to be
reassembled from however many TCP chunks the sender chose to
split it into. The peer chooses that split, which used to
be the expensive half: the decoder rebuilt its accumulator on
every chunk, so assembling one frame cost work quadratic in
the chunk count. A frame just under the 16 MiB cap, delivered
in TCP-sized ~1400-byte writes, is about 12 000 chunks and
≈ 100 GB of memory copying — roughly 6000× amplification on
bytes the attacker never had to send, and reachable before the
hello gate, so no membership was needed. The decoder now
appends into a buffer it grows by doubling, so each byte is
copied once on arrival and the amplification is 1×.
Three bounds cover what an unauthenticated socket may still hold while it is connected:
- An incomplete frame may sit for 30 seconds without another byte arriving before the connection is closed. It is a stall bound, not a budget for the frame: the deadline is re-armed on every chunk, so a peer shipping a large frame over a congested link is never punished for being slow — only for going silent.
- The handshake itself has 5 seconds, from the moment the
socket is accepted, and an accepted socket that has not sent
its
helloby then is closed. This is the one an incomplete frame does not cover: a socket that sends no bytes at all is not stalled mid-frame, so nothing about it is being tracked. It is the same 5 seconds the dialling side already gives itself, and that clock starts earlier — before the TCP connect and the TLS handshake — so a peer that is still trying has always given up first. Once thehellolands the deadline is gone: an established peer between two gossip rounds is idle by design and is never dropped for it. - Inbound connections are capped at 1024. A fully-meshed cluster needs one per peer, so the cap is set far above any realistic topology; what it removes is the ability to multiply the per-connection cost by opening sockets in a loop. The newest connection is refused rather than an established one evicted — eviction would let an attacker push real peers off the node.
The second and third are a pair. A cap is only worth what its
slots’ turnover is worth, and without a handshake deadline the
cap becomes the attack: 1024 connections that send nothing
hold every slot for the life of the process, and every peer
and ClusterClient after them is refused. On Bun that is
reachable even under mTLS, because a socket’s open callback
fires before the TLS handshake completes — the slot would be
taken while there is still no certificate to check.
The first and third are what bound inbound decode memory at
all: the product of the two, with the per-connection term
already capped by actor-ts.remote.max-frame-bytes. Lowering
that key is also the lever for a deployment that never sends
large envelopes — the resident cost per connection follows it
directly.
What a well-formed frame still cannot grow
Section titled “What a well-formed frame still cannot grow”Shape validation says a frame is readable, not that acting on
it is free. Pub-sub gossip is the clean example: a peer naming
100 000 topics it claims subscribers for sends one perfectly
legal frame, and the receiver used to allocate 100 000 map
entries for it — no local Subscribe, no malformed field, no
log line. The receptionist had the same shape on its own gossip
path.
Membership itself had the same shape, and it is the one registry nobody has to opt into: every clustered node keeps a map of members, and gossip is what fills it. The rules below decide whether a claim is believable; none of them bounded how many believable claims one peer may make. A sender announcing its own address is waved through by design — refusing that would mean no node could ever join — so naming a fresh address per frame allocated an entry per name.
These registries are capped now, and the caps apply to the gossip path as much as to local calls:
| Registry | Capped by | Default |
|---|---|---|
| Live cluster members | cluster.max-members | 1000 |
removed tombstones | cluster.max-tombstones | 10000 |
| Pub-sub subscribers per topic | cluster.pub-sub.max-subscribers-per-topic | 10000 |
| Pub-sub topics | cluster.pub-sub.max-topics | 10000 |
| Pub-sub remote claimants per topic | cluster.pub-sub.max-remote-nodes-per-topic | 1000 |
| Receptionist subscribers per key | cluster.receptionist.max-subscribers-per-key | 1000 |
| Receptionist subscriptions total | cluster.receptionist.max-subscriptions-total | 10000 |
A claim over a cap is dropped and logged rather than refused on
the wire — the frame is well-formed, and dropping the connection
over it would let one noisy peer cost a healthy link. A local
Subscribe over a cap is answered with SubscribeRejected
instead, because the caller is in a position to do something
about it. Both pub-sub registries also watch their subscribers,
so the other half of the old growth — refs that stopped without
ever unsubscribing — is reclaimed rather than capped.
Lower the defaults for a deployment where the legitimate numbers are far below them: a cap is only a bound on the damage, and one set far above real usage bounds very little.
Why membership needs two caps, and why the second is the real one
Section titled “Why membership needs two caps, and why the second is the real one”Splitting members from tombstones is not tidiness. A phantom
member in up / joining / unreachable is a member the
failure detector is watching, so it is downed and dropped
failure-detector.down-after after the attacker stops feeding it
— five seconds, at the default. A record gossiped as removed
is watched by nothing: only cluster.tombstone.time-to-live
reclaims it, a day later. So the flood that persists is the
tombstone flood, and max-tombstones is the cap doing the work.
Refusing one costs nothing either, because a tombstone for an
address this node holds no record of suppresses nothing that
exists — while a refused live record costs a legitimate member
one gossip round, and comes with a WARN naming the cap.
Both caps are charged on the bucket a record moves into, not
on whether it creates an entry. A gossiped record that keeps an
entry where it already is — up → unreachable, or a newer
tombstone over an older one — is a free in-place update. One that
crosses between the buckets is charged to the bucket it enters:
a tombstone re-incarnated as up needs room among the live
members, a live member gossiped as removed needs room among the
tombstones. Charging only entry creation let the two caps trade
headroom with each other for free — a re-incarnation vacated the
tombstone bucket without giving up a map slot, so the next flood
of tombstones was admitted too, and alternating the two grew the
map without bound while respecting both caps at every individual
step. A refused conversion leaves the member live, where the
failure detector reclaims it the slower way.
Tombstones this node mints itself — a peer’s leave, a downing
decision, an operator down() — convert a record it already
holds and are never subject to the cap. Capping its own
bookkeeping would drop the suppression that stops stale gossip
resurrecting an evicted address, which is a liveness bug wearing
a security fix’s clothes.
The ceiling worth knowing is not the heap. Gossip carries the
whole member list, so at roughly 110 000 entries a node’s own
frame outgrows remote.max-frame-bytes and every peer terminates
the connection on the length prefix — the node evicts itself from
the cluster while still running, long before anything runs out of
memory.
Who may say what
Section titled “Who may say what”Passing the shape check does not make a frame believable.
Gossip merges used to be decided purely by version magnitude,
and versions are seeded from Date.now() — so an attacker
could always pick a winning number and rewrite any member’s
status, including the receiving node’s own. Two rules now
sit in front of the merge:
- Nobody downgrades us. A claim about this node’s own
address is refused. The one exception is promotion out of
joining/weakly-upintoup, which has to come from outside because it is the leader’s decision — and which is harmless to accept, since a joining node is already trying to becomeup. Everything else about our own record we decide ourselves. - Third-party claims need a sender with standing. Saying something about another node requires the connection’s peer to be a member this node already considers active. A sender may always announce its own record — that is how joining works.
Both rules key on the connection’s peer, not on the
payload’s from field, since that field is the one thing an
attacker fully controls. The same reasoning covers three
neighbours. A leave is only accepted from the node that is
leaving. A heartbeat refreshes the failure detector for the
peer that sent it rather than the address it names — that one
previously let a peer keep a dead node looking healthy, and
made the receiver dial an attacker-chosen host. And a
ClusterClient ask is answered down the connection the
envelope arrived on: routing on the payload let any connected
party choose the address this node replied to, and opened a
connection to.
The client envelope no longer carries a sender field at all —
after the rule above it had exactly one correct value, which
is the one the receiver already holds. An envelope that still
names an address other than its own connection is counted on
cluster_envelope_from_mismatch_total{frame} and otherwise
ignored; see ClusterClient.
Unreachability is deliberately not covered by these rules: “I cannot reach C” is inherently a third-party observation, and every node must converge on the same view before a downing provider decides. Refusing those claims would leave each node with only its own reachability picture.
A sharding directive needs the coordinator behind it
Section titled “A sharding directive needs the coordinator behind it”The same reasoning reaches one layer up, into the extensions
whose frames pass the wire edge and validate their own
payloads. A ShardRegion used to treat any message whose
kind began with sharding. as a framework directive —
including the five only the coordinator may issue. HandOff
stops every entity under a shard; ShardHome moves ownership;
RememberedEntities pre-creates entities; ShardMapUpdate
publishes an allocation map to every local subscriber, DevTools
panel and application listener included. One small frame from
anyone who had completed hello did any of that, repeatably.
The region could not have checked: sharding registered no
per-path envelope handler, so an inbound frame reached the
actor through generic path resolution, which delivers with no
sender at all. It now claims its own path on the envelope
router, which hands the handler the connection’s peer, and
wraps the frame in a class instance before putting it in its
own mailbox — a shape a JSON wire body cannot mint, so the
directive arms can tell a routed frame from a forged one.
A tagged { kind } object would not have been enough: it is
reproducible verbatim from a payload.
Both halves are load-bearing. Without the wrapper check, a frame addressed non-canonically — a trailing slash, a doubled separator — misses the exact-string handler lookup, still resolves to the same region through the actor tree, and arrives unwrapped. And without the origin check, an authenticated peer is any cluster member, so every member could issue directives to every region.
Refused frames are dropped and logged at WARN. The
coordinator’s own local leg builds the same wrapper, so a
single-node cluster rebalances exactly as before.
A region can only speak for its own node
Section titled “A region can only speak for its own node”The same gap ran in the other direction, and it was the more
expensive one. Every message a ShardCoordinator accepts is a
claim about the sender’s own node: which shards its region
hosts, that its region is gone, that a hand-off finished, where
to send a shard home or a statistics reply. The coordinator read
all of that out of the payload. Its only gate was isLeader()
plus, with a lease, holding the lease — which answers “am I the
authoritative coordinator?”, never “may this sender speak for
that region?”.
So one well-formed sharding.Register naming somebody else’s
address seized every shard of a type, and one
sharding.RegionTerminated evicted its region. The eviction is
the worse half: it sends no HandOff, so the victim keeps its
shard actors and its entities running. The result is two owners
for one live shard — the same entity id instantiated twice, and
for a persistent entity two writers on one persistenceId.
The region-side origin check above does not blunt this, because
the attacker never sends a ShardHome. It poisons the
coordinator’s allocation map with one forged frame and the
genuine coordinator then emits the redirect itself, from the
leader’s own node, inside a wrapper every region rightly
accepts. One forgery, every subsequent hop authentic.
The coordinator now claims its own well-known path on the envelope router and applies the same two conditions: the frame arrived inside the wrapper, and the address its payload names is the peer’s own. Both directions of the register loop are therefore attributed, which also fixed how a region addresses the coordinator — by the leader’s system name, not its own, since an actor path carries the system it belongs to.
hostedShards is bounded and capped on top of that. It was the
only caller-sized input the coordinator had — one Register
wrote an allocation entry per array element, with no range check
and no length cap, into state that is broadcast to every region
and persisted to the coordinator-state store, so the growth
survived restarts. Entries outside 0 .. numShards - 1 are now
dropped and the accepted set cannot exceed numShards.
A version cannot claim an address in advance
Section titled “A version cannot claim an address in advance”Version is a logical clock seeded from Date.now(), so
“highest version wins” also decides what happens the first
time an address is mentioned at all. That made an address
claimable before the node owning it exists. A stranger
announces itself under the address the next pod is about to
get — announcing your own record is the claim the rules above
never refuse — dates it close to the 24 h skew cap, and
attaches whatever roles it likes. The leader’s promotion loop
lifts the record into the active set, and the node that
really owns the address loses every merge afterwards, because
it seeds its version from its own clock and that is lower.
Roles are what routing, sharding placement, singleton hosting
and downing quorums are computed from, so the phantom is not a
cosmetic row in a member list.
A gossiped member version is therefore held to a tight clock-skew budget — 5 minutes by default:
const clusterOptions = ClusterOptions.create() .withHost('10.0.0.5') .withPort(2552) .withMaxVersionSkewMs(30 * 60 * 1000);The budget applies to every merge, not only to the record that introduces an address. It was originally the narrower rule — a tight cap on a first sighting, a generous 24 h one on every update — and that split could be stepped around by introducing the address first: two records for the same address in one frame, or a frame with no member records at all, which still makes the receiver file the sender’s own address. Any rule that lets a record earn the wider budget fails the same way one step later, because without peer certificates every step of the earning is something the attacker can produce.
Raise it for a deployment whose clocks are known to run loose;
24 * 60 * 60 * 1000 restores the old single-cap behaviour.
A refusal is not exclusion — a node announcing itself is still
recorded, at version 1 and without roles — but it is durable:
a node whose clock runs further ahead than the budget stays in
the member list without roles until its clock comes back. That
was always this cap’s verdict on such a node; what changed is
that the verdict now sticks instead of being reversed by the
node’s second gossip frame.
Refused records are reported once per frame rather than once per
record — a WARN naming the peer and the count, and the counter
cluster_gossip_records_refused_total{reason}, whose reason
label is closed at version-skew, map-cap, timestamp-skew,
replayed-frame and self-claim. Logging per record would hand a
peer log amplification in place of the growth it just lost.
self-claim is the rule above, seen from the counter: a peer
asserting a status for the receiving node that is neither the one
it already holds nor the leader promoting it to up. A peer
echoing the status the node does hold is the ordinary content of
every gossip round — its own record travels back to it in every
frame — so that one is refused without being counted or logged, and
the counter still reads zero on a healthy cluster.
A recorded frame cannot be played back to a receiver that has heard from its sender
Section titled “A recorded frame cannot be played back to a receiver that has heard from its sender”A gossip frame is a snapshot of the member map, and a member’s version only moves when its status does — so a frame captured off the wire stays valid indefinitely. Against a converged receiver that costs nothing: every record in it loses the “higher version wins” comparison to the record already on file. What makes it an exploit is an entry the receiver has deleted. The failure detector’s down path deletes outright rather than leaving a tombstone, so that a partition followed by a heal can re-discover the peer — and an expired tombstone is pruned for the same reason. Either way there is nothing left to compare against, and the branch that files a first sighting has no lower version bound at all.
So replaying a downed member’s own pre-down record brought it
back: same address, same version, up, carrying the roles it
had. Shard placement and singleton hosting follow the roles, so
the cluster starts routing work to a node it had evicted.
Every gossip frame therefore carries a sequence its author stamps — seeded from that node’s wall clock at startup and incremented once per frame — and a receiver remembers the highest one it has accepted from each connection peer. A frame that does not out-number that mark is dropped whole, before any record in it is looked at. The comparison itself needs no knob: it is between a peer and itself, so no clock-skew budget comes into it.
The mark is written in exactly one place — the gossip path, as it accepts a frame — so it exists for a peer this node has accepted a frame from, and for no other address. That is the bound, and the third residual below is what falls outside it.
A sequence must also be plausible — a finite number no
further ahead of the receiver’s clock than maxVersionSkewMs,
the same budget a gossiped version gets, because the number is
stamped from the author’s wall clock. A frame outside that budget
is refused, exactly like a repeat.
That started out as the weaker rule: the frame was merged, and
only adopting its number as the mark was refused, on the
argument that a frame numbered Number.MAX_SAFE_INTEGER cannot
be a recording of a real frame. Only the sequence is fabricated
in that attack. The members array is still the recording, and
nothing on this wire binds the two together — so a captured frame
with one field rewritten was merged, left the mark untouched, and
was therefore merged again on every delivery, without limit,
against a receiver that held a mark for a sender that was still
alive. Refusing it loses nothing: the mark stays where the last
plausible frame put it, so the real node’s next frame still
out-numbers it and still lands. Silencing the real node forever —
the exploit the version cap above exists to prevent — is closed
from both sides rather than one.
An address additionally carries an incarnation: which
process is answering at that system@host:port, minted once
per Cluster.join. It is deliberately carried and not yet acted
on. The field is optional on the wire, so that a node running
either version understands one running the other — and an
optional field is bypassed by stripping it, which means a
refusal keyed on a mismatch would be one an attacker opts out of
while a legitimate peer of the previous version walks into it.
Requiring the field breaks all eight address-bearing frame fields
at once and waits on the wire-protocol versioning work. It stays
out of toString, equals and compareTo, so nothing keyed on
a node’s identity changes.
The one comparison that needs no agreement is made: a record a peer sends about this node keeps this node’s own incarnation. The leader’s promotion is the single claim about itself a node accepts, and it is merged wholesale, address included — so without that substitution the local record’s identifier would be whatever the last peer to promote it happened to say.
Three things this does not close. First, a peer that has earned standing and composes a fresh frame naming a deleted address at its old version: that is a forgery rather than a replay, and refusing it needs the incarnation to be required rather than carried, which is the wire break above.
Second, a missing mark admits everything, and there are three ways
to be missing one — only one of which is the sender’s own
eviction. An evicted member’s mark is dropped with it, for the
reason above — and an ordinary partition takes out sender and
subject together, so that route arrives ready to use. A fresh or
restarted process starts with none at all. And a member learned
third-party never had one: gossip is epidemic, so a node files
C as up on B’s word and has still never seen a frame from C.
In that last case the sender is a full member the whole time and
nothing anywhere was evicted — which is why this guard cannot be
described as holding “while its sender is a member”. The earlier
wording of this section, and of the changelog and roadmap entries,
said exactly that, and it was wrong.
Refusing a frame from a peer with no mark is not the missing check: the first frame from every peer is one, so a receiver that refused them would never learn a cluster exists. What a recording gets from an empty mark is a two-frame bootstrap — the first frame buys standing and installs the mark, off its own recorded number, and the second out-numbers it and speaks for the subject. Against a member learned third-party it is shorter still, because the standing is already there and one frame does it.
Third, and this is why the two above stay open: nothing keyed on the sender’s own counter can separate a recording from a live frame, because both were stamped by the same counter. What separates them is which process emitted them, and the only receiver-checkable statement of that is the incarnation — which has to be required before a refusal can rest on it, and that waits on the wire-protocol versioning work.
Requiring it would close a recording of a previous incarnation
outright, which is the ordinary restart case and the bulk of the
exposure. Two edges would survive even then, and they are worth
stating so nobody expects more: a node downed while still running
answers under the same incarnation its recording carries; and a
freshly started receiver holds no earlier incarnation of the
subject to compare a first sighting against, so a record about a
member it has never heard of remains admissible. On a correctly
configured mTLS cluster neither edge is reachable by an outsider —
hello is bound to the peer certificate — so the exposure is
plaintext and server-only-TLS deployments, plus a compromised peer.
All three residuals are tracked separately.
Per-deployment recipe
Section titled “Per-deployment recipe”import { TcpTransport, Cluster, ClusterOptions } from 'actor-ts/cluster';
const tlsOptionsType = { cert: fs.readFileSync(process.env.TLS_CERT_PATH!), key: fs.readFileSync(process.env.TLS_KEY_PATH!), ca: fs.readFileSync(process.env.TLS_CA_PATH!), rejectUnauthorized: true,};
const transport = new TcpTransport(self, log, tlsOptionsType);
const clusterOptions = ClusterOptions.create() .withHost(host) .withPort(port) .withSeeds(seeds) .withTransport(transport);await Cluster.join(system, clusterOptions);Env vars carry paths; the cert-manager / vault mounts the files. Code stays generic across environments.
Where to next
Section titled “Where to next”- Operations overview — production-readiness checklist.
- TLS everywhere — TLS for HTTP + brokers + journals.
- Master key rotation — rotating data-at-rest encryption keys.
- Transports — the underlying transport interface.
- Configuration — the HOCON keys for TLS settings.
