Upgrade strategies
Two kinds of production upgrade:
| Kind | Pattern |
|---|---|
| Code-only upgrade | New binary, same schemas. Rolling deployment — old + new versions coexist briefly, except across a release that changes the cluster wire (see below). |
| Schema-breaking upgrade | New shapes for events / state / messages. Migration first, then rolling deployment. |
Pick the kind, follow the pattern. Mixing them naively breaks production — old nodes can’t read new schemas or vice versa.
Cluster-wire compatibility per release
Section titled “Cluster-wire compatibility per release”A rolling deployment assumes the mixed-version window is survivable — for a few minutes, old and new nodes talk to each other. That assumption is a property of the release, not of your code, and two consecutive releases break it. Neither is a schema change in the sense this page uses the word — your events, state and config are untouched — so neither is caught by the additive patterns below.
| Upgrade | Mixed-version window | What it looks like when it fails |
|---|---|---|
| v0.15.x → v0.16.0 | Not safe (#112) | Every gossip frame carries a sequence, and GossipMessage gains a required field. An upgraded peer refuses an old node’s frames; an old node ignores an upgraded peer’s. Nothing errors — membership silently never converges while both versions run, so the cluster never agrees on who is in it. |
| v0.16.0 → v0.17.0 | Not safe (#450) | Frames are a tagged JSON tree instead of bare JSON.stringify output. Most legacy traffic decodes unchanged; a legacy body that already had a reserved tag’s shape does not. __map__, __set__, __regexp__, __bigint__, __url__, __number__ and __error__ throw at any depth, and a decoder throw costs the whole connection along with every frame batched into the same chunk. Two fail silently instead: __bytes__ decodes to a Uint8Array, __date__ to an Invalid Date. The other direction is plainly lossy — an older node reads the tag wrapper as ordinary data. |
Upgrade the cluster in one step. Stop every node, then start every node on the new version. No intermediate version makes either leg rollable, so a hop from v0.15.x straight to v0.17.0 crosses both breaks and needs the same treatment — once, rather than twice.
When this ends. Both breaks exist for the same reason: the cluster protocol has no version negotiation, so a node cannot ask a peer what it speaks and adapt. #823 adds that handshake. Until it lands, a mixed-version window is a hazard rather than a supported state, and this table is what to check before an upgrade.
Code-only upgrades
Section titled “Code-only upgrades”The common case. Bug fixes, refactors, behavior tweaks without changing persisted-data shapes.
1. Build the new binary (tag v1.2.3).2. Deploy via rolling update.3. K8s replaces pods one at a time.4. Each pod: SIGTERM → coordinated-shutdown → drain → new pod spawns → cluster-rejoin.5. Done.The cluster’s gossip + sharding rebalance + coordinated-shutdown handle the choreography. Total downtime: zero (if configured right; see Kubernetes deployment).
Requirements:
- Replicas ≥ 2. Single-replica clusters can’t drain cleanly.
- Coordinated shutdown configured with sane phase timeouts.
- Health checks correctly gate readiness.
Schema-breaking upgrades
Section titled “Schema-breaking upgrades”Whenever the upgrade changes:
- Event shapes in a journal.
- State shapes in a durable-state store.
- Message shapes that nodes might send each other during the rolling window.
- Configuration keys that move between major versions.
The pattern: make the change additive, then upgrade.
Pattern — additive event shapes
Section titled “Pattern — additive event shapes”Old code wrote:
type DepositedV1 = { kind: 'deposited'; amount: number };New code wants:
type DepositedV2 = { kind: 'deposited'; amount: number; currency: string };Step 1: deploy intermediate code that accepts both shapes.
class Account extends PersistentActor<...> { override eventAdapter() { return defaultsAdapter<DepositedV2>({ manifest: 'Deposited', currentVersion: 2, defaults: { 1: { currency: 'USD' } }, }); }}This step:
- Writes V2 events under the new shape.
- Reads V1 events with
currencydefaulted to USD. - Works in old + new clusters because old code reads its own shape and ignores envelope wrapping.
Roll this out via standard rolling deployment.
Step 2 (optional later) — drop the defaultsAdapter once
all old events are aged out or snapshotted. Usually keep it
indefinitely for safety.
See migration recipes for the per-pattern walkthrough.
Pattern — non-additive schema changes
Section titled “Pattern — non-additive schema changes”For renames, restructures, removed fields, the mechanics need more steps:
1. Deploy code that READS old + writes NEW. (`migratingAdapter`)2. Roll out fully. All new events are now in the new shape.3. Deploy code that READS NEW only (no longer supports old). Drops the migrating step.4. Optional: bulk-migration to rewrite still-extant old events into the new shape if you'd like to drop adapter complexity.See migratingAdapter for the implementation.
Inter-actor message changes
Section titled “Inter-actor message changes”// v1 message: { kind: 'request' }// v2 message: { kind: 'request', traceId: string }During a rolling deployment, old nodes might send v1 to new nodes (or vice versa). The new code must tolerate both versions of incoming messages.
Strategy:
- Add the new field as optional in the message type.
- New code can handle messages missing the field (default it).
- Deploy. Old → new sends without the field, works. New → old sends with the field, old ignores it.
Once everything’s on v2, the field can become required in a later deployment.
Configuration changes
Section titled “Configuration changes”# v1 → v2: renamed config keyactor-ts.cluster.gossip-interval = 1s # v1actor-ts.cluster.gossip-interval-ms = 1000 # v2 (renamed)The framework’s config system doesn’t auto-migrate renamed keys. Two strategies:
- Read both in the code that loads config; honor either name until you can require the new one.
- Run migration scripts that rewrite
application.confto the new key names.
Easier: avoid renaming config keys. When you must, deprecate the old name + warn at startup for one release before removing.
What if you can’t avoid downtime?
Section titled “What if you can’t avoid downtime?”Sometimes the schema-break is bad enough that an online migration is genuinely impossible — different storage backend, fundamental restructuring. Then plan downtime:
1. Announce maintenance window.2. Coordinated shutdown of the entire cluster.3. Run offline migration scripts (sometimes hours).4. Bring up the new version.5. Smoke-test before user traffic.Frequency: ideally never. But occasionally inevitable.
Rollback strategy
Section titled “Rollback strategy”Always have a rollback plan:
- Code-only: previous binary still works against same schemas. Roll back via K8s rolling-back.
- Schema-breaking: the previous code reads the new shapes (because step 1 was additive). Rollback is safe.
- Non-additive change: harder — the rollback step needs to also know the new shape. Avoid; if necessary, use feature flags to gate the new code path while keeping the old one reachable.
Where to next
Section titled “Where to next”- Operations overview — the production checklist.
- Rolling migration — the practical step-by-step recipe.
- Migration overview — schema evolution for events + state.
- Migration recipes — the cookbook.
- Coordinated shutdown — the graceful-stop machinery rolling deploys rely on.
