Troubleshooting
Esta página aún no está disponible en tu idioma.
When something’s wrong in production, this page is the starting point. Each section is a symptom, the likely cause(s), and what to check / log / measure to confirm.
For the diagnostic tools themselves:
- Logging — structured logs.
- Stock metrics — built-in actor / cluster metrics.
- Tracing — per-request flow.
Cluster
Section titled “Cluster”Symptom: cluster won’t form
Section titled “Symptom: cluster won’t form”Pods start, but never reach Up.
Causes to check (in order):
- Seeds unreachable — wrong addresses / DNS / RBAC.
- Cluster port firewalled — pods can’t talk on 2552.
- Different system names —
ActorSystem.create('app-a')on one node,'app-b'on another. - TLS misconfiguration — handshake fails silently.
Diagnostics:
cluster.subscribe((evt) => { if (evt instanceof MemberJoined) { system.log.info(`saw join: ${evt.member.address}`); } else if (evt instanceof SelfUp) { system.log.info('self up'); }});Check whether SelfUp fires. If not, the local node hasn’t
even bootstrapped. Look at the seed-provider output:
kubectl logs pod-1 | grep -i seed# → "discovered N seeds: ..." ← should not be emptyIf empty: seed provider’s selector / DNS / config is wrong.
Symptom: cluster flaps
Section titled “Symptom: cluster flaps”Members oscillate between reachable and unreachable.
Causes:
- Tight failure-detector thresholds for the network’s jitter. See failure-detector tuning.
- Network partitions (real ones — diagnose at the infrastructure level).
- GC pauses longer than the unreachable threshold.
Diagnostics:
rate(cluster_unreachable_duration_ms_count[5m]) > 0.1# Frequent transitionsIn logs:
[INFO ] cluster — node-X marked unreachable[INFO ] cluster — node-X marked reachable[INFO ] cluster — node-X marked unreachablePersistent flapping signals threshold tuning.
Sharding
Section titled “Sharding”Symptom: sharded entities don’t spawn
Section titled “Symptom: sharded entities don’t spawn”Messages to a sharding region don’t reach entities.
Causes:
- Coordinator not yet started — sharding is async; messages sent before the coordinator is ready get buffered.
- No nodes match the
role— empty up-member set, so no shards are allocated. - extractEntityId returns undefined — the message routing has nothing to hash.
Diagnostics:
sharding_shards_hosted{type="entity-type"}# Should equal numShards (total) across the clusterLogs:
kubectl logs pod-1 | grep -i shard# → "coordinator: allocating shard X to node-Y"If you see “no candidates,” your role filter rejects every
node. Check role tags in Cluster.join.
Symptom: rebalance storm
Section titled “Symptom: rebalance storm”Sharding metrics show continuous rebalances; entities flicker.
Causes:
- Aggressive allocation strategy —
LeastShardAllocationStrategywithrebalanceThreshold: 1and frequent membership changes. - Flapping cluster (see above) — every change triggers rebalance.
Fix:
new LeastShardAllocationStrategy(/* threshold */ 5, /* max */ 3);Higher threshold = less sensitivity to small imbalances.
Persistence
Section titled “Persistence”Symptom: actor takes 30 seconds to start
Section titled “Symptom: actor takes 30 seconds to start”preStart takes a long time; the actor’s first
onReceive is delayed.
Cause: deep journal without snapshots. Recovery reads every event ever.
Diagnostics:
histogram_quantile(0.99, persistence_recovery_duration_ms_bucket)If P99 recovery is > 1 second, set a snapshot policy:
override snapshotPolicy() { return everyNEvents(100); }See Snapshots.
Symptom: an entity’s state depends on which node answers
Section titled “Symptom: an entity’s state depends on which node answers”State jumps back after a rebalance or deploy; remembered entities vanish after a coordinator failover; a projection processes events twice.
Cause: each node reads its own database. Per-node storage (a SQLite file each, separate in-memory stores, two nodes each pointing at their own Postgres) means a moved entity replays whatever its new node holds. Nothing errors — the optimistic append check runs against the local database, so it structurally cannot fire across nodes.
Diagnostics — two log needles, both once per node:
node-local storage # the store can never be shared (#1356)storage identity differs # shared-capable, but two instances (#1358)The second one also catches misconfiguration you can’t see in the backend name: stale connection strings, a restored backup, two “identical” database containers.
Fix: point every node at the same database instance — see
storage locality &
identity — or
use replicated event
sourcing where
per-node journals are the intended design. There is no
automatic repair for histories that already diverged: pick the
surviving database per persistence id and re-seed the others from
it (migrateBetweenJournals), then fix the wiring before
restarting.
Symptom: events recovered but state is wrong
Section titled “Symptom: events recovered but state is wrong”On restart, the actor’s state doesn’t match what it should be from the journal.
Causes:
onEventhas a side effect — runs during replay and somehow alters the state path.onEventusesDate.now()or random — non-deterministic; each replay produces different state.- Schema change without an adapter — old events have a
different shape than what
onEventexpects.
Diagnostics: compare the actor’s state after recovery to what’s in the journal. Replay manually in a test to isolate.
Memory + performance
Section titled “Memory + performance”Symptom: memory grows unboundedly
Section titled “Symptom: memory grows unboundedly”Heap grows linearly with uptime; eventually OOM.
Causes (most common first):
- A mailbox with a slow consumer — mailboxes are unbounded by
default, so a producer that outruns its consumer grows the heap
until one of them stops.
actor_mailbox_sizeshows which actor, and the log carries a backlog warning from 10 000 messages up. - Subscriber set leaks — actors registering with the event stream or DistributedPubSub but never unsubscribing on stop.
- DistributedData keys accumulating —
LWWMapwith millions of keys. - Persistent buffers — stash buffers, ask reply-to refs.
Diagnostics:
actor_mailbox_size{class=...}# Find actors with persistently large queues
histogram_quantile(0.99, rate(actor_mailbox_wait_seconds_bucket[5m]))# …and how long messages are waiting before anyone gets to them.# A series only exists above 10 000 queued messages, so this is the# signal that moves first — a backlog shows up as wait long before# it shows up as size.
histogram_quantile(0.99, rate(actor_mailbox_depth_bucket[5m]))# How deep the queues actually got. Unlike the gauge this has no floor,# so a burst of a few thousand shows up here and nowhere else — and# unlike the gauge it is a distribution, so a spike between two 2-second# samples is still recorded.
histogram_quantile(0.99, rate(actor_dispatcher_queue_delay_seconds_bucket[5m]))# One level up: turns waiting for the dispatcher rather than messages# waiting for the actor. High here and low above means the actors are# fine and the scheduling is the queue.# Heap dump via runtime tools:node --inspect / Bun's profilerCheck mailbox sizing + the leak patterns in event stream.
Symptom: HTTP latency high under load
Section titled “Symptom: HTTP latency high under load”Actor work is fine; HTTP responses are slow.
Cause: an actor monopolizes the event loop, starving HTTP handlers.
Fix: per-actor ThroughputDispatcher on the heavy actor.
Tests + dev
Section titled “Tests + dev”Symptom: tests hang at the end
Section titled “Symptom: tests hang at the end”Bun / Vitest test process doesn’t exit.
Cause: await system.terminate() not called in a fixture
teardown. Leaked schedulers / actor cells keep the event
loop alive.
Fix:
afterEach(async () => { await tk.shutdown();});See TestKit.
Symptom: shutdown dropped my messages
Section titled “Symptom: shutdown dropped my messages”Work that was queued when terminate() was called never ran, and
turns up in the dead-letter stream instead.
terminate() drains the actors under /user first, so a plain
backlog is handled. Four things it will not wait for:
- The drain budget ran out. It is
actor-ts.system.shutdown-drain-timeout, 2 s by default. An actor that keeps producing new work while draining — a self-tell loop, a rally between two actors — never goes quiet, so only the budget ends it. Raise the budget, and raiseactor-ts.coordinated-shutdown.default-phase-timeoutwith it if you shut down through the pipeline. - The mailbox was parked.
context.throttle(...)and a supervisor-suspended mailbox both count as quiet, because neither drains at a rate a shutdown can wait for. Cancel the throttle before shutting down if the backlog matters. - The work was not in a mailbox yet. A
context.timerstick that had not fired, or atellfrom a promise a handler started withoutawaiting it, arrives after the tree looks quiet. Await it, or fire it as a message instead. - It was addressed to
/system. Framework actors are never quiet by design, so they are not part of the drain.
To confirm which, subscribe to DeadLetter on the event stream before
shutting down — the dead letter names the message and the recipient.
Subscribing has to happen before, because publishing is all that
happens by default: nothing keeps a dead letter, so there is nothing to
go back and read afterwards. Set
actor-ts.dead-letters.store = "persistent" if you would rather have
the record without knowing in advance that you will want it — see
Dead letters.
Symptom: flaky timing-sensitive tests
Section titled “Symptom: flaky timing-sensitive tests”Tests pass locally, fail in CI.
The usual cause is real-clock timing: a fixed sleep long enough on an idle machine and too short on a loaded CI runner. Use ManualScheduler to control time deterministically, or wait on the observable state rather than on a duration.
That is not the only cause, and guessing between them wastes an afternoon. Diagnosing test flakes has the repeat-run harness that separates a flake from a broken test, plus the catalogued causes — a timer quantum that fires early, a dispatcher hop no poll interval can win, and the three multi-node suites CI does not run at all.
Where to start when nothing’s obvious
Section titled “Where to start when nothing’s obvious”1. Check logs for ERROR-level entries. Filter by time window around the issue.2. Check stock metrics — what changed? (rate of restarts, mailbox depth, member count).3. Check cluster events on the event stream.4. If a request is failing, follow the trace ID through logs / trace backend.5. Reproduce in a multi-node-spec test if you can.Where to next
Section titled “Where to next”- Operations overview — the broader production checklist.
- FAQ — common questions and pitfalls.
- Stock metrics — what to read when symptoms emerge.
- Logging — how to make logs actually useful.
- Tracing — per-request flow when logs aren’t enough.
