콘텐츠로 이동
한국어

Troubleshooting

이 콘텐츠는 아직 번역되지 않았습니다.

When something’s wrong in production, this page is the starting point. Each section is a symptom, the likely cause(s), and what to check / log / measure to confirm.

For the diagnostic tools themselves:

Pods start, but never reach Up.

Causes to check (in order):

  1. Seeds unreachable — wrong addresses / DNS / RBAC.
  2. Cluster port firewalled — pods can’t talk on 2552.
  3. Different system names — ActorSystem.create('app-a') on one node, 'app-b' on another.
  4. TLS misconfiguration — handshake fails silently.

Diagnostics:

cluster.subscribe((evt) => {
if (evt instanceof MemberJoined) {
system.log.info(`saw join: ${evt.member.address}`);
} else if (evt instanceof SelfUp) {
system.log.info('self up');
}
});

Check whether SelfUp fires. If not, the local node hasn’t even bootstrapped. Look at the seed-provider output:

Terminal window
kubectl logs pod-1 | grep -i seed
# → "discovered N seeds: ..." ← should not be empty

If empty: seed provider’s selector / DNS / config is wrong.

Members oscillate between reachable and unreachable.

Causes:

  1. Tight failure-detector thresholds for the network’s jitter. See failure-detector tuning.
  2. Network partitions (real ones — diagnose at the infrastructure level).
  3. GC pauses longer than the unreachable threshold.

Diagnostics:

rate(cluster_unreachable_duration_ms_count[5m]) > 0.1
# Frequent transitions

In logs:

[INFO ] cluster — node-X marked unreachable
[INFO ] cluster — node-X marked reachable
[INFO ] cluster — node-X marked unreachable

Persistent flapping signals threshold tuning.

Messages to a sharding region don’t reach entities.

Causes:

  1. Coordinator not yet started — sharding is async; messages sent before the coordinator is ready get buffered.
  2. No nodes match the role — empty up-member set, so no shards are allocated.
  3. extractEntityId returns undefined — the message routing has nothing to hash.

Diagnostics:

sharding_shards_hosted{type="entity-type"}
# Should equal numShards (total) across the cluster

Logs:

Terminal window
kubectl logs pod-1 | grep -i shard
# → "coordinator: allocating shard X to node-Y"

If you see “no candidates,” your role filter rejects every node. Check role tags in Cluster.join.

Sharding metrics show continuous rebalances; entities flicker.

Causes:

  1. Aggressive allocation strategy — LeastShardAllocationStrategy with rebalanceThreshold: 1 and frequent membership changes.
  2. Flapping cluster (see above) — every change triggers rebalance.

Fix:

new LeastShardAllocationStrategy(/* threshold */ 5, /* max */ 3);

Higher threshold = less sensitivity to small imbalances.

preStart takes a long time; the actor’s first onReceive is delayed.

Cause: deep journal without snapshots. Recovery reads every event ever.

Diagnostics:

histogram_quantile(0.99, persistence_recovery_duration_ms_bucket)

If P99 recovery is > 1 second, set a snapshot policy:

override snapshotPolicy() { return everyNEvents(100); }

See Snapshots.

Symptom: an entity’s state depends on which node answers

Section titled “Symptom: an entity’s state depends on which node answers”

State jumps back after a rebalance or deploy; remembered entities vanish after a coordinator failover; a projection processes events twice.

Cause: each node reads its own database. Per-node storage (a SQLite file each, separate in-memory stores, two nodes each pointing at their own Postgres) means a moved entity replays whatever its new node holds. Nothing errors — the optimistic append check runs against the local database, so it structurally cannot fire across nodes.

Diagnostics — two log needles, both once per node:

node-local storage # the store can never be shared (#1356)
storage identity differs # shared-capable, but two instances (#1358)

The second one also catches misconfiguration you can’t see in the backend name: stale connection strings, a restored backup, two “identical” database containers.

Fix: point every node at the same database instance — see storage locality & identity — or use replicated event sourcing where per-node journals are the intended design. There is no automatic repair for histories that already diverged: pick the surviving database per persistence id and re-seed the others from it (migrateBetweenJournals), then fix the wiring before restarting.

Symptom: events recovered but state is wrong

Section titled “Symptom: events recovered but state is wrong”

On restart, the actor’s state doesn’t match what it should be from the journal.

Causes:

  1. onEvent has a side effect — runs during replay and somehow alters the state path.
  2. onEvent uses Date.now() or random — non-deterministic; each replay produces different state.
  3. Schema change without an adapter — old events have a different shape than what onEvent expects.

Diagnostics: compare the actor’s state after recovery to what’s in the journal. Replay manually in a test to isolate.

Heap grows linearly with uptime; eventually OOM.

Causes (most common first):

  1. A mailbox with a slow consumer — mailboxes are unbounded by default, so a producer that outruns its consumer grows the heap until one of them stops. actor_mailbox_size shows which actor, and the log carries a backlog warning from 10 000 messages up.
  2. Subscriber set leaks — actors registering with the event stream or DistributedPubSub but never unsubscribing on stop.
  3. DistributedData keys accumulating — LWWMap with millions of keys.
  4. Persistent buffers — stash buffers, ask reply-to refs.

Diagnostics:

actor_mailbox_size{class=...}
# Find actors with persistently large queues
histogram_quantile(0.99, rate(actor_mailbox_wait_seconds_bucket[5m]))
# …and how long messages are waiting before anyone gets to them.
# A series only exists above 10 000 queued messages, so this is the
# signal that moves first — a backlog shows up as wait long before
# it shows up as size.
histogram_quantile(0.99, rate(actor_mailbox_depth_bucket[5m]))
# How deep the queues actually got. Unlike the gauge this has no floor,
# so a burst of a few thousand shows up here and nowhere else — and
# unlike the gauge it is a distribution, so a spike between two 2-second
# samples is still recorded.
histogram_quantile(0.99,
rate(actor_dispatcher_queue_delay_seconds_bucket[5m]))
# One level up: turns waiting for the dispatcher rather than messages
# waiting for the actor. High here and low above means the actors are
# fine and the scheduling is the queue.
Terminal window
# Heap dump via runtime tools:
node --inspect / Bun's profiler

Check mailbox sizing + the leak patterns in event stream.

Actor work is fine; HTTP responses are slow.

Cause: an actor monopolizes the event loop, starving HTTP handlers.

Fix: per-actor ThroughputDispatcher on the heavy actor.

Bun / Vitest test process doesn’t exit.

Cause: await system.terminate() not called in a fixture teardown. Leaked schedulers / actor cells keep the event loop alive.

Fix:

afterEach(async () => {
await tk.shutdown();
});

See TestKit.

Work that was queued when terminate() was called never ran, and turns up in the dead-letter stream instead.

terminate() drains the actors under /user first, so a plain backlog is handled. Four things it will not wait for:

  • The drain budget ran out. It is actor-ts.system.shutdown-drain-timeout, 2 s by default. An actor that keeps producing new work while draining — a self-tell loop, a rally between two actors — never goes quiet, so only the budget ends it. Raise the budget, and raise actor-ts.coordinated-shutdown.default-phase-timeout with it if you shut down through the pipeline.
  • The mailbox was parked. context.throttle(...) and a supervisor-suspended mailbox both count as quiet, because neither drains at a rate a shutdown can wait for. Cancel the throttle before shutting down if the backlog matters.
  • The work was not in a mailbox yet. A context.timers tick that had not fired, or a tell from a promise a handler started without awaiting it, arrives after the tree looks quiet. Await it, or fire it as a message instead.
  • It was addressed to /system. Framework actors are never quiet by design, so they are not part of the drain.

To confirm which, subscribe to DeadLetter on the event stream before shutting down — the dead letter names the message and the recipient. Subscribing has to happen before, because publishing is all that happens by default: nothing keeps a dead letter, so there is nothing to go back and read afterwards. Set actor-ts.dead-letters.store = "persistent" if you would rather have the record without knowing in advance that you will want it — see Dead letters.

Tests pass locally, fail in CI.

The usual cause is real-clock timing: a fixed sleep long enough on an idle machine and too short on a loaded CI runner. Use ManualScheduler to control time deterministically, or wait on the observable state rather than on a duration.

That is not the only cause, and guessing between them wastes an afternoon. Diagnosing test flakes has the repeat-run harness that separates a flake from a broken test, plus the catalogued causes — a timer quantum that fires early, a dispatcher hop no poll interval can win, and the three multi-node suites CI does not run at all.

1. Check logs for ERROR-level entries. Filter by time window
around the issue.
2. Check stock metrics — what changed? (rate of restarts,
mailbox depth, member count).
3. Check cluster events on the event stream.
4. If a request is failing, follow the trace ID through logs /
trace backend.
5. Reproduce in a multi-node-spec test if you can.
  • Operations overview — the broader production checklist.
  • FAQ — common questions and pitfalls.
  • Stock metrics — what to read when symptoms emerge.
  • Logging — how to make logs actually useful.
  • Tracing — per-request flow when logs aren’t enough.