Ir al contenido
Español

Backoff supervisor

Esta página aún no está disponible en tu idioma.

The framework’s default supervisor strategy restarts a child up to 10 times a minute. For transient failures, that can mean hammering a broken dependency — a broker that’s reconnecting, a DB that’s recovering — with restart-after-restart, each crashing identically.

BackoffSupervisor is the alternative. It wraps a single child actor and reschedules its restart with an exponential backoff (200 ms, 400, 800, …, clamped at a max), plus jitter so a herd of clients doesn’t synchronize.

import { ActorSystem, Actor, BackoffSupervisor } from 'actor-ts';
class Flaky extends Actor<{ kind: 'do-it' }> {
override preStart(): void {
if (Math.random() < 0.7) throw new Error('upstream not ready');
}
override onReceive(message: { kind: 'do-it' }): void {
this.log.info('ok');
}
}
const system = ActorSystem.create('demo');
const supervisor = system.spawn(
BackoffSupervisor.factory({
child: Flaky,
minBackoff: 200,
maxBackoff: 10_000,
randomFactor: 0.2,
}),
'flaky-supervisor',
);
// Send messages to the supervisor — they're forwarded to the
// current child, or stashed during a backoff window.
supervisor.tell({ kind: 'do-it' });

The supervisor:

  1. Spawns a Flaky child under stoppingStrategy (so a crash = a clean stop, not a default Restart).
  2. Death-watches the child.
  3. On Terminated, schedules a one-shot timer to spawn a fresh child after policy.delayFor(restartCount) ms.
  4. Buffers messages arriving during the backoff window.

When the child eventually starts successfully and processes messages, the buffered messages get flushed to it (with original sender refs preserved for ask-style replies).

Five steps, in execution order:

BackoffSupervisor

spawn child

child crashes → Terminated

schedule next spawn

after backoff.delayFor(n) ms

spawn child #2 → drain stash → ...

incoming messages buffered

(stash or drop)

The framework names successive children child-1, child-2, child-3, … so old terminations don’t collide with new spawns.

A respawn needs a death, not a message about one

Section titled “A respawn needs a death, not a message about one”

The supervisor accepts arbitrary user messages — that is the whole point, it forwards them to the child — so a Terminated naming the current child is something anything holding the supervisor’s ref could send it. Acting on that would retire a live child, bump the backoff counter, and spawn a replacement beside a child-1 nobody stopped and nobody watches any more.

Two things stop it. The runtime refuses a Terminated it did not emit before the supervisor ever sees it. And the supervisor asks the child’s own cell whether it has actually terminated before it retires it, so even a genuine notification about a running actor is declined with a warning rather than obeyed. If a predecessor is somehow still alive when its replacement is about to be spawned, it is stopped rather than left orphaned.

None of this is configurable, and none of it costs anything on the ordinary path: a child that really died reports terminated before its watchers hear about it, so the check is a boolean read.

The BackoffOptions<T> shape:

import { ActorOptions } from 'actor-ts';
type BackoffOptions<T> = {
child: ActorClassOrFactory<T>;
childOptions?: ActorOptions<T>;
childName?: string;
minBackoff: number;
maxBackoff: number;
randomFactor?: number; // default 0.2
policy?: BackoffPolicy;
resetCounter?: ResetCounter; // default 'after-min-stable'
forward?: ForwardStrategy; // default 'stash'
triggerOn?: TerminationTrigger; // default 'any'
maxStashSize?: number; // default 1000
drainGraceMs?: number; // default min(50, minBackoff)
forwardDuringGrace?: boolean; // default true
clock?: () => number;
};

The most interesting fields:

ValueWhen to respawn
'any' (default)Respawn on every termination — both crashes and clean stops.
'failure'Respawn only on crashes. A clean context.stopSelf() means “this child is done”; the supervisor stops itself afterwards.
'stop'Respawn only on clean stops (e.g. a transient connection actor that periodically tears itself down). Crashes propagate up.

'failure' is the right default if you’re modelling “restart on unexpected death” — a clean self-stop is a deliberate choice the supervisor should honor. 'any' restarts on every termination — the broadest possible policy, useful when the supervisor doesn’t care why the child stopped.

forward — what to do with messages while the child is dead

Section titled “forward — what to do with messages while the child is dead”
forward: 'stash', // buffer up to maxStashSize, drain after respawn
// or
forward: 'drop', // discard silently (debug-logged)

Stashing preserves sender refs so ask-replies continue to work after the respawn — a message asked while the child was down still gets its reply once the new child handles it.

Dropping is the right call for “transient pings that aren’t worth keeping” — telemetry, heartbeats, where stale messages are worse than lost ones.

When the stash fills. A backoff window longer than the traffic can wait through eventually pushes the stash past maxStashSize, and the oldest buffered message makes room for the newest. That message is not discarded silently: it becomes a dead letter, carrying its payload and the sender that would have received the reply, so an overflow is reconstructible afterwards instead of being a gap in the trace.

The accompanying warning is aggregated, not one line per message — it fires on the first eviction and again at each doubling, and carries the running total. A flood against a supervisor stuck in a long backoff would otherwise turn a message flood into a log flood at exactly the same rate.

resetCounter: 'after-min-stable', // reset when child alive >= minBackoff (default)
resetCounter: 'never', // never reset (counter grows monotonically)
resetCounter: { kind: 'after-time', ms: 60_000 }, // reset after 60s alive

Without resetting, a child that fails after a long-stable period gets the same long backoff as after a recent crash — which is usually wrong (the long-running success suggests the failure is fresh). 'after-min-stable' resets the count when the child has been alive for at least minBackoff, so a normal short backoff restarts after a long-running success.

After a respawn, the supervisor waits up to drainGraceMs (50 ms default) before draining the stash to the new child. This protects against children that crash in preStart:

  • If the child dies during the grace window, the stash is held back for the next incarnation — stashed messages aren’t lost to dead-letters when the child keeps crashing on startup.

forwardDuringGrace: true (default) sends new messages immediately during the grace; forwardDuringGrace: false stashes them until grace expires. The default trades a tiny risk of dead-lettering during a preStart-crash for lower latency on the happy path.

import { BackoffSupervisor, linearBackoff } from 'actor-ts';
BackoffSupervisor.factory({
child: ...,
minBackoff: 500,
maxBackoff: 10_000,
policy: linearBackoff({ minMs: 500, maxMs: 10_000, stepMs: 500 }),
});

Override the default exponential backoff with any BackoffPolicy — linear, fibonacci, custom. minBackoff / maxBackoff are still required (they’re advisory caps; the framework uses them for the resetCounter heuristic), but the policy controls the actual delay computation.

Three good fits:

  1. Broker connections (Kafka, NATS, AMQP) where a transient broker outage means the actor connect() fails for a few seconds before recovering. Default defaultStrategy would restart aggressively; backoff smooths it out.
  2. Database actors that hold a connection pool — when the DB hiccups, the actor crashes, and backoff buys time before re-establishing.
  3. Third-party API actors with rate-limit-aware retries — when a vendor returns 429, the actor crashes; backoff waits before re-trying.

OneForOneStrategy(decider, { maxRetries, withinTimeRangeMs }) caps restarts at N per window but doesn’t delay between them — the framework restarts immediately after each crash.

BackoffSupervisor adds the delay-between-restarts piece plus a message-buffering layer. The two are complementary:

  • For non-transient bugs, plain supervision with a low maxRetries is fine (give up after a few attempts and let the failure escalate).
  • For transient infrastructure issues, backoff supervision is worth the extra moving parts.

You can combine them — wrap a BackoffSupervisor’s own strategy with a OneForOneStrategy(..., { maxRetries: 10 }) to say “back off between restarts, but give up entirely after 10 attempts.”

  • Backoff policy — the exponentialBackoff / linearBackoff primitives that produce the policy value.
  • Supervision — the plain-supervision baseline this builds on.
  • Circuit breaker — for backing off before a call fails (not after).
  • Retry — per-call retry with similar backoff math, but outside the actor world.

The BackoffSupervisor API reference covers all options.