Перейти к содержимому
Русский

Mailbox sizing

Это содержимое пока не доступно на вашем языке.

The default actor mailbox is unbounded. Nothing is discarded on the way in, and the queue grows until the actor drains it. Bounding one is a per-actor decision, because it is a decision to lose messages — you take it where you can say which messages are safe to lose.

This page is the decision guide for production mailbox sizing.

const ref = system.spawnAnonymous(Worker);
// ↑ unbounded FIFO mailbox — nothing is ever dropped on the way in

Unbounded is the honest default: an actor framework cannot know which of your messages is expendable, and between v0.10 and v0.15 this one guessed. It bounded every mailbox at 10 000 with drop-head, so an actor that fell behind silently lost its oldest queued message — which is right for a sensor reading and wrong for a Terminated signal, a delivery confirmation or a WebSocket close. All three went through the same queue, and each became a filed defect.

One of the three has since been taken off the table for good: a death-watch Terminated now arrives through a lane no overflow policy can shed, and no throttle can drop. So a bound you set costs the actor its backlog, never the deaths it is watching for. The other two are application messages, and there the sentence above still holds in full.

The ceiling that trade bought was not real either: only the user queue was bounded. System messages were never capped, so the process could exhaust its heap regardless.

What you get instead is growth you are told about:

  • An actor whose queue reaches 10 000 messages logs a warning, and again at each doubling — 20 000, 40 000, and so on. This is always on and needs no metrics stack.
  • With metrics enabled, actor_mailbox_size{class, path} reports the depth of any mailbox at or above that same mark.

So the failure mode an unbounded queue can still reach — heap exhaustion, long GC pauses, a producer that never learns there is a problem — announces itself well before it arrives. Watch for the warning; bound the actors that produce it.

Depth is not a throughput problem. Both queues inside a mailbox are ring buffers, so taking the next message advances an index rather than reindexing the backlog behind it: a queue of a million costs the same per message as a queue of ten. What a deep queue does cost is memory and latency — every message in it is retained, and the one at the back waits for everything ahead of it. Those are the reasons to bound, and neither is fixed by a faster queue.

Three patterns where an unbounded queue is the wrong answer. In each, the capacity and the policy are both deliberate choices:

1. Producer/consumer mismatch known in advance

Section titled “1. Producer/consumer mismatch known in advance”
import { ActorOptions } from 'actor-ts';
// Slow consumer: writes to disk at 10/sec; producer pushes 1000/sec
const writerOptions = ActorOptions.create()
.withMailbox(() => new BoundedMailbox({
capacity: 1_000,
overflow: 'reject',
}));
const slowWriter = system.spawnAnonymous(SlowWriter, writerOptions);

Bound at the worst-case-acceptable buffer. reject propagates backpressure to the sender — they see MailboxFullError and adapt (retry, drop, alert).

2. Telemetry-style actors (stale data is wrong)

Section titled “2. Telemetry-style actors (stale data is wrong)”
import { ActorOptions } from 'actor-ts';
const telemetryOptions = ActorOptions.create()
.withMailbox(() => new BoundedMailbox({
capacity: 5_000,
overflow: 'drop-head',
}));
const telemetry = system.spawnAnonymous(MetricsAggregator, telemetryOptions);

For metrics, sensor readings, status pings — fresher is better. drop-head discards the oldest pending message when new ones arrive, keeping the queue full of recent data.

import { ActorOptions } from 'actor-ts';
const authOptions = ActorOptions.create()
.withMailbox(() => new BoundedMailbox({
capacity: 10_000,
overflow: 'drop-new',
}));
const auth = system.spawnAnonymous(AuthActor, authOptions);

drop-new discards incoming messages when full — preserves already-queued work. Right when “the queue I have is the work I care about” — partial denial of service is preferable to processing nothing.

Three factors:

  1. Worst-case burst size — how many messages arrive in the worst-case window before the consumer can drain.
  2. Per-message memory — capacity × bytes_per_message bounds the memory cost.
  3. Latency budget — capacity / drain_rate bounds the worst-case latency a message waits before processing.

For a worker processing 100 msg/sec, expecting bursts up to 1000 msg arriving in 1 second:

capacity = 1000 # worst-case burst
worst-case latency = 1000 / 100 = 10s # if fully queued

If 10 seconds of queue is acceptable, capacity 1000 is fine. If not, reduce capacity or accept that producers will see MailboxFullError.

Stock metrics (Stock metrics) expose mailbox depth:

actor_mailbox_size{class="Worker", path="..."}
actor_mailbox_depth_bucket
actor_mailbox_wait_seconds_bucket
actor_mailbox_dropped_total{class="Worker", reason="drop-head"}

Watch:

  • actor_mailbox_size — the depth of any mailbox at or above 10 000 queued messages. A series only exists once an actor crosses that mark, so its presence is the signal; on a healthy system the metric is empty. A drained mailbox loses its series rather than standing at its last spike, so alert on the series existing, not on a transition to 0.
  • actor_mailbox_depth — the same quantity as a distribution, covering the range the gauge is silent on. This is the metric to size a capacity from: a capacity is a bet on the worst-case burst, and the histogram’s p99 is the measurement of that burst. Its last bucket boundary is the gauge’s 10 000 floor, so the two cover everything between them — the histogram tells you how deep, the gauge tells you who. It has no labels, so it costs the same whether you run ten actors or ten thousand sharded entities.
  • actor_mailbox_wait_seconds — how long messages actually waited before delivery. This is the earlier of the two backlog signals, and the one to alert on: depth has a 10 000-message floor before it reports anything, while wait starts climbing as soon as an actor stops keeping up. It has no labels, so it tells you that the system is falling behind, not which actor — pair it with depth, which has the path, once it fires.
  • actor_mailbox_dropped_total — non-zero with drop-head / drop-new is by design for the actors you bounded; on any other actor it should not appear at all. Aggregated per class, not per actor — the two metrics deliberately differ here, because a depth series only exists for an actor already in trouble while a drop series would exist for every actor doing its job. Use an onDrop of your own if you need per-instance drop counts; it receives the discarded envelope as well as the reason, so it can say what was lost and not only how much.
  • MailboxFullError rate at the sender — usually surfaces as supervisor restarts of the sending actor.

The same 10 000 threshold produces a log warning, repeated at each doubling, whether or not metrics are enabled. That is the line to alert on if you run no metrics stack.

A counter says how much was shed. It does not say what: a dropped telemetry sample and a dropped command with an unsettled callback are the same increment. deadLetterDrops closes that gap by routing each discarded envelope to dead letters, where it keeps its payload and its sender:

import { ActorOptions, BoundedMailbox, BoundedMailboxOptions } from 'actor-ts';
const commandMailbox = BoundedMailboxOptions.create()
.withCapacity(1_000)
.withOverflow('drop-new')
.withDeadLetterDrops(true);
const commandOptions = ActorOptions.create()
.withMailbox(() => new BoundedMailbox(commandMailbox));

PriorityMailboxOptions has the same switch, and there it is worth the cost more often: a message this mailbox sheds is one your priority function ranked last, and only the payload can confirm that was the right call.

Off by default, and the default is not timidity. A drop happens on the sender’s stack, and a dead letter is a durable capture followed by a synchronous publish to every event-stream subscriber — so an always-on version would turn load shedding into per-message work at precisely the moment the bound exists to make things cheaper. The rule of thumb: turn it on where you would have to explain the loss afterwards, leave it off for the firehose you bounded on purpose.

Two things it does not change. It reaches only mailboxes you construct yourself — withMailboxCapacity has no door onto it — and the letter carries the message, the sender and the recipient, the same three fields every other loss path in the framework records. The MDC context and the tracing span travelling on the envelope are not among them.

PolicyWhen
rejectBackpressure surfaces to sender. Sender must handle.
drop-headTelemetry / metrics — newest wins.
drop-newCritical work — preserve queued, drop incoming.

Pick by what the right answer is on overflow:

  • “Sender should retry / alert” → reject.
  • “Stale data is wrong” → drop-head.
  • “Queued work is precious” → drop-new.

There’s no “best” — context-dependent.

For actors with mixed urgency:

import { ActorOptions, PriorityMailbox } from 'actor-ts';
const workerOptions = ActorOptions.create<Message>()
.withMailbox(() => new PriorityMailbox<Message>({
priorityFor: (m) => m.kind === 'urgent' ? 0 : 5,
}));
const worker = system.spawnAnonymous(Worker, workerOptions);

Lower numbers = higher priority. System messages always trump.

Use for:

  • HTTP responses (urgent) vs batch jobs (deferrable).
  • Health pings vs bulk metrics.

Ordering and a ceiling are not an either/or. withMailbox and withMailboxCapacity cannot be combined, so the capacity goes on the mailbox’s own options:

import { ActorOptions, PriorityMailbox, PriorityMailboxOptions } from 'actor-ts';
const triageOptions = PriorityMailboxOptions.create<Message>()
.withPriorityFor((m) => m.kind === 'urgent' ? 0 : 5)
.withCapacity(10_000)
.withOverflow('drop-lowest-priority');
const workerOptions = ActorOptions.create<Message>()
.withMailbox(() => new PriorityMailbox<Message>(triageOptions));

The policy set is 'drop-lowest-priority' / 'drop-new' / 'reject', defaulting to reject. drop-head is absent on purpose: a priority queue’s head is the message you called most important. drop-lowest-priority sheds from the other end instead, which is the version of “shed load” that a priority mailbox exists to express — the urgent tier survives and the bulk tier pays.

Drops land in actor_mailbox_dropped_total like any other, so the same alerting applies.

One shape to watch when you size this: a message whose priorityFor cannot be evaluated is ranked last, which on a full mailbox makes it the one that gets shed. That is the right answer, but it means a broken priority callback shows up here as a rising drop-new rather than as an error — wire onPriorityError if you want to tell the two apart.

See Mailboxes for the full PriorityMailbox surface.

producer → reject backpressure → sender slows down
producer → drop-head → producer keeps going; reader sees latest
producer → drop-new → producer keeps going; reader processes earliest

Bounded mailboxes are one layer in a backpressure story. For end-to-end backpressure (the upstream system slowing down), you’d combine:

  • Bounded mailbox at the actor.
  • Sender retry logic.
  • Upstream rate-limiting (HTTP 429, broker push-back).

The mailbox enforces the local boundary; the rest is your protocol design.