Health checks
The management routes expose two health endpoints:
GET /health— liveness. Returns 200 if the process is operational.GET /ready— readiness. Returns 200 if the pod is ready to receive traffic.
Both aggregate one registry per ActorSystem, reached with
healthChecksOf(system). Framework components register into it as
they start; you register your own checks on the same object:
import { healthChecksOf, managementRoutes } from 'actor-ts/management';
const health = healthChecksOf(system);
health.addReadiness(async () => { const ok = await database.ping(); return { name: 'database', status: ok, detail: ok ? undefined : 'database unreachable' };});
health.addReadiness(async () => { try { await cache.ping(); return { name: 'cache', status: true }; } catch (e) { return { name: 'cache', status: false, detail: (e as Error).message }; }});
await system.http(8558).bind(managementRoutes(system, cluster));When any check returns status: false, the corresponding
endpoint returns 503 with a JSON body listing every check’s
result.
The built-in checks
Section titled “The built-in checks”The framework registers these itself — you get them without asking,
and they are in the aggregate /ready reads. The gRPC
grpc.health.v1.Health service reads the registry you
hand it, so pass healthChecksOf(system) and the two can never
disagree about what “ready” means; a fresh HealthCheckRegistry
there forks them.
| Check | Kind | Registered by | Fails when |
|---|---|---|---|
actor-system | liveness | the registry itself | system.terminate() has completed |
cluster-membership | readiness | Cluster.join | this node is not up in its own member view |
cluster-transport | readiness | Cluster.join | the node can reach none of the peers it still expects |
cluster-transport is deliberately a total isolation test, not
“every peer is reachable”. A partial partition leaves the node able
to gossip, converge and route, so dropping it from the load balancer
would remove capacity from a cluster that is coping. Being cut off
from all of it is the case where continuing to serve is the
split-brain hazard — and the case a node cannot see from its
membership view alone, because no peer is allowed to downgrade this
node’s own record. A single-node cluster expects nobody and always
passes.
“Reachable” means an open connection to a peer the failure detector
has not written off, and it needs both halves. An open socket on
its own proves nothing: the canonical partition — iptables -j DROP,
a black-holed route, a wedged peer — takes the traffic away without
producing a FIN or an RST, so the sockets stay established for as long
as the kernel keeps retrying while nothing at all is being exchanged.
Requiring the peer to be reachable as well hands that judgement to the
failure detector, which is the component that notices silence.
What the check does not catch, so you can plan around it:
- the failure detector’s own latency — between the partition starting and the detector firing, the node still reports ready. That is the same window the rest of the cluster reacts on.
- a one-way partition in which this node still receives. Its
peers keep looking
Up, so it keeps looking ready while nothing it sends arrives. Catching that needs an acknowledged round trip.
Leaving does not silence the checks
Section titled “Leaving does not silence the checks”cluster.leave() leaves both checks registered and failing —
/ready answers 503 for the rest of the process’s life, which is
exactly what drains a pod before it stops. Nothing in the framework
removes a readiness check on its way out of service: an empty
aggregate reads as healthy (see below), so un-registering would make
“everything passes” and “nothing is reporting any more” the same
answer. A later Cluster.join on the same system retires the old
pair and installs its own, so a process that leaves and re-joins
recovers normally.
The check signature
Section titled “The check signature”type HealthCheckFunction = () => Promise<HealthCheckResult> | HealthCheckResult;
type HealthCheckResult = { name: string; // identifies the check in the response status: boolean; // true = healthy detail?: string; // human-readable note, usually on failure};A check may be sync or async. addLiveness / addReadiness each
return an unsubscribe function — call it to remove the check
again:
const remove = health.addReadiness(() => ({ name: 'warmup', status: warmedUp }));// ... once warm-up is permanently done:remove();Long-running checks block the response, so keep them fast (sub-second, ideally < 100 ms).
Liveness vs readiness
Section titled “Liveness vs readiness”| Probe | What it answers | What K8s does on failure |
|---|---|---|
Liveness (/health) | “Would restarting this process help?” | Restart the pod. |
Readiness (/ready) | “Should a load balancer send this pod traffic?” | Stop routing to this pod (keep it running). |
A check is liveness or readiness depending on which method you
call — addLiveness or addReadiness. They are separate
lists; to run the same check for both, register it with both.
The difference is not a matter of taste — it decides what may go in each list:
- Liveness may depend on nothing outside this process. A failing liveness check gets the pod killed, so a check that goes red when a shared database blinks turns one dependency’s outage into a fleet-wide restart storm, and the restarts cannot fix what broke. This is why the only framework liveness check is “the actor system has not shut down”.
- Readiness is the right home for exactly what liveness must not touch — the dependencies a request needs. It takes the node out of rotation and leaves it running, so a database blip, a warming cache or a cluster rejoin belong here.
A readiness probe that answers 200 while the node cannot reach what it needs is worse than no probe at all: it keeps taking traffic it cannot serve.
Cluster readiness
Section titled “Cluster readiness”/ready reports a clusterReady flag alongside the check list. It
is the cluster-membership check’s own result, read back out of the
aggregate rather than computed separately — so the flag and the
check can never contradict each other. It is false until the local
node reaches the Up state, and always true on a system that never
joined a cluster:
GET /ready→ 503{ "status": "DOWN", "clusterReady": false, "checks": [ { "name": "cluster-membership", "status": false, "detail": "this node is 'joining' in its own member view, not 'up'" }, { "name": "cluster-transport", "status": true } ]}The in-process counterpart is cluster.awaitReady() /
cluster.isReady() — the same up-only membership semantics, awaited
from code instead of scraped over HTTP (see
Cluster bootstrap → Waiting for readiness).
The two are deliberately not coupled: awaitReady looks at membership
alone, never at the readiness aggregate, because app-registered checks
may only pass after initialisation that runs after bootstrap — waiting
on them from inside Cluster.bootstrap could deadlock the start.
/ready remains the load balancer’s view.
Multiple checks
Section titled “Multiple checks”health.addReadiness(databaseCheck);health.addReadiness(cacheCheck);health.addReadiness(downstreamApiCheck);All readiness checks run in parallel when /ready is hit.
The response lists each check’s result:
{ "status": "DOWN", "clusterReady": true, "checks": [ { "name": "cluster-membership", "status": true }, { "name": "cluster-transport", "status": true }, { "name": "database", "status": false, "detail": "connection refused" }, { "name": "cache", "status": true }, { "name": "downstream-api", "status": true } ]}The aggregate is UP iff every check’s status is true.
An empty check list is UP
Section titled “An empty check list is UP”Registering no checks is a statement that nothing gates the probe, so
a system with an empty readiness list answers 200 — otherwise every
plain, cluster-free service behind managementRoutes would be 503 for
its whole life. /ready, /health and the gRPC health service all
apply that rule through the same exported function:
import { isHealthy } from 'actor-ts/management';
isHealthy([]); // trueisHealthy([{ name: 'database', status: false }]); // falseThe rule is only safe because nothing empties the list. If a
component could un-register on its way down, “healthy” and “no longer
reporting” would be indistinguishable — a component going quiet
reports status: false and stays registered instead. Use the undo
returned by addReadiness for replacement, or when the thing the
check observes is being torn down, never to signal trouble.
Testing checks
Section titled “Testing checks”import { HealthCheckRegistry } from 'actor-ts/management';
it('readiness fails when the database is down', async () => { const health = new HealthCheckRegistry(); health.addReadiness(async () => ({ name: 'database', status: false, detail: 'mock' }));
const results = await health.checkReadiness(); expect(results).toEqual([{ name: 'database', status: false, detail: 'mock' }]);});A bare new HealthCheckRegistry() is the right thing in a unit test
— it holds only what the test puts in it, with none of the
framework’s own checks. In production, always go through
healthChecksOf(system): a second registry is a second notion of
“ready”, and only one of them is the one the endpoints read.
checkLiveness() / checkReadiness() run the registered checks
and return the HealthCheckResult[] the endpoints aggregate —
handy for unit-testing checks in isolation. A check that throws is
caught and reported as { name: 'unknown', status: false, detail }.
The isHealthy(results) helper is the same all-pass predicate the
endpoints use.
Timeouts
Section titled “Timeouts”There is no built-in per-check timeout — a hung check blocks
the whole /health (or /ready) response. For a check that can
stall, race it against your own deadline:
health.addReadiness(async () => { const status = await Promise.race([ slowProbe().then(() => true), new Promise<boolean>((r) => setTimeout(() => r(false), 2_000)), ]); return { name: 'downstream', status };});Without a guard, a stuck check eventually trips K8s’s own probe timeout (10 s default) and triggers a restart. Keep checks fast.
Where to next
Section titled “Where to next”- Management overview — the bigger picture.
- HTTP endpoints — the full endpoint reference.
- Kubernetes deployment — the probe configuration this pairs with.
