Ir al contenido
Español

Health checks

Esta página aún no está disponible en tu idioma.

The management routes expose two health endpoints:

  • GET /health — liveness. Returns 200 if the process is operational.
  • GET /ready — readiness. Returns 200 if the pod is ready to receive traffic.

Both aggregate one registry per ActorSystem, reached with healthChecksOf(system). Framework components register into it as they start; you register your own checks on the same object:

import { healthChecksOf, managementRoutes } from 'actor-ts/management';
const health = healthChecksOf(system);
health.addReadiness(async () => {
const ok = await database.ping();
return { name: 'database', status: ok, detail: ok ? undefined : 'database unreachable' };
});
health.addReadiness(async () => {
try {
await cache.ping();
return { name: 'cache', status: true };
} catch (e) {
return { name: 'cache', status: false, detail: (e as Error).message };
}
});
await system.http(8558).bind(managementRoutes(system, cluster));

When any check returns status: false, the corresponding endpoint returns 503 with a JSON body listing every check’s result.

The framework registers these itself — you get them without asking, and they are in the aggregate /ready reads. The gRPC grpc.health.v1.Health service reads the registry you hand it, so pass healthChecksOf(system) and the two can never disagree about what “ready” means; a fresh HealthCheckRegistry there forks them.

CheckKindRegistered byFails when
actor-systemlivenessthe registry itselfsystem.terminate() has completed
cluster-membershipreadinessCluster.jointhis node is not up in its own member view
cluster-transportreadinessCluster.jointhe node can reach none of the peers it still expects

cluster-transport is deliberately a total isolation test, not “every peer is reachable”. A partial partition leaves the node able to gossip, converge and route, so dropping it from the load balancer would remove capacity from a cluster that is coping. Being cut off from all of it is the case where continuing to serve is the split-brain hazard — and the case a node cannot see from its membership view alone, because no peer is allowed to downgrade this node’s own record. A single-node cluster expects nobody and always passes.

“Reachable” means an open connection to a peer the failure detector has not written off, and it needs both halves. An open socket on its own proves nothing: the canonical partition — iptables -j DROP, a black-holed route, a wedged peer — takes the traffic away without producing a FIN or an RST, so the sockets stay established for as long as the kernel keeps retrying while nothing at all is being exchanged. Requiring the peer to be reachable as well hands that judgement to the failure detector, which is the component that notices silence.

What the check does not catch, so you can plan around it:

  • the failure detector’s own latency — between the partition starting and the detector firing, the node still reports ready. That is the same window the rest of the cluster reacts on.
  • a one-way partition in which this node still receives. Its peers keep looking Up, so it keeps looking ready while nothing it sends arrives. Catching that needs an acknowledged round trip.

cluster.leave() leaves both checks registered and failing — /ready answers 503 for the rest of the process’s life, which is exactly what drains a pod before it stops. Nothing in the framework removes a readiness check on its way out of service: an empty aggregate reads as healthy (see below), so un-registering would make “everything passes” and “nothing is reporting any more” the same answer. A later Cluster.join on the same system retires the old pair and installs its own, so a process that leaves and re-joins recovers normally.

type HealthCheckFunction = () => Promise<HealthCheckResult> | HealthCheckResult;
type HealthCheckResult = {
name: string; // identifies the check in the response
status: boolean; // true = healthy
detail?: string; // human-readable note, usually on failure
};

A check may be sync or async. addLiveness / addReadiness each return an unsubscribe function — call it to remove the check again:

const remove = health.addReadiness(() => ({ name: 'warmup', status: warmedUp }));
// ... once warm-up is permanently done:
remove();

Long-running checks block the response, so keep them fast (sub-second, ideally < 100 ms).

ProbeWhat it answersWhat K8s does on failure
Liveness (/health)“Would restarting this process help?”Restart the pod.
Readiness (/ready)“Should a load balancer send this pod traffic?”Stop routing to this pod (keep it running).

A check is liveness or readiness depending on which method you call — addLiveness or addReadiness. They are separate lists; to run the same check for both, register it with both.

The difference is not a matter of taste — it decides what may go in each list:

  • Liveness may depend on nothing outside this process. A failing liveness check gets the pod killed, so a check that goes red when a shared database blinks turns one dependency’s outage into a fleet-wide restart storm, and the restarts cannot fix what broke. This is why the only framework liveness check is “the actor system has not shut down”.
  • Readiness is the right home for exactly what liveness must not touch — the dependencies a request needs. It takes the node out of rotation and leaves it running, so a database blip, a warming cache or a cluster rejoin belong here.

A readiness probe that answers 200 while the node cannot reach what it needs is worse than no probe at all: it keeps taking traffic it cannot serve.

/ready reports a clusterReady flag alongside the check list. It is the cluster-membership check’s own result, read back out of the aggregate rather than computed separately — so the flag and the check can never contradict each other. It is false until the local node reaches the Up state, and always true on a system that never joined a cluster:

GET /ready
→ 503
{
"status": "DOWN",
"clusterReady": false,
"checks": [
{ "name": "cluster-membership", "status": false, "detail": "this node is 'joining' in its own member view, not 'up'" },
{ "name": "cluster-transport", "status": true }
]
}

The in-process counterpart is cluster.awaitReady() / cluster.isReady() — the same up-only membership semantics, awaited from code instead of scraped over HTTP (see Cluster bootstrap → Waiting for readiness). The two are deliberately not coupled: awaitReady looks at membership alone, never at the readiness aggregate, because app-registered checks may only pass after initialisation that runs after bootstrap — waiting on them from inside Cluster.bootstrap could deadlock the start. /ready remains the load balancer’s view.

health.addReadiness(databaseCheck);
health.addReadiness(cacheCheck);
health.addReadiness(downstreamApiCheck);

All readiness checks run in parallel when /ready is hit. The response lists each check’s result:

{
"status": "DOWN",
"clusterReady": true,
"checks": [
{ "name": "cluster-membership", "status": true },
{ "name": "cluster-transport", "status": true },
{ "name": "database", "status": false, "detail": "connection refused" },
{ "name": "cache", "status": true },
{ "name": "downstream-api", "status": true }
]
}

The aggregate is UP iff every check’s status is true.

Registering no checks is a statement that nothing gates the probe, so a system with an empty readiness list answers 200 — otherwise every plain, cluster-free service behind managementRoutes would be 503 for its whole life. /ready, /health and the gRPC health service all apply that rule through the same exported function:

import { isHealthy } from 'actor-ts/management';
isHealthy([]); // true
isHealthy([{ name: 'database', status: false }]); // false

The rule is only safe because nothing empties the list. If a component could un-register on its way down, “healthy” and “no longer reporting” would be indistinguishable — a component going quiet reports status: false and stays registered instead. Use the undo returned by addReadiness for replacement, or when the thing the check observes is being torn down, never to signal trouble.

import { HealthCheckRegistry } from 'actor-ts/management';
it('readiness fails when the database is down', async () => {
const health = new HealthCheckRegistry();
health.addReadiness(async () => ({ name: 'database', status: false, detail: 'mock' }));
const results = await health.checkReadiness();
expect(results).toEqual([{ name: 'database', status: false, detail: 'mock' }]);
});

A bare new HealthCheckRegistry() is the right thing in a unit test — it holds only what the test puts in it, with none of the framework’s own checks. In production, always go through healthChecksOf(system): a second registry is a second notion of “ready”, and only one of them is the one the endpoints read.

checkLiveness() / checkReadiness() run the registered checks and return the HealthCheckResult[] the endpoints aggregate — handy for unit-testing checks in isolation. A check that throws is caught and reported as { name: 'unknown', status: false, detail }. The isHealthy(results) helper is the same all-pass predicate the endpoints use.

There is no built-in per-check timeout — a hung check blocks the whole /health (or /ready) response. For a check that can stall, race it against your own deadline:

health.addReadiness(async () => {
const status = await Promise.race([
slowProbe().then(() => true),
new Promise<boolean>((r) => setTimeout(() => r(false), 2_000)),
]);
return { name: 'downstream', status };
});

Without a guard, a stuck check eventually trips K8s’s own probe timeout (10 s default) and triggers a restart. Keep checks fast.