HTTP endpoints
The management server exposes a small set of operational
endpoints. Most are read-only; the admin ones (/cluster/down,
/cluster/leave) are opt-in and off by default.
Security (security audit #8): the read-only endpoints still reveal
internal topology — member addresses (host:port), roles, the leader, and the
shard map. Treat the management server as sensitive: bind it to an internal
interface and gate it with the ipAllowlist (and/or auth) options of
managementRoutes(...) rather than exposing it publicly.
Health probes
Section titled “Health probes”GET /health
Section titled “GET /health”Liveness — is this process fundamentally healthy?
GET /health→ 200 OK{ "status": "UP", "checks": [ { "name": "actor-system", "status": true } ]}503 with { "status": "DOWN", ... } if any liveness check fails.
See Health checks.
GET /ready
Section titled “GET /ready”Readiness — should this pod receive traffic?
GET /ready→ 200 OK{ "status": "UP", "clusterReady": true, "checks": [ { "name": "cluster-membership", "status": true }, { "name": "cluster-transport", "status": true }, { "name": "database", "status": true } ]}The checks array holds every readiness check on the system’s
registry — the framework’s cluster-membership and
cluster-transport plus whatever you registered. clusterReady is
cluster-membership’s own result, read back out of that array
rather than computed separately, so it is false until the local
node reaches Up and true when there is no cluster at all. 503
if any check fails.
Cluster info
Section titled “Cluster info”Available when a cluster is passed to managementRoutes.
GET /cluster/members
Section titled “GET /cluster/members”GET /cluster/members→ 200 OK{ "members": [ { "address": "actor-ts://my-app@10.0.0.5:2552", "status": "up", "version": 3, "roles": ["compute"] }, ... ], "self": "actor-ts://my-app@10.0.0.5:2552"}Full membership snapshot. Useful for:
- Manual cluster-state checks during incidents.
- External dashboards that visualize the cluster.
- Tests verifying cluster joining.
GET /cluster/leader
Section titled “GET /cluster/leader”GET /cluster/leader→ 200 OK{ "leader": "actor-ts://my-app@10.0.0.5:2552", "isSelf": true }The leader’s address plus isSelf (is the local node the leader).
leader is null before a leader is elected. Useful for monitoring
leadership churn.
GET /cluster/shards?type=<typeName>
Section titled “GET /cluster/shards?type=<typeName>”GET /cluster/shards?type=cart→ 200 OK{ "typeName": "cart", "leader": "my-app@10.0.0.5:2552", "version": 7, "takenAt": 1710000000000, "regions": [ { "key": "my-app@10.0.0.5:2552|/system/cluster/sharding/region-cart", "address": "my-app@10.0.0.5:2552", "path": "/system/cluster/sharding/region-cart", "proxy": false, "shards": [0, 1, 2, ..., 33] }, ... ], "shardHome": [ { "shard": 0, "regionKey": "my-app@10.0.0.5:2552|/system/cluster/sharding/region-cart" }, ... ]}Shows the shard-to-region allocation for a sharded type. Use for:
- Verifying even distribution after a rebalance.
- Diagnosing hot regions (one region with too many shards).
- Manual rebalance triggers in development.
Where the answer comes from. The node answers from
ClusterSharding.shardMap(typeName) — the last map the coordinator
broadcast to this node’s region. Nothing needs configuring: no
DistributedData extension, no coordinatorStateStore, no round trip
to the leader. The one precondition is that the answering node takes
part in the type — it has called sharding.start(...) or
sharding.startProxy(...) for it — because the coordinator broadcasts
only to regions that registered. A node that never mentioned cart
has nothing to report and answers 404.
This endpoint used to read the coordinator’s DistributedData
snapshot, which made a 200 unreachable in a default configuration:
nothing in the framework starts that extension, and the snapshot is
only written when you opt into a coordinatorStateStore. The store
is still available and still optional — it shortens leader failover —
but this endpoint no longer depends on it.
Reading the body:
versioncounts coordinator broadcasts, not shard assignments — one bump per publish, and a publish coalesces a burst of moves. Compare it across two readouts to tell “the map moved” from “the map is the same”.takenAtandleaderare stamped by the answering node when the map arrived, so two nodes can report the sameversionwith slightly different timestamps.regions[].keyis<address>|<region path>, the same keyshardHome[].regionKeyuses. A region that hosts nothing is still listed, with an emptyshards— “registered and given no shards” is the interesting half of a rebalance question.- A 200 with an empty
shardHomeis normal, not an error:regionsfills as soon as regions register, while a shard only gets a home once an entity in it is addressed.
Metrics
Section titled “Metrics”GET /metrics
Section titled “GET /metrics”Opt-in via enableMetricsEndpoint: true.
GET /metrics→ 200 OKContent-Type: text/plain; version=0.0.4; charset=utf-8
# HELP actor_messages_processed_total ...# TYPE actor_messages_processed_total counteractor_messages_processed_total{class="Worker",path="..."} 12345...Prometheus text format. See Prometheus exporter.
Answers 503 instead when the installed MetricsRegistry cannot be
read back through collect() — the
prom-client bridge is
the one in the framework that cannot, because it forwards its writes
to prom-client and keeps no snapshot:
GET /metrics→ 503 Service Unavailable
metrics endpoint unavailable: the installed MetricsRegistry does notsupport collect() — ...A 200 with an empty body would be worse than the error: an empty
exposition is a valid scrape, so Prometheus records the target as
up and every framework series quietly stops existing. The check runs
per request, because the registry can be swapped at any point in the
system’s life.
Admin (opt-in)
Section titled “Admin (opt-in)”POST /cluster/leave
Section titled “POST /cluster/leave”Opt-in via enableLeaveEndpoint: true.
POST /cluster/leave→ 202 AcceptedleavingThe 202 body is the plain-text word leaving (not JSON) — the leave
is fire-and-forget.
Triggers graceful cluster-leave. The node transitions through
leaving → exiting → removed; shards rebalance away; the
actor system terminates (configurable).
Use for:
- Pod retirement before rolling-update.
- Manual node-out during incidents.
Gate behind authentication — anyone with port access can drain your node.
POST /cluster/down
Section titled “POST /cluster/down”Opt-in via enableDownEndpoint: true.
POST /cluster/down{ "address": "actor-ts://my-app@10.0.0.5:2552" }
→ 202 Accepted{ "downed": "actor-ts://my-app@10.0.0.5:2552" }Returns 202 with { downed } when the member was downed, or 404
if the address is unknown or already terminal.
Force-downs a remote member by address. Destructive — the target’s actors stop, its shards reallocate elsewhere.
Use for:
- Split-brain recovery when no downing strategy is configured.
- Removing stuck
unreachablemembers that refuse to recover.
High-risk endpoint — gate behind auth + audit logs.
Custom routes
Section titled “Custom routes”import { path, get, concat, completeJson } from 'actor-ts/http';import { managementRoutes } from 'actor-ts/management';
const baseRoutes = managementRoutes(system, cluster);
const customRoutes = concat( baseRoutes, path('admin', get(async () => completeJson(200, { appVersion: '1.2.3' })), ),);
await system.http(port, { host }).bind(customRoutes);managementRoutes(system, cluster, options) returns the base
routes; combine with your own using concat. Useful for
adding app-specific admin endpoints alongside the standard set.
Response format
Section titled “Response format”The read endpoints (/health, /ready, /cluster/members,
/cluster/leader, /cluster/shards) and /cluster/down return
JSON. /metrics returns Prometheus text/plain, and
/cluster/leave returns the plain-text body leaving.
Errors are plain-text bodies (the framework’s complete(status, message)), not a structured object — e.g. a missing type query
param yields a 400 whose body is missing query param \type“.
Status codes follow HTTP conventions: 200 for reads, 202 Accepted
for the admin actions (/cluster/leave, /cluster/down), 4xx for
client errors (404 when a /cluster/down address is unknown), and
503 when the cluster or a health check is unavailable.
Authentication & IP allowlisting (#312)
Section titled “Authentication & IP allowlisting (#312)”The privileged endpoints (/cluster/down, /cluster/leave,
/cluster/members, /cluster/shards, /metrics) accept anonymous
requests by default. That’s only safe when the management port
is on a network-isolated bind (a separate Service, a different
port, 127.0.0.1-only) or behind a reverse-proxy that does
authentication for you.
For production deployments that expose management endpoints to a
broader network, attach the built-in middlewares via
managementRoutes(system, cluster, { auth, ipAllowlist }):
import { BearerTokenAuth, IpAllowlist,} from 'actor-ts/http';import { managementRoutes,} from 'actor-ts/management';
const routes = managementRoutes(system, cluster, { enableLeaveEndpoint: true, enableDownEndpoint: true, enableMetricsEndpoint: true, // Shared-secret bearer token (rotation supported via multiple entries). auth: BearerTokenAuth({ tokens: [process.env.MGMT_TOKEN!], realm: 'my-app-mgmt', }), // Network-level fence — only requests from these CIDRs reach // the application at all. ipAllowlist: IpAllowlist({ allow: ['10.0.0.0/8', '127.0.0.1/32'], // Behind a reverse proxy, name the proxies: the client address // is then resolved from the forwarded chain's socket end // inwards. Without this the default reads the socket peer, // which behind a proxy is the proxy. // trustedProxies: ['10.9.9.0/24'], }),});Policy split:
authwraps the privileged subset (/cluster/*,/metrics)./healthand/readystay anonymous — Kubernetes liveness/readiness probes can’t easily attach an Authorization header. SetauthProtectHealth: truewhen your deployment guarantees the probes can present credentials.ipAllowlistwraps EVERY endpoint including/healthand/ready. Network-level isolation is independent of who’s authenticated.
Failure modes:
- Missing / wrong
Authorizationheader →401 UnauthorizedwithWWW-Authenticate: Bearer realm="...". - Client IP outside the allowlist (or unresolvable) →
403 Forbidden.
The middlewares are general-purpose — you can use them outside the
management subtree via the withMiddleware(mw, route) builder.
Where to next
Section titled “Where to next”- Management overview — the bigger picture.
- Health checks — custom check registration.
- Cluster overview — the
membership semantics behind the
/cluster/*endpoints. - Kubernetes deployment — the K8s recipe using these endpoints.
- Prometheus exporter — the metrics format details.
