콘텐츠로 이동
한국어

HTTP endpoints

이 콘텐츠는 아직 번역되지 않았습니다.

The management server exposes a small set of operational endpoints. Most are read-only; the admin ones (/cluster/down, /cluster/leave) are opt-in and off by default.

Security (security audit #8): the read-only endpoints still reveal internal topology — member addresses (host:port), roles, the leader, and the shard map. Treat the management server as sensitive: bind it to an internal interface and gate it with the ipAllowlist (and/or auth) options of managementRoutes(...) rather than exposing it publicly.

Liveness — is this process fundamentally healthy?

GET /health
→ 200 OK
{
"status": "UP",
"checks": [
{ "name": "actor-system", "status": true }
]
}

503 with { "status": "DOWN", ... } if any liveness check fails. See Health checks.

Readiness — should this pod receive traffic?

GET /ready
→ 200 OK
{
"status": "UP",
"clusterReady": true,
"checks": [
{ "name": "cluster-membership", "status": true },
{ "name": "cluster-transport", "status": true },
{ "name": "database", "status": true }
]
}

The checks array holds every readiness check on the system’s registry — the framework’s cluster-membership and cluster-transport plus whatever you registered. clusterReady is cluster-membership’s own result, read back out of that array rather than computed separately, so it is false until the local node reaches Up and true when there is no cluster at all. 503 if any check fails.

Available when a cluster is passed to managementRoutes.

GET /cluster/members
→ 200 OK
{
"members": [
{
"address": "actor-ts://my-app@10.0.0.5:2552",
"status": "up",
"version": 3,
"roles": ["compute"]
},
...
],
"self": "actor-ts://my-app@10.0.0.5:2552"
}

Full membership snapshot. Useful for:

  • Manual cluster-state checks during incidents.
  • External dashboards that visualize the cluster.
  • Tests verifying cluster joining.
GET /cluster/leader
→ 200 OK
{ "leader": "actor-ts://my-app@10.0.0.5:2552", "isSelf": true }

The leader’s address plus isSelf (is the local node the leader). leader is null before a leader is elected. Useful for monitoring leadership churn.

GET /cluster/shards?type=cart
→ 200 OK
{
"typeName": "cart",
"leader": "my-app@10.0.0.5:2552",
"version": 7,
"takenAt": 1710000000000,
"regions": [
{
"key": "my-app@10.0.0.5:2552|/system/cluster/sharding/region-cart",
"address": "my-app@10.0.0.5:2552",
"path": "/system/cluster/sharding/region-cart",
"proxy": false,
"shards": [0, 1, 2, ..., 33]
},
...
],
"shardHome": [
{ "shard": 0, "regionKey": "my-app@10.0.0.5:2552|/system/cluster/sharding/region-cart" },
...
]
}

Shows the shard-to-region allocation for a sharded type. Use for:

  • Verifying even distribution after a rebalance.
  • Diagnosing hot regions (one region with too many shards).
  • Manual rebalance triggers in development.

Where the answer comes from. The node answers from ClusterSharding.shardMap(typeName) — the last map the coordinator broadcast to this node’s region. Nothing needs configuring: no DistributedData extension, no coordinatorStateStore, no round trip to the leader. The one precondition is that the answering node takes part in the type — it has called sharding.start(...) or sharding.startProxy(...) for it — because the coordinator broadcasts only to regions that registered. A node that never mentioned cart has nothing to report and answers 404.

This endpoint used to read the coordinator’s DistributedData snapshot, which made a 200 unreachable in a default configuration: nothing in the framework starts that extension, and the snapshot is only written when you opt into a coordinatorStateStore. The store is still available and still optional — it shortens leader failover — but this endpoint no longer depends on it.

Reading the body:

  • version counts coordinator broadcasts, not shard assignments — one bump per publish, and a publish coalesces a burst of moves. Compare it across two readouts to tell “the map moved” from “the map is the same”.
  • takenAt and leader are stamped by the answering node when the map arrived, so two nodes can report the same version with slightly different timestamps.
  • regions[].key is <address>|<region path>, the same key shardHome[].regionKey uses. A region that hosts nothing is still listed, with an empty shards — “registered and given no shards” is the interesting half of a rebalance question.
  • A 200 with an empty shardHome is normal, not an error: regions fills as soon as regions register, while a shard only gets a home once an entity in it is addressed.

Opt-in via enableMetricsEndpoint: true.

GET /metrics
→ 200 OK
Content-Type: text/plain; version=0.0.4; charset=utf-8
# HELP actor_messages_processed_total ...
# TYPE actor_messages_processed_total counter
actor_messages_processed_total{class="Worker",path="..."} 12345
...

Prometheus text format. See Prometheus exporter.

Answers 503 instead when the installed MetricsRegistry cannot be read back through collect() — the prom-client bridge is the one in the framework that cannot, because it forwards its writes to prom-client and keeps no snapshot:

GET /metrics
→ 503 Service Unavailable
metrics endpoint unavailable: the installed MetricsRegistry does not
support collect() — ...

A 200 with an empty body would be worse than the error: an empty exposition is a valid scrape, so Prometheus records the target as up and every framework series quietly stops existing. The check runs per request, because the registry can be swapped at any point in the system’s life.

Opt-in via enableLeaveEndpoint: true.

POST /cluster/leave
→ 202 Accepted
leaving

The 202 body is the plain-text word leaving (not JSON) — the leave is fire-and-forget.

Triggers graceful cluster-leave. The node transitions through leaving → exiting → removed; shards rebalance away; the actor system terminates (configurable).

Use for:

  • Pod retirement before rolling-update.
  • Manual node-out during incidents.

Gate behind authentication — anyone with port access can drain your node.

Opt-in via enableDownEndpoint: true.

POST /cluster/down
{ "address": "actor-ts://my-app@10.0.0.5:2552" }
→ 202 Accepted
{ "downed": "actor-ts://my-app@10.0.0.5:2552" }

Returns 202 with { downed } when the member was downed, or 404 if the address is unknown or already terminal.

Force-downs a remote member by address. Destructive — the target’s actors stop, its shards reallocate elsewhere.

Use for:

  • Split-brain recovery when no downing strategy is configured.
  • Removing stuck unreachable members that refuse to recover.

High-risk endpoint — gate behind auth + audit logs.

import { path, get, concat, completeJson } from 'actor-ts/http';
import { managementRoutes } from 'actor-ts/management';
const baseRoutes = managementRoutes(system, cluster);
const customRoutes = concat(
baseRoutes,
path('admin',
get(async () => completeJson(200, { appVersion: '1.2.3' })),
),
);
await system.http(port, { host }).bind(customRoutes);

managementRoutes(system, cluster, options) returns the base routes; combine with your own using concat. Useful for adding app-specific admin endpoints alongside the standard set.

The read endpoints (/health, /ready, /cluster/members, /cluster/leader, /cluster/shards) and /cluster/down return JSON. /metrics returns Prometheus text/plain, and /cluster/leave returns the plain-text body leaving.

Errors are plain-text bodies (the framework’s complete(status, message)), not a structured object — e.g. a missing type query param yields a 400 whose body is missing query param \type“.

Status codes follow HTTP conventions: 200 for reads, 202 Accepted for the admin actions (/cluster/leave, /cluster/down), 4xx for client errors (404 when a /cluster/down address is unknown), and 503 when the cluster or a health check is unavailable.

The privileged endpoints (/cluster/down, /cluster/leave, /cluster/members, /cluster/shards, /metrics) accept anonymous requests by default. That’s only safe when the management port is on a network-isolated bind (a separate Service, a different port, 127.0.0.1-only) or behind a reverse-proxy that does authentication for you.

For production deployments that expose management endpoints to a broader network, attach the built-in middlewares via managementRoutes(system, cluster, { auth, ipAllowlist }):

import {
BearerTokenAuth,
IpAllowlist,
} from 'actor-ts/http';
import {
managementRoutes,
} from 'actor-ts/management';
const routes = managementRoutes(system, cluster, {
enableLeaveEndpoint: true,
enableDownEndpoint: true,
enableMetricsEndpoint: true,
// Shared-secret bearer token (rotation supported via multiple entries).
auth: BearerTokenAuth({
tokens: [process.env.MGMT_TOKEN!],
realm: 'my-app-mgmt',
}),
// Network-level fence — only requests from these CIDRs reach
// the application at all.
ipAllowlist: IpAllowlist({
allow: ['10.0.0.0/8', '127.0.0.1/32'],
// Behind a reverse proxy, name the proxies: the client address
// is then resolved from the forwarded chain's socket end
// inwards. Without this the default reads the socket peer,
// which behind a proxy is the proxy.
// trustedProxies: ['10.9.9.0/24'],
}),
});

Policy split:

  • auth wraps the privileged subset (/cluster/*, /metrics). /health and /ready stay anonymous — Kubernetes liveness/readiness probes can’t easily attach an Authorization header. Set authProtectHealth: true when your deployment guarantees the probes can present credentials.
  • ipAllowlist wraps EVERY endpoint including /health and /ready. Network-level isolation is independent of who’s authenticated.

Failure modes:

  • Missing / wrong Authorization header → 401 Unauthorized with WWW-Authenticate: Bearer realm="...".
  • Client IP outside the allowlist (or unresolvable) → 403 Forbidden.

The middlewares are general-purpose — you can use them outside the management subtree via the withMiddleware(mw, route) builder.