Ir al contenido
Español

Kubernetes deployment

Esta página aún no está disponible en tu idioma.

Kubernetes is the most common deployment target. The framework plays well with K8s once you get a few things right: stable identity (for stateful actors), seed discovery (so nodes find each other), and a clean shutdown path (so rolling updates don’t drop traffic).

This page is a working recipe — copy, adapt, deploy.

apiVersion: v1
kind: ServiceAccount
metadata:
name: actor-ts
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: actor-ts-pod-reader
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: actor-ts
subjects:
- kind: ServiceAccount
name: actor-ts
roleRef:
kind: Role
name: actor-ts-pod-reader
apiGroup: rbac.authorization.k8s.io
---
apiVersion: v1
kind: Service
metadata:
name: actor-ts-cluster
spec:
clusterIP: None # headless — DNS returns pod IPs
selector:
app: actor-ts
ports:
- name: cluster
port: 2552
targetPort: 2552
---
apiVersion: v1
kind: Service
metadata:
name: actor-ts
spec:
selector:
app: actor-ts
ports:
- name: http
port: 80
targetPort: 8080
- name: management
port: 8558
targetPort: 8558
---
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: actor-ts
spec:
serviceName: actor-ts-cluster
replicas: 3
selector:
matchLabels:
app: actor-ts
template:
metadata:
labels:
app: actor-ts
spec:
serviceAccountName: actor-ts
terminationGracePeriodSeconds: 30
containers:
- name: app
image: ghcr.io/your-org/your-app:1.2.3
ports:
- name: cluster
containerPort: 2552
- name: http
containerPort: 8080
- name: management
containerPort: 8558
env:
- name: ACTOR_TS_HOSTNAME
valueFrom:
fieldRef:
fieldPath: status.podIP
- name: ACTOR_TS_PORT
value: "2552"
- name: K8S_NAMESPACE
valueFrom:
fieldRef:
fieldPath: metadata.namespace
- name: K8S_LABEL_SELECTOR
value: "app=actor-ts"
- name: DB_PASSWORD
valueFrom:
secretKeyRef:
name: actor-ts-secrets
key: db-password
readinessProbe:
httpGet:
path: /ready
port: management
initialDelaySeconds: 5
periodSeconds: 5
livenessProbe:
httpGet:
path: /health
port: management
initialDelaySeconds: 15
periodSeconds: 10
lifecycle:
preStop:
exec:
# Drain LB before SIGTERM hits the app
command: ["/bin/sh", "-c", "sleep 10"]
ServiceAccount: actor-ts
Role: actor-ts-pod-reader # get + list pods
RoleBinding: binds them

The K8s API seed provider needs to list pods matching a label selector to discover peers. Without this RBAC grant, the seed provider gets 403 and the cluster never forms.

clusterIP: None

A headless service returns the pod IPs directly via DNS (no ClusterIP virtual address). Useful when nodes need stable, direct peer identities — the failure detector’s heartbeats target specific pod IPs, not a load-balanced abstraction.

clusterIP: <default>

A normal Service for HTTP and the management endpoint — these benefit from load-balancing. Cluster traffic goes through the headless service, app traffic through this one.

UseWhen
StatefulSetStable pod names (actor-ts-0, actor-ts-1, …). Useful when you want predictable identity for entity placement, or when persistent volumes are mounted per-pod.
DeploymentPod names are random. Fine if your app is stateless (no per-pod identity required) and persistence is external (Cassandra journal, shared S3 snapshot store).

For sharded actors with remember-entities = true on persistent volumes per pod, StatefulSet is the right choice. For externally-persisted state (cluster talking to a shared Postgres/Cassandra), Deployment is fine and simpler.

terminationGracePeriodSeconds: 30

K8s sends SIGTERM, then waits this long before SIGKILL. Sized based on:

  • HTTP drain — typically 5-10 s.
  • Cluster leave gossip — 5-15 s for convergence.
  • Journal flush — depends on the journal.

30 s is a reasonable default. Bump it if your cluster is large or the failure-detector window is long.

preStop:
exec:
command: ["/bin/sh", "-c", "sleep 10"]

Critical for clean rolling updates. The flow:

  1. K8s marks the pod terminating and starts the preStop hook in parallel with the load-balancer-deregistration.
  2. sleep 10 — gives the load balancer time to stop sending new traffic to this pod.
  3. After the sleep, K8s sends SIGTERM.
  4. The app’s coordinated-shutdown hooks drain in-flight requests, leave the cluster, etc.

Without the sleep, SIGTERM races with LB deregistration — in-flight requests can see “draining” responses.

readinessProbe: /ready
livenessProbe: /health

The framework’s management routes expose these endpoints.

  • /ready — “should a load balancer send this pod traffic?” Already gated on the framework’s own checks: cluster-membership (this node is up) and cluster-transport (it is not cut off from every peer). Add your per-app checks — database reachable, dependencies warm — with healthChecksOf(system).addReadiness. In-process, Cluster.bootstrap gates the same way: actor-ts.cluster.bootstrap.minimum-members (or awaitReady: { minimumMembers } in code) keeps a resolved bootstrap from meaning anything less than “the cluster reached its expected size” — set it to the replica count, exactly like required-contact-points.
  • /health — “would restarting this pod help?” Failing means K8s restarts it, so it depends on nothing outside the process — the framework’s only liveness check is actor-system. Never put a database or a downstream service here: a shared outage would restart the whole fleet, and the restarts would not fix it.
import { ActorSystem } from 'actor-ts';
import { Cluster, ClusterOptions } from 'actor-ts/cluster';
import { KubernetesApiSeedProvider, KubernetesApiSeedProviderOptions } from 'actor-ts/discovery';
import { managementRoutes } from 'actor-ts/management';
const system = ActorSystem.create('my-app');
// 1. Cluster join with K8s API seed discovery
const kubernetesApiSeedProviderOptions = KubernetesApiSeedProviderOptions.create()
.withNamespace(process.env.K8S_NAMESPACE!)
.withServiceName(process.env.K8S_SERVICE_NAME!)
.withSystemName(system.name)
.withPort(2552);
const seeds = await new KubernetesApiSeedProvider(kubernetesApiSeedProviderOptions).lookup();
const clusterOptions = ClusterOptions.create()
.withHost(process.env.ACTOR_TS_HOSTNAME!)
.withPort(parseInt(process.env.ACTOR_TS_PORT!))
.withSeeds(seeds.map(a => a.toString()))
.withRoles(['compute']);
const cluster = await Cluster.join(system, clusterOptions);
// 2. Management endpoints
const mgmtRoutes = managementRoutes(system, cluster);
await system.http(8558).bind(mgmtRoutes);
// 3. App HTTP server
const http = system.extension(HttpExtensionId);
await http.newServerAt('0.0.0.0', 8080).bind(routes);
// 4. Run until the kubelet says stop, then shut down in order
await system.runUntilTerminated();

Four pieces in order:

  1. Cluster join with seed discovery — the K8s API seed provider queries pods matching app=actor-ts and uses their IPs as seeds. On the first pod, the seed list is just itself (the auto-promote-to-leader path).
  2. Management endpoints — /health + /ready for K8s probes, /cluster/members for debugging.
  3. App HTTP — your routes, separate port from management.
  4. Run until terminated — installs the SIGTERM/SIGINT handlers, blocks, then runs the pipeline and resolves when the system is down. There is no separate teardown step to write: both bind() calls registered their unbind in the service-unbind phase and Cluster.join registered the leave in cluster-leave, so the pod drops out of the Service endpoints and out of the cluster before the actors stop.

See Discovery — Kubernetes API for the seed provider’s full options.

Terminal window
kubectl rollout restart statefulset/actor-ts

For each pod, in order (StatefulSet) or arbitrary (Deployment):

  1. K8s marks the pod terminating + starts preStop.
  2. 10-second LB drain.
  3. SIGTERM lands.
  4. Coordinated-shutdown runs:
    • Stop accepting new HTTP requests (service-unbind, wired by bind()).
    • Drain in-flight requests (service-requests-done, yours).
    • Close broker connections (service-stop, wired by each broker actor).
    • Issue cluster.leave() (cluster-leave, wired by Cluster.join).
    • Terminate the actor system (actor-system-terminate).
  5. runUntilTerminated() resolves, its signal handlers come off, and the process exits cleanly.
  6. K8s starts a new pod from the new image.
  7. New pod joins the cluster via the seed provider.

For sharded entities, rebalancing happens automatically — the leaving node’s shards are reallocated; new entities re-spawn on the new pod from the journal.