콘텐츠로 이동
한국어

Master key rotation

이 콘텐츠는 아직 번역되지 않았습니다.

For at-rest encryption (object-storage encryption, durable-DD-encryption), the framework’s keys live in a MasterKeyRing — a versioned set of keys: one active key plus any number of retired ones, each tagged with a numeric version (a single byte, 0-255). Rotation is online: new writes go under the active key; old reads still work under whatever version they were encrypted with; a background sweep eventually re-encrypts older data.

import type { MasterKeyRing } from 'actor-ts/persistence';
// Each key is 32 raw bytes (AES-256) — decode from a base64 env
// var, a mounted secret, or a KMS unwrap (see "Storage of master
// keys" below).
const keyRing: MasterKeyRing = {
active: { version: 2, key: Buffer.from(process.env.MASTER_KEY_V2!, 'base64') }, // new — current
retired: [{ version: 1, key: Buffer.from(process.env.MASTER_KEY_V1!, 'base64') }],
};

MasterKeyRing is a type, not a class — you build the plain object above; there is nothing to new. active is the key used for new writes. Reads dispatch on the version byte stored in each blob’s manifest, matching it against active or one of the retired entries.

Because that byte is the reader’s only handle on the key, each entry needs its own version number. A ring with the same version on two entries is rejected — at registration, at the store, and at the start of a sweep — rather than left to be resolved by lookup order. See every version is used once.

The one-byte field caps how many versions can be live at once, not how often you may rotate: once a sweep has moved the whole corpus onto the active version, the retired entries go away and their numbers become reusable. Registration warns from version 240 on so there is time to schedule that.

Three triggers:

  1. Scheduled rotation — a security policy (every 90 days, yearly).
  2. Suspected compromise — leaked key material; rotate immediately.
  3. Compliance — regulatory requirements mandating periodic rotation.

Even without a specific trigger, periodic rotation is good practice — limits blast radius of an undetected leak.

1. Generate a fresh key with the next version number. Add it
to keyRing.retired; do NOT yet promote it to active.
2. Roll out the updated ring to all nodes. Verify reads work
— old data still decrypts, writes still use the current
active key.
3. Promote the fresh key to active (move the previous active
into retired). Roll out. New writes use the new key; old
data still readable.
4. Run the re-encryption sweep. Old data is read, decrypted,
re-encrypted under the active key.
5. Once the sweep completes, drop the old key from retired.
Keep a rollback window of ~7 days where the old key is
still available; after that, drop it.

The framework supports each of these steps without downtime.

const keyRing: MasterKeyRing = {
active: { version: 1, key: Buffer.from(process.env.MASTER_KEY_V1!, 'base64') }, // still v1 — active unchanged
retired: [{ version: 2, key: Buffer.from(process.env.MASTER_KEY_V2!, 'base64') }], // v2 added to the ring, not yet active
};

Roll out this config to every node. Reads of v1-encrypted data still work; writes still use v1 — but every node now holds v2 and can decrypt under it.

This step is safe and reversible — if v2 isn’t actually needed yet, revert by removing it from retired.

const keyRing: MasterKeyRing = {
active: { version: 2, key: Buffer.from(process.env.MASTER_KEY_V2!, 'base64') }, // ← now v2
retired: [{ version: 1, key: Buffer.from(process.env.MASTER_KEY_V1!, 'base64') }],
};

Roll out. New writes go under v2. Reads consult the ring and find the right key (v1 or v2) based on the version byte in the stored blob’s manifest.

After this step, gradually new data accumulates under v2 as the workload writes. Old data stays under v1 until re-encrypted.

import { reEncryptObjectStorage } from 'actor-ts/persistence';
// `backend` is the ObjectStorageBackend your store writes through
// (a FilesystemObjectStorageBackend, S3ObjectStorageBackend, …).
const result = await reEncryptObjectStorage(backend, {
keyPrefix: 'snapshots/', // which objects to sweep
keyring: keyRing, // active = v2, retired = [v1]
info: 'acme/prod/snapshot/v1', // the store's HKDF context
});
console.log(`re-encrypted ${result.rewrote} of ${result.scanned} objects`);
console.log(`skipped, still under the old key: ${result.skippedMalformedKey}`);

keyPrefix must be the store’s own prefix, exactly. A shorter one shifts every key’s persistence-id segment by a level, and the sweep refuses the whole corpus rather than salt on the wrong segment.

info is required and must be the exact value the encrypting store uses. It is the HKDF context, i.e. half of what the subkey is derived from — a wrong value fails every decrypt rather than silently producing wrong output.

The sweep lists every object under keyPrefix and, for each one not already at the ring’s active version, decrypts it and re-encrypts under active. Objects already current are skipped without a write.

Useful options:

  • skip — a (key) => boolean predicate; matching keys are left untouched (exclude non-body objects, or scope a partial rotation).
  • onProgress — per-object callback for logging or an operator dashboard on long sweeps.
  • progress — a ReEncryptProgressStore (e.g. InMemoryReEncryptProgressStore) for crash-resume of very large sweeps.
  • verifyKeyringCompleteness — on by default: samples some blobs and refuses to start if any references a version missing from the ring.
  • newInfo — rotates the HKDF context instead of (or alongside) the key; see rotating the HKDF context.
  • integrity — the integrity configuration the corpus was written under, in the same shape the store takes. Required on a bucket that has the integrity HMAC turned on: the tag covers the manifest bytes, so the sweep verifies it on read and recomputes it on write, and without the key it refuses before the first write rather than aborting mid-corpus. See sweeping an integrity-protected corpus.
  • allowUntaggedBodies — re-admits untagged bodies while integrity is set, for a bucket still part-way through turning integrity on. Off by default, deliberately: from the sweep’s position a tag that was never written and a tag that was stripped look the same.

The sweep is idempotent + resumable — an object already at the active version is skipped without a write, so re-running after an interruption is safe. By default a resumed run re-lists and re-checks every key; pass a progress store to skip straight past the objects already done.

The sweep does not hand back a result you have to audit. If any object’s key could not yield a usable persistence id, it finishes rotating everything else and then throws ReEncryptIncompleteError — because every one of those objects is still encrypted under the key Step 4 is about to drop, and a counter you have to remember to read is not a safeguard.

import { ReEncryptIncompleteError } from 'actor-ts/persistence';
try {
const result = await reEncryptObjectStorage(backend, { /* … */ });
// Reaching here means skippedMalformedKey === 0. Step 4 is cleared.
} catch (thrown) {
if (thrown instanceof ReEncryptIncompleteError) {
console.error(thrown.malformedKeys); // a bounded sample of the offenders
console.error(thrown.result); // what the pass did manage to rotate
}
throw thrown; // do NOT proceed to Step 4
}

Reaching that catch means one of three things: a key written out-of-band, a key written by a version older than the one that started rejecting control characters on put, or a keyPrefix that does not match the store’s own prefix. Correct the keys, or exclude genuinely foreign objects with skip, then re-run.

Precondition: the Step 3 sweep returned. It throws rather than returns while any object is still under the retired key, so a sweep that completed normally — and only that — clears this step. If you are inspecting a result object instead, the field to check is skippedMalformedKey, and it must be 0.

After the sweep completes (every item encrypted under v2):

const keyRing: MasterKeyRing = {
active: { version: 2, key: Buffer.from(process.env.MASTER_KEY_V2!, 'base64') },
// retired dropped — v1 is gone
};

Drop v1 entirely. Any data still encrypted under v1 (e.g., backups that haven’t been re-encrypted) is now unreadable.

Wait a rollback window before dropping. ~7 days lets you recover from “oh no, the sweep didn’t actually cover all the backups.” After confirmed migration, drop v1.

The keyRing doesn’t ship a key-storage backend. Common patterns:

SourcePattern
Env varsprocess.env.MASTER_KEY_V2 — simplest, fine for tests.
K8s secretsMounted as files; read at startup.
HashiCorp VaultPull dynamically at startup; refresh periodically.
AWS KMS / GCP KMS / Azure Key VaultCloud KMS APIs. Decrypt-on-load via the KMS encryption keys.

For production, KMS is the right answer — keys never leave the secure boundary in plaintext form.

If multiple clusters share the same encrypted store (e.g., a DR replica that reads the primary’s backups), all clusters need the same keyRing. Rotate them together; don’t let one cluster fall behind on key generations.