Aller au contenu
Français

Encryption

Ce contenu n’est pas encore disponible dans votre langue.

Object-storage payloads support client-side AES-GCM encryption — the framework encrypts before put, decrypts on get, keys managed as a single master key or a versioned key ring for rotation.

import {
ObjectStorageDurableStateStore,
ObjectStorageDurableStateStoreOptions,
S3ObjectStorageBackend,
S3ObjectStorageOptions,
} from 'actor-ts/persistence';
const masterKey = Buffer.from(process.env.MASTER_KEY_V1!, 'base64'); // 32 bytes
const objectStorageDurableStateStoreOptions = ObjectStorageDurableStateStoreOptions.create()
.withBackend(new S3ObjectStorageBackend(S3ObjectStorageOptions.create() /* .withRegion(...).withBucket(...) */))
.withEncryption({ mode: 'client-aes256-gcm', masterKey, info: 'acme/prod/durable-state/v1' });
const store = new ObjectStorageDurableStateStore(objectStorageDurableStateStoreOptions);

Now every persisted state is AES-GCM-encrypted before upload — under a per-persistenceId subkey derived from the master key. Reads transparently decrypt.

info is required and deployment-specific. It is not decoration — see Choosing info below.

Object stores (S3, GCS, Azure Blob) offer server-side encryption built-in. Why also encrypt client-side?

ThreatServer-sideClient-side
Storage compromise (someone reads from disk)✓✓
Account compromise — reading (someone has S3 credentials)✗✓
Account compromise — writing (see below)✗partly
Cloud provider compromise✗✓
Audit / compliance requiring “we hold the keys”✗✓

Client-side encryption protects against more threats but costs more (CPU per op, key-management overhead). Most apps should use both: server-side as baseline + client-side for sensitive payloads.

An attacker who can write to the bucket is a different attacker from one who can only read it, and encryption answers only half of what they can do. They cannot forge a body — that needs the key — but they can move and replay the authentic ones they find, and the auth tag travels with the bytes.

Two shapes of that, both closed since #612 and neither by encryption alone:

  • Replaying a body onto a different object. One persistenceId’s state copied onto another’s key, or one snapshot copied over a later sequence number. The storage key is now authenticated alongside the body, so the copy no longer verifies — see Binding a body to its key.
  • Replaying an older body over a newer one. Same key, same everything, just stale. No authenticator can see this: the revision lives inside the authenticated bytes. The durable-state store keeps an in-process floor instead — see Rollback protection.
type EncryptionConfig =
| { mode: 'none' }
| { mode: 'sse-s3' } // server-side, S3-managed
| { mode: 'sse-kms'; kmsKeyId: string } // server-side, KMS-managed
| { mode: 'client-aes256-gcm'; masterKey: Uint8Array; info: string }
| { mode: 'client-aes256-gcm'; masterKeys: MasterKeyRing; info: string };
type MasterKeyRing = {
active: MasterKeyRingEntry; // new writes encrypt under this
retired?: MasterKeyRingEntry[]; // older keys, kept for decryption
};
type MasterKeyRingEntry = {
version: number; // 0..255, embedded in the body manifest
key: Uint8Array; // 32 bytes (AES-256)
};

Client-side encryption uses mode: 'client-aes256-gcm' with either a single masterKey (32 bytes) or a masterKeys ring. The ring carries:

  • active — the key new writes encrypt under.
  • retired — older keys, kept to decrypt historical blobs.

Each entry’s version byte travels in the body manifest so decrypt can pick the matching key. Carrying the active key plus retired keys is the foundation of key rotation.

info is HKDF’s context-binding input (RFC 5869 §3.2). The subkey for a blob is derived from three things: the master key, the persistenceId (as HKDF salt), and info. Change any one and you get an unrelated key.

It is required, and there is no default, deliberately. A shared default would mean any two deployments holding the same master key derive byte-for-byte the same subkey for the same persistenceId — so a staging environment restored from a production dump, or a DR region, could read production’s blobs and nothing in the config would say so. That is a decision only the operator can make, so the framework insists you make it.

Encode environment + purpose + version, most specific first:

'acme/prod/snapshot/v1'
'acme/staging/snapshot/v1'
'acme/prod/durable-state/v1'
  • Different environments MUST differ, even on the same master key. This is the whole point.
  • Different payload kinds SHOULD differ (snapshots vs. durable state), so one compromised derivation context does not extend to the other.
  • A trailing version gives a future context rotation somewhere to go.

info is not recorded on the wire. Unlike the key version, no manifest byte records which info a blob was written under. Changing it makes every existing blob undecryptable until a sweep rewrites them — see rotating the context. Pick the value before the first write.

On put:
serialize value → compress → derive per-pid subkey (HKDF from active key)
→ AES-GCM(bytes, subkey, iv) → ciphertext
→ body manifest "ATS1" { flags, keyVersion, iv, ciphertext }
→ S3.put(body) // no key-id metadata header
On get:
S3.get → body manifest { flags, keyVersion, iv, ciphertext }
→ pick master key by keyVersion (active or a retired entry)
→ derive per-pid subkey (HKDF) → AES-GCM decrypt → decompress
→ deserialize

The key version is embedded in the body manifest — every blob records which key version it was encrypted under. This lets the framework decrypt old payloads using the right key even after the active key has been rotated (retired keys stay in the ring).

  • The body — serialized state / event / snapshot.
  • Not — the object key, object metadata headers, the bucket name.

For object metadata that shouldn’t leak (sensitive persistenceIds), use a separate naming scheme (hash IDs before they become object keys).

The master key bytes come from somewhere. Common patterns:

const masterKey = Buffer.from(process.env.MASTER_KEY_V1!, 'base64'); // 32 bytes
// → .withEncryption({ mode: 'client-aes256-gcm', masterKey, info: 'acme/prod/snapshot/v1' })

Simplest. Each key is a 32-byte buffer (for AES-256-GCM), base64-encoded in the env.

Risk: env vars are visible to anything that can read the process environment. Use only when the env is itself secured (K8s secrets, etc.).

import { KMS } from '@aws-sdk/client-kms';
const kms = new KMS();
const decrypted = await kms.decrypt({
KeyId: 'alias/master',
CiphertextBlob: Buffer.from(process.env.WRAPPED_KEY!, 'base64'),
});
const masterKey = decrypted.Plaintext!; // 32 bytes, held in memory
// → .withEncryption({ mode: 'client-aes256-gcm', masterKey, info: 'acme/prod/snapshot/v1' })

Master key is stored encrypted under a cloud KMS key. The app fetches it on startup, decrypts via KMS, holds in memory.

Better than raw env vars — only KMS access is required to recover keys.

Similar pattern: pull master keys from Vault at startup.

AES-GCM is fast — modern CPUs have hardware support.

Per 100 KB encrypt + decrypt:

  • ~0.5-1 ms on modern x86 / Apple Silicon.
  • Effectively free on small objects.

For most workloads, encryption is invisible in profiles.

encrypt → compress → S3.put # ✗ no compression benefit on ciphertext
compress → encrypt → S3.put # ✓ this is what the framework does

The framework’s order is compress first, then encrypt — compressed bytes are still compressible (not random); after encryption, they’re effectively random and uncompressable.

If you set both compression and encryption, you get this order automatically.

After enabling encryption on a previously-unencrypted bucket:

state/cart-42 ← old, plaintext (encrypted flag unset in the manifest)
state/cart-43 ← new, encrypted (encrypted flag set, key version 0)

The framework detects each per-payload from the body manifest:

  • Encrypted flag unset → plaintext path.
  • Encrypted flag set → decrypt using the key version the manifest records.

Means you can enable encryption gradually — new writes get encrypted, old reads still work, and a background re-encryption sweep can migrate the rest.

See key rotation for the rotation flow.

Encryption is not integrity. An unencrypted body is plain JSON sitting in a bucket: anyone who can write the object can edit it — and for ObjectStorageDurableStateStore that includes the revision field the compare-and-swap check is built on.

Opt in with an integrity config. Its key is 32 bytes and separate from the encryption master key, because the threat here is tampering, not disclosure:

const integrityKey = Buffer.from(process.env.INTEGRITY_KEY_V1!, 'base64'); // 32 bytes
const objectStorageDurableStateStoreOptions = ObjectStorageDurableStateStoreOptions.create()
.withBackend(backend)
.withIntegrity({ mode: 'hmac-sha256', integrityKey });
const store = new ObjectStorageDurableStateStore(objectStorageDurableStateStoreOptions);

Every write now appends an HMAC-SHA256 tag (truncated to 16 bytes) over the whole framed body — the manifest header included, so the compression and encryption flags are covered as well — and every read verifies it before a single payload byte reaches the store layer.

ObjectStorageSnapshotStore takes the same option, and the one-call wiring hands it to both stores at once:

const objectStoragePluginOptions = ObjectStoragePluginOptions.create()
.withBackend({ kind: 's3', bucket: 'my-app', region: 'eu-central-1' })
.withIntegrity({ mode: 'hmac-sha256', integrityKey });
const { durableStateStore } = await registerObjectStoragePlugins(ext, objectStoragePluginOptions);

A snapshot is arguably the more valuable target of the two: recovery folds events on top of it, so whoever can rewrite one dictates the state an actor comes back as. Replay bounds the sequenceNr a snapshot may claim — a snapshot ahead of the journal is refused — but nothing else authenticates the state payload.

Configuring integrity makes a tag mandatory on read. A body that arrives without one is refused:

BodyCodec: body carries no integrity tag but an integrityKey was
supplied for decoding.

That is the point of the control rather than a rough edge. The bit that records “this body has a tag” lives in the body, so whoever can rewrite the object can also clear it and drop the 16 tag bytes. If an absent tag meant “skip verification”, the check would protect nobody: stripping a tag is far easier than forging one.

Migrating a bucket written before integrity

Section titled “Migrating a bucket written before integrity”

Bodies written before you turned integrity on carry no tag, so they meet the same rule. Open the migration window explicitly:

const objectStorageDurableStateStoreOptions = ObjectStorageDurableStateStoreOptions.create()
.withBackend(backend)
.withIntegrity({ mode: 'hmac-sha256', integrityKey })
.withAllowUntaggedBodies(true);

Then rewrite every object — a load followed by an upsert per persistenceId re-frames it with a tag — and remove the option again. While it is set, an untagged body is accepted from anyone; a body that does carry a tag is still verified, but the downgrade route is open for as long as the window is.

Snapshots migrate on their own: keepN prunes the untagged ones as new tagged snapshots are taken, so the window can usually close after keepN saves per persistenceId instead of after an explicit sweep.

The two compose. Hand the sweep the same integrity configuration the store has — a flat config or the very same per-persistenceId resolver — and every tag is verified on the way in and recomputed on the way out:

await reEncryptObjectStorage(backend, {
keyPrefix: 'state/',
keyring,
info: 'acme/prod/durable-state/v1',
integrity: { mode: 'hmac-sha256', integrityKey },
});

Leave integrity out on a tagged corpus and the sweep refuses before it touches anything, naming an offending key: without the key it can neither read a tagged body nor write one back, and finding that out object by object would leave the corpus half-rotated (#739).

Two things the sweep deliberately does not do:

  • It never adds a tag to an untagged body. It re-seals only what already carried one. Promoting the rest would make every rewritten object unreadable to any reader you have not yet given the integrity key to — and a rotation runs while the application is serving. Turning integrity on corpus-wide is the migration above, not a side effect of rotating a key.

  • It does not rotate the integrity key itself, and there is no window that lets you. There is no version byte for the integrity key, so a half-swept corpus would be indistinguishable from a tampered one — and allowUntaggedBodies does not help, because it re-admits bodies carrying no tag. A body tagged under the old key still carries FLAG_INTEGRITY_HMAC and fails the comparison before that branch is reached, with the new key and without it alike. The sweep is fail-closed here rather than silently half-rolling, which is the right direction, but it means there is no sweep-shaped way to change the key.

    What does work is the per-call PersistenceOptions.integrity override, read and write independently: load(pid, { integrity: old }) then upsert(pid, revision, state, { integrity: new }). It is per-persistenceId, it bumps every entity’s revision through the CAS, and it has no snapshot equivalent. #1354 carries the shape a real roll would need.

Sweeping a corpus that is mid-integrity-migration — some bodies tagged, some not — needs allowUntaggedBodies: true on the sweep as well, for the same reason the store needs it. Without it the sweep stops at the first untagged body rather than assume the tag was never there.

A tag answers “did the holder of the key produce these bytes?” and nothing else. It does not answer “were they produced for this object?” — so an authentic body moved to another storage key used to verify there just as well.

That mattered most in the plain-HMAC configuration. integrityKey is one flat secret for the whole deployment, with no per-persistenceId derivation of any kind, and load returns the requested persistenceId with the body’s state. So one account’s object copied onto another account’s key came back as that other account’s state. Client-side encryption narrowed this without closing it: HKDF salts the subkey with the persistenceId, which separates two pids but not two objects of one pid — so a snapshot could still be replayed onto a different sequence number of the same actor.

Every write now binds the storage key. It goes into AES-GCM’s additional authenticated data on an encrypted body and, length-prefixed, into the HMAC input on a tagged one; a manifest bit records that it was done. Nothing to configure — both object-storage stores know the key they are writing to.

Bodies written before this existed carry no binding, and they still decode: an old bucket keeps working. That is also the gap, because the bit that says “this body is bound” is a manifest byte like any other, so one authentic pre-binding body remains a replay token for every key in the bucket until unbound bodies stop being accepted:

const objectStorageDurableStateStoreOptions = ObjectStorageDurableStateStoreOptions.create()
.withBackend(backend)
.withIntegrity({ mode: 'hmac-sha256', integrityKey })
.withRequireContextBinding();

Rewrite the corpus first, or reads will start failing. For durable state that is a load + upsert per persistenceId; snapshots reach it through keepN pruning; a bucket that is mid-rotation gets there via reEncryptObjectStorage, which rebinds every body it passes and no longer skips one merely because its key version is current. That includes the unencrypted-plus-HMAC configuration, where there is no master key to rotate at all: such a body is re-framed for its key and then left alone on every later pass.

One replay stays invisible to every authenticator: the same body at the same key, just an older one. It was written by the legitimate writer, its tag is genuine, and the revision an attacker wants to roll back sits inside the bytes the tag covers. There is nothing to detect — except that this process has already seen a higher revision.

ObjectStorageDurableStateStore remembers that highest revision per persistenceId and refuses to load below it. It is on by default:

const objectStorageDurableStateStoreOptions = ObjectStorageDurableStateStoreOptions.create()
.withBackend(backend)
.withRejectRevisionRollback(false); // opt out

Two limits worth knowing:

  • It is per store instance, not per process. A restart forgets the floor — and so does anything else that builds a new store, because every node constructs its own. A shard rebalance therefore hands the next writer a cold floor with no restart anywhere. A durable floor needs trusted state outside the bucket, which the framework has no seam for yet.
  • It fires on a legitimate delete-and-recreate from another writer. A recreated record restarts at revision 1. Deleting through this store drops the floor with it, so this only applies across processes — and it is the case withRejectRevisionRollback(false) exists for.