Cloudflare Containers have added a public‑beta snapshot API that lets you capture the entire filesystem of a running container and later restore it when the container is started again. This capability is directly relevant to AI, cloud/platform, DevOps/SRE, and security engineers because it provides a deterministic way to pause, migrate, or restart workloads without losing any on‑disk state.
How the container snapshot API works
The API is exposed through the Durable Object Container interface. Calling snapshotContainer() on the container object returns a handle that represents the current filesystem. You can store that handle in Durable Object storage and pass it back to start() when you want to re‑hydrate the container.
import { DurableObject } from "cloudflare:workers";
export class MyDurableObject extends DurableObject {
async saveSnapshot() {
// Capture the running container's filesystem.
const containerSnapshot = await this.ctx.container.snapshotContainer({});
await this.ctx.storage.put("containerSnapshot", containerSnapshot);
}
async restoreSnapshot() {
const containerSnapshot = await this.ctx.storage.get("containerSnapshot");
if (!containerSnapshot) return;
// Restart the container with the saved snapshot.
this.ctx.container.start({ containerSnapshot, enableInternet: false });
}
}
The snapshot handle is opaque; the platform manages the underlying storage and restoration mechanics. The example stores the handle under the key "containerSnapshot" and disables outbound internet access on restore, but those options are configurable per use case.
Architectural and operational implications
Persisting a container’s filesystem changes the way you think about state. Previously, containers were effectively stateless between restarts, requiring external storage for any durable data. With snapshots, you can treat the container itself as a stateful unit, which simplifies patterns such as:
- Graceful shutdown for long‑running AI inference jobs that need to resume exactly where they left off.
- Hand‑off of a workload from one Durable Object instance to another for load‑balancing or geographic relocation.
- Testing and debugging workflows that need a reproducible filesystem snapshot after a failure.
Because the snapshot is stored in Durable Object storage, it inherits the same durability guarantees and scaling characteristics. However, you now have an additional artifact to manage, back up, and possibly prune, which should be incorporated into your operational playbooks.
Security and isolation considerations
The snapshot captures everything that exists on the container’s disk at the moment of capture, including any temporary files, logs, or secrets that may have been written there. Storing the snapshot handle in Durable Object storage means that any code with read access to that storage can retrieve and re‑materialize the filesystem. Practitioners should therefore:
- Avoid writing long‑lived credentials or sensitive data to the container’s filesystem; prefer environment variables or secret‑management APIs.
- Restrict access to the Durable Object storage key that holds the snapshot handle.
- Consider lifecycle policies that delete snapshots after they are no longer needed to reduce exposure.
Enabling enableInternet: false on restore, as shown in the example, is a simple way to limit outbound network exposure for restored containers, but the flag is optional and should be set according to your threat model.
Related CloudNinjas coverage: hands-on guides.
What This Means For Practitioners
- Integrate snapshot capture into your graceful‑shutdown or scaling logic to preserve on‑disk state without external databases.
- Update your CI/CD pipelines to include snapshot validation steps if you rely on filesystem state for correctness.
- Audit your container code for accidental persistence of secrets, and adjust secret handling to use platform‑provided mechanisms.
- Define retention policies for snapshot handles in Durable Object storage to balance recoverability against storage cost and security risk.
