Live
Self‑Managing Context in LLMs Reduces Compute Overhead and Improves ThroughputAI‑Generated OSS Vulnerability Scans Overwhelm Human Review – Implications for Security OpsBootstrapping Claude Code with Dependency Records Eliminates Initial Memory RequirementsEnterprise Copilot model control and MCP startup options in JetBrains pluginMicrosoft‑Decision‑1 Arrives on Foundry: What Engineers Need to KnowIntegrating Production Feedback into the AI Agent Lifecycle: Practical Architecture and Ops GuidanceOpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model BackendSelf‑Managing Context in LLMs Reduces Compute Overhead and Improves ThroughputAI‑Generated OSS Vulnerability Scans Overwhelm Human Review – Implications for Security OpsBootstrapping Claude Code with Dependency Records Eliminates Initial Memory RequirementsEnterprise Copilot model control and MCP startup options in JetBrains pluginMicrosoft‑Decision‑1 Arrives on Foundry: What Engineers Need to KnowIntegrating Production Feedback into the AI Agent Lifecycle: Practical Architecture and Ops GuidanceOpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model Backend
Google Cloud

Persistent AI Agents on Cloud Run Instances: Architecture, Deployment, and Ops

AI SummaryPowered by AI

Google Cloud added Cloud Run instances, a singleton, long‑lived compute option for Cloud Run. This gives engineers a low‑cost, serverless‑style way to run personal AI agents without managing VMs.

Google Cloud announced Cloud Run instances, a new mode for Cloud Run that runs a single, long‑lived container instead of scaling to zero. The change lets engineers host personal AI agents without the overhead of a full VM while keeping costs predictable.

What Cloud Run Instances Provide

Each instance runs exactly one container and does not autoscale. The runtime can stay active for up to seven days, after which the platform automatically restarts it using a default restart policy. A stable HTTPS URL is assigned to the instance and persists across updates and restarts. Operators can manually stop an instance when idle and start it again on demand.

Deploying a Personal AI Agent

Deploying an open‑source agent such as OpenClaw follows a single gcloud command. The command pulls the container image, exposes the required port, makes the endpoint publicly reachable, mounts a Cloud Storage bucket for persistent configuration, and injects environment variables for secrets.

gcloud beta run instances create openclaw-instance \
  --image ghcr.io/openclaw/openclaw:latest \
  --port 18789 \
  --public \
  --add-volume mount-path=/home/node/.openclaw,type=cloud-storage,mount-options=\"uid=1000;gid=1000;file-mode=0700;dir-mode=0700\",bucket=${BUCKET} \
  --set-env-vars \"OPENCLAW_GATEWAY_PASSWORD=${PASSWORD},GEMINI_API_KEY=${GEMINI_API_KEY}\"

After the instance is created, the agent remains reachable via its HTTPS URL and can be contacted through messaging platforms such as Telegram or WhatsApp. The instance continues to run until the operator stops it or the seven‑day limit triggers a restart.

Operational and Security Implications

Because the instance is always on, the cost model is based on a shared vCPU with burst capacity. The source cites a price of $5.70 for a 1 vCPU / 1 GiB configuration running continuously for 30 days, which is substantially lower than a dedicated VM.

Running a public HTTPS endpoint means the surface area for inbound traffic is larger than a private VM behind a firewall. Secrets are supplied via environment variables, so standard secret‑management hygiene (rotation, limited exposure) remains important. The mounted Cloud Storage bucket provides persistent state, but the bucket’s IAM configuration must be reviewed to avoid unintended access.

Future SSH access for instances is announced but not yet available; practitioners should monitor the rollout if they need shell‑level debugging.

Related CloudNinjas coverage: Google Cloud.

What This Means For Practitioners

Cloud Run instances give you a serverless‑style, low‑cost host for stateful AI agents that need to stay alive continuously. Evaluate whether the singleton model, seven‑day runtime limit, and public endpoint fit your security posture and cost targets. If they do, adopt the instance model for personal agents, and plan for secret handling and bucket permissions accordingly.

Originally published atGoogle Cloud Blog