GCP

Cloud Run instances (Preview): dedicated singleton runtimes that run up to seven days

Cloud Run instances (Preview) give dedicated, addressable runtimes that run up to seven days with automatic restarts—good for tiny agents, alters HA tradeoffs.

September 30, 2026·3 min read·AI researched · AI written · AI reviewed

Google Cloud just added an officially supported way to run single, addressable serverless processes for days at a time: Cloud Run instances (Preview). These are not just larger containers or a configuration flag — they’re dedicated, singleton runtimes that do not autoscale, run up to seven days continuously, and restart automatically by default.

Operationally this is a different primitive from the usual Cloud Run model. Traditional Cloud Run endpoints are ephemeral, autoscaling stateless containers. Cloud Run instances are intentionally stateful-in-practice: one always-on process you can address directly, with predictable pricing and a restart policy rather than horizontal scaling. Example pricing: Google published that a continuously running instance with 1 vCPU and 1 GiB of memory (shared vCPU) costs roughly $5.70 for 30 days. That kind of low baseline cost makes running small agent runtimes or long-lived connectors plausible without the overhead of a full VM or GKE deployment.

What you get

  • Addressability and persistence: single, addressable process for WebSocket connections, agent loops, or in-memory caches.
  • Long lived: up to seven days of continuous runtime before the platform restarts the instance. Restart is automatic by default.
  • Predictable economics: a low-cost baseline (Google’s example: ~$5.70/30 days for 1 vCPU/1 GiB shared), removing much of the "serverless tax" for always-on tiny services.

What you lose (and must design for)

  • No autoscaling: the platform will not create more instances to absorb load. If you need concurrency, you must handle it in-process or run multiple instances manually.
  • Availability tradeoffs: the platform will restart the instance (up to the seven-day cap). That’s different from multi-replica HA — your availability model becomes process-level healing plus restart, not redundancy. Treat these like lightweight short-lived VMs, not HA front ends.

This is the right call from Google. Platform teams have been improvising long-lived serverless processes for years — using Cloud Run hacks, tiny GKE clusters, or VMs for things that are conceptually simple agents. Offering a first-class primitive with predictable billing and lifecycle semantics will reduce messy homebrew solutions. But teams who adopt it and expect traditional autoscale semantics will be surprised.

How this sits with other Google announcements

GKE is also moving toward agent-first AI isolation: Rapid-channel defaults are on Kubernetes 1.36, and Google is promoting the GKE Agent Sandbox add-on (based on the open-source Agent Sandbox controller) for isolated, stateful, single-replica AI-agent workloads. If your workload needs kernel isolation, fine-grained node controls, GPUs, or true stateful pod affinity, GKE Agent Sandbox is the right primitive. If you want tiny, cheap, addressable agents with minimal ops, Cloud Run instances are the obvious choice. I linked deeper coverage of both: Cloud Run instances (Preview): persistent, addressable runtimes at $5.70/mo and GKE Agent Sandbox for AI agents: isolated, stateful runtimes on GKE.

Cost governance moved forward too. Google expanded billing caps and budget actions that can alert and, if configured, restrict or pause service usage, including Cloud Run and AI services such as Gemini. That’s overdue: long-lived agent processes plus cheap price points are a recipe for silent runaway spend if you don’t guard them. But be blunt about the operational tradeoff — automatic spending limits are effectively another availability control. Teams that wire billing caps to kill traffic should treat them like circuit breakers and own the fallback behavior.

Final take: this is the beginning of a sensible split in primitives. Cloud Run instances make tiny, addressable agents cheap and manageable; GKE Agent Sandbox gives you the isolated, stateful single-replica option when you need more control. What platform teams must do now is define the failure model for each primitive and bake the right health, restart, and billing-circuit behaviors into deployment templates. If you treat a Cloud Run instance like a multi-AZ service, you’ll discover the cost of being lazy — fast.

Sources

cloud-rungkegke-agent-sandboxcloud-cost-controls
← All articles
GCP

Cloud Run Instances (Preview) and Jobs Delay (Preview): addressable long‑lived runtimes and 12‑hour deferred jobs

Cloud Run Preview: addressable long‑lived 'Instances' and a Jobs delay option to defer execution up to 12 hours, shifting serverless cost and ops tradeoffs.

Sep 29, 2026·3mcloud-rungoogle-cloud
GCP

GKE Agent Sandbox for AI agents: isolated, stateful runtimes on GKE

GKE Agent Sandbox adds isolated, stateful runtimes for AI agents on GKE, forcing platform teams to rethink scheduling, identity and inference security.

Sep 28, 2026·3mgkecloud-run
GCP

Cloud Run instances (Preview): persistent, addressable runtimes at $5.70/mo

Cloud Run instances (Preview) add persistent, addressable runtimes - Google lists $5.70/mo for 1 vCPU + 1 GiB continuous - a shift in ops and metering.

Sep 27, 2026·3mcloud-rungoogle-cloud