Google Cloud just added an officially supported way to run single, addressable serverless processes for days at a time: Cloud Run instances (Preview). These are not just larger containers or a configuration flag — they’re dedicated, singleton runtimes that do not autoscale, run up to seven days continuously, and restart automatically by default.
Operationally this is a different primitive from the usual Cloud Run model. Traditional Cloud Run endpoints are ephemeral, autoscaling stateless containers. Cloud Run instances are intentionally stateful-in-practice: one always-on process you can address directly, with predictable pricing and a restart policy rather than horizontal scaling. Example pricing: Google published that a continuously running instance with 1 vCPU and 1 GiB of memory (shared vCPU) costs roughly $5.70 for 30 days. That kind of low baseline cost makes running small agent runtimes or long-lived connectors plausible without the overhead of a full VM or GKE deployment.
What you get
- Addressability and persistence: single, addressable process for WebSocket connections, agent loops, or in-memory caches.
- Long lived: up to seven days of continuous runtime before the platform restarts the instance. Restart is automatic by default.
- Predictable economics: a low-cost baseline (Google’s example: ~$5.70/30 days for 1 vCPU/1 GiB shared), removing much of the "serverless tax" for always-on tiny services.
What you lose (and must design for)
- No autoscaling: the platform will not create more instances to absorb load. If you need concurrency, you must handle it in-process or run multiple instances manually.
- Availability tradeoffs: the platform will restart the instance (up to the seven-day cap). That’s different from multi-replica HA — your availability model becomes process-level healing plus restart, not redundancy. Treat these like lightweight short-lived VMs, not HA front ends.
This is the right call from Google. Platform teams have been improvising long-lived serverless processes for years — using Cloud Run hacks, tiny GKE clusters, or VMs for things that are conceptually simple agents. Offering a first-class primitive with predictable billing and lifecycle semantics will reduce messy homebrew solutions. But teams who adopt it and expect traditional autoscale semantics will be surprised.
How this sits with other Google announcements
GKE is also moving toward agent-first AI isolation: Rapid-channel defaults are on Kubernetes 1.36, and Google is promoting the GKE Agent Sandbox add-on (based on the open-source Agent Sandbox controller) for isolated, stateful, single-replica AI-agent workloads. If your workload needs kernel isolation, fine-grained node controls, GPUs, or true stateful pod affinity, GKE Agent Sandbox is the right primitive. If you want tiny, cheap, addressable agents with minimal ops, Cloud Run instances are the obvious choice. I linked deeper coverage of both: Cloud Run instances (Preview): persistent, addressable runtimes at $5.70/mo and GKE Agent Sandbox for AI agents: isolated, stateful runtimes on GKE.
Cost governance moved forward too. Google expanded billing caps and budget actions that can alert and, if configured, restrict or pause service usage, including Cloud Run and AI services such as Gemini. That’s overdue: long-lived agent processes plus cheap price points are a recipe for silent runaway spend if you don’t guard them. But be blunt about the operational tradeoff — automatic spending limits are effectively another availability control. Teams that wire billing caps to kill traffic should treat them like circuit breakers and own the fallback behavior.
Final take: this is the beginning of a sensible split in primitives. Cloud Run instances make tiny, addressable agents cheap and manageable; GKE Agent Sandbox gives you the isolated, stateful single-replica option when you need more control. What platform teams must do now is define the failure model for each primitive and bake the right health, restart, and billing-circuit behaviors into deployment templates. If you treat a Cloud Run instance like a multi-AZ service, you’ll discover the cost of being lazy — fast.