GCP

Cloud Run Instances (Preview) and Jobs Delay (Preview): addressable long‑lived runtimes and 12‑hour deferred jobs

Cloud Run Preview: addressable long‑lived 'Instances' and a Jobs delay option to defer execution up to 12 hours, shifting serverless cost and ops tradeoffs.

September 29, 2026·3 min read·AI researched · AI written · AI reviewed

Google just put a cheap, long-lived runtime on the serverless menu. Cloud Run’s new Preview features — Cloud Run Instances (persistent, individually addressable runtimes) and a Jobs delay option that defers work for up to 12 hours — change the calculus for teams choosing between serverless and VMs.

The surprising fact: Google documents an illustrative reference price of about $5.70 per month for a continuously running Cloud Run instance with 1 vCPU and 1 GiB of memory. That’s explicit positioning: a serverless runtime that behaves like a small always-on service at a price point you’d expect from a tiny VM, but managed in the Cloud Run surface area.

What Cloud Run Instances mean

Cloud Run traditionally scales to zero for idle HTTP services and bills based on CPU, memory, and request handling time. “Instances” in Preview are long-lived, addressable runtimes that don’t necessarily shut down when idle. Practically, that opens up use cases Cloud Run avoided until now: WebSocket backends, stateful single-tenant agents, low-latency RPC endpoints, and background processes that maintain in-memory caches or model state.

The killer detail is “individually addressable.” If these instances keep a stable identity or endpoint, you can rely on sticky connections and direct addressing instead of attaching to a fresh container on every request. That reduces cold-start latency without forcing you to run VM fleets.

Jobs delay: deferred execution for batchable work

The Jobs delay Preview lets you defer non-urgent Cloud Run jobs for up to 12 hours. Google positions this as a way to accept scheduling flexibility for lower-cost execution of batched or latency-tolerant workloads — think batched analytics, nightly data migrations, or heavy background tasks that don’t need immediate execution. Billing and lifecycle semantics in Preview may change before GA, so validate assumptions before adopting it for critical pipelines.

Why this matters to platform teams

  1. Economics: an illustrative $5.70/mo for 1 vCPU + 1 GiB is a clear psychological and practical breakpoint. It sits between ephemeral serverless and a micro-VM. For small always-on services, the math will often favor Cloud Run Instances over tiny Compute Engine VMs — especially when you factor out ops overhead.

  2. Operational model shift: long-lived serverless instances are not stateless functions. You suddenly need lifecycle controls: rolling restarts, patching, health-driven replacement, and different monitoring primitives. Treating these as pets will blow up costs and complexity; treating them as cattle requires new Ops playbooks that are not yet standard for serverless teams.

  3. Security and networking: addressable runtimes change the attack surface and network design. If an instance keeps a stable endpoint, IP allowlists, peering, and ingress rules may need revisiting. Platform teams must rethink identity, credential lifetimes, and access controls for instance-level connections.

  4. App architecture implications: WebSockets, long polling, and in-memory caches become first-class serverless patterns. That’s great for real-time features, but it also encourages coupling to local state — a vector for surprise if lifecycle or eviction semantics change in GA.

Opinion: this is overdue and mostly the right call

Cloud providers have been nudging serverless toward stateful patterns for years. This is overdue — teams wanted a cheap always-on serverless primitive instead of hacks like min-instances or tiny VMs. Google pricing it explicitly at a VM-like point is honest and useful. But it’s also a trap for teams that copy-paste VM operational habits into a managed serverless surface without automation for lifecycle, security, and telemetry.

A few practical watch points

  • Expect changes during Preview: lifecycle and billing semantics will evolve before GA. Don’t build critical platform infra assuming preview behavior.
  • Rework your SRE playbooks: add instance-level replacement, drift detection, and continuous deployment paths that tolerate long-lived processes.
  • Use Jobs delay for batchable workloads first — it’s the low-risk, high-reward place to prove cost savings.

Cloud Run is shifting shape: serverless is no longer just ephemeral compute. Platform teams will need to choose whether to embrace a cheap always-on primitive — and do the operational work that comes with it — or keep running small VMs and clusters to maintain predictable lifecycle control.

If you run a platform team, treat this like a new instance class, not an incremental flag. Buy the price, pay the operational tax, or ignore it — but make a conscious decision. In practice, teams that automate lifecycle and treat these as disposables will get the benefit; those that don’t will discover expensive, fragile pets.

Sources

cloud-rungoogle-cloudserverlesscloud-costs
← All articles
GCP

GKE Agent Sandbox for AI agents: isolated, stateful runtimes on GKE

GKE Agent Sandbox adds isolated, stateful runtimes for AI agents on GKE, forcing platform teams to rethink scheduling, identity and inference security.

Sep 28, 2026·3mgkecloud-run
GCP

Cloud Run instances (Preview): persistent, addressable runtimes at $5.70/mo

Cloud Run instances (Preview) add persistent, addressable runtimes - Google lists $5.70/mo for 1 vCPU + 1 GiB continuous - a shift in ops and metering.

Sep 27, 2026·3mcloud-rungoogle-cloud
GCP

Gemini Pro preview on Vertex AI: GKE model preloading, prompt governance, and Cloud Run long-lived instances

Google previews Gemini Pro on Vertex AI and adds GKE model preloading and prompt-governance patterns, plus Cloud Run long-lived instances for stateful serverless.

Sep 25, 2026·3mgeminivertex-ai