GCP

Cloud Run Instances (Preview): Dedicated singleton runtimes for long‑lived AI agents

Cloud Run now offers singleton instances that run up to seven days at $5.70/mo for 1 vCPU/1 GiB, enabling cheap in‑process agent state but adding steady costs.

October 3, 2026·3 min read·AI researched · AI written · AI reviewed

For $5.70 a month you can now run a 1 vCPU / 1 GiB container continuously and never hit scaletozero. Google Cloud's Cloud Run instances (preview) make that primitive first-class: a dedicated singleton instance that runs one container without autoscaling for up to seven days of continuous runtime, aimed squarely at longlived, stateful AI agents.

This is significant because it's the clearest signal yet that major clouds expect teams to keep agent state in-process. A warm, addressable runtime that persists for hours (and can be renewed) removes a huge amount of engineering complexity — no external memory bank for ephemeral context, fewer serialization edge cases, and drastically lower cold-start pain for chained reasoning.

But the cheap headline masks two real operational shifts.

First, cost and metering are becoming more explicit for persistent agent workloads. Cloud Run dedicated singleton instances are billed as continuous instance runtime rather than per-request invocations, so platform and FinOps teams will see steady, recurring line items for long-lived instances and their memory use instead of only ephemeral compute bursts. That is the right call — clouds should stop pretending long-running agent affinity is free — but teams that ignore these lines will be surprised when agent fleets scale beyond a handful of instances.

Second, you now have two distinct primitives for stateful agents on Google Cloud: Cloud Run instances for a lightweight, managed singleton runtime, and a Kubernetes-native option (GKE Agent Substrate) for teams that need node-level control, GPUs, sidecars, or custom networking. Substrate is positioned as the Kubernetes-native choice and is available for evaluation/preview; production availability and channels vary.

A few concrete things to note right away:

  • Cloud Run instances avoid scaletozero and run a single instance without autoscaling for up to seven days; thats explicitly built for agents that hold memory or persistent in-process models.
  • Pricing example: continuous 1 vCPU + 1 GiB = $5.70 per 30 days; scale that by vCPU/RAM and instance count and you get predictable recurring cost instead of opaque ephemeral compute bursts.
  • Billing will surface dedicated line items for persistent instance runtime and memory usage (distinct from per-request Cloud Run invoicing), so instrumenting sessions and memory use matters for cost attribution.

If you think this is a mere UX convenience, you're missing the platform rulebook rewrite. Giving teams an inexpensive, addressable runtime nudges application architecture toward more in-process state and longer lived service lifecycles. That simplifies agent design and speeds development — and it centralizes risk. Observability, cost attribution, ROI on model warmup, and IAM/credential boundaries all move from transient, hard-to-trace spikes into steady, long-lived signals.

Practical stance: use Cloud Run instances for small fleets or edge cases where low ops and fast iteration matter. Use GKE Agent Substrate (once appropriately available) when you need node control, GPUs, or more complex networking. Instrument agent sessions and memory-bank usage now — the billing visibility makes it straightforward to see growth, and you'll want alerts before the monthly bill becomes a surprise.

This matters because “serverless but persistent” is the pattern the next wave of agent architectures will adopt. Cloud Run's $5.70 figure won't stay symbolic forever — it’s a market nudge, not charity. Platform teams who treat persistent instances as free will be the ones retrofitting quotas, telemetry, and cost guardrails six months from now. Build those guardrails first; the tools to measure and control agent costs are finally arriving, and you'll need them.

Sources

cloud-rungke-agent-substrateagent-platformai-agents
← All articles
GCP

GKE Agent Substrate: high-density AI-agent sandboxes with sub-500ms resume

GKE Agent Substrate and Cloud Run Instances (Preview) outline a Google Cloud agent-first stack: dense, resumable AI-agent sandboxes plus persistent singletons.

Oct 2, 2026·3mgke-agent-substratecloud-run-instances
GCP

Cloud Run instances (Preview): dedicated singleton runtimes that run up to seven days

Cloud Run instances (Preview) give dedicated, addressable runtimes that run up to seven days with automatic restarts—good for tiny agents, alters HA tradeoffs.

Sep 30, 2026·3mcloud-rungke
GCP

Cloud Run Instances (Preview) and Jobs Delay (Preview): addressable long‑lived runtimes and 12‑hour deferred jobs

Cloud Run Preview: addressable long‑lived 'Instances' and a Jobs delay option to defer execution up to 12 hours, shifting serverless cost and ops tradeoffs.

Sep 29, 2026·3mcloud-rungoogle-cloud