Azure

Anyscale on Azure GA: Managed Ray control plane with workers running in customer AKS clusters

Anyscale on Azure GA: managed Ray control plane; Ray workers run in your AKS clusters—compute stays in your tenancy. You must manage GPUs, network, and storage.

October 11, 2026·3 min read·AI researched · AI written · AI reviewed

Microsoft's October Azure updates quietly shift responsibility for enterprise AI: Anyscale on Azure is GA as a managed Ray control plane whose worker processes run in customers' AKS clusters. That means the orchestration, autoscaling, and multi-node scheduling logic is supplied by Anyscale, but the actual compute, networking, and storage live in your tenancy and under your constraints.

This is the right call from a compliance perspective. Sensitive model training and inference stay inside VNet boundaries and your subscription. But it's also a practical pivot: platform teams now get a powerful managed runtime and a new operational boundary to manage — one that behaves less like a SaaS and more like another critical internal platform.

Two other Azure pieces that matter here: AKS multi-NIC / multi-networking (public preview) addresses the networking needs of high-performance and AI workloads, and improved network perimeter telemetry has been announced as generally available, giving teams better visibility into how perimeter controls affect flows. Those two are not incidental — they're the plumbing you need if you're going to run distributed Ray workloads at scale on AKS.

The new operational boundary platform teams must own

Anyscale's approach splits trust and responsibility: Anyscale operates the control plane and supplies Ray operator bundles, but execution happens on your nodes. That model avoids data egress and supports corporate governance, yet it forces teams to own:

  • Node pool sizing and GPU topology: Ray's actor and placement groups are sensitive to node shapes and fragmentation. If you don't reserve GPUs cleanly and configure eviction/taint behavior, you'll get noisy neighbor and fragmentation failures.
  • CNI and throughput: Ray RPCs and object-store traffic are chatty. Azure CNI without careful tuning, or wrong MTU/route setups, will become a bottleneck. That's why AKS multi-NIC/multi-networking (public preview) matters — you can separate management, RDMA/high-throughput, and service traffic.
  • Storage and memory pressure: Ray object stores need large ephemeral or networked stores. Platform policies for local SSDs, ephemeral volumes, and kubelet eviction thresholds suddenly matter more.

If you thought installing a Helm chart was "doing Ray", you're not ready for production. Treat Anyscale as a federated control plane that needs platform-level goals: node-pool isolation, GPU quota enforcement, taints/tolerations, and observability wired into Ray's eviction/placement events.

Networking and resiliency: signals and new surfaces

Better perimeter telemetry is a welcome counterweight. Being able to quantify how perimeter policies affect service reachability and overlay flows helps debug Ray's cross-node behaviour and policy regressions. But metrics are only useful when teams act on them — expect to tune NSGs, route tables, and CNI plugin settings as you onboard distributed training jobs.

AKS multi-NIC / multi-networking addresses an immediate, practical need: AI platforms want segregated NICs (management vs. RDMA vs. egress), and Azure is starting to give you that. If you plan SR-IOV, accelerated networking, or isolated egress paths for model telemetry, build your node pools and pod annotations around multi-NIC expectations. See our earlier piece on AKS Multi-NIC public preview for implementation gotchas.

Azure is also shipping AI-driven resiliency tooling and preview agents that surface reliability risks. Useful — but they add an agent surface that needs onboarding and least-privilege. Combine those agents with automation and remediation workflows and you have rapid loops that can act inside your tenant; coordinate IAM and audit trails before you flip them on.

Final take

Azure has drawn a clear line: they will manage the brain of distributed model orchestration, but execution remains your job. That's the correct trade for regulated enterprises, and it will accelerate adoption — provided platform teams accept the operational debt. If you treat this like a skinny SaaS, you'll get surprised by GPU fragmentation, noisy-neighbor networking, and automation runbooks that don't exist. If you treat it like a built-in horizontal platform, and bake multi-NIC, observability, and hardened node pools into your AKS offering, you'll gain the benefits without the disasters.

Expect other clouds to copy the split control-plane/tenant-execution model. The real winners will be the platform teams that plan for it now — not the teams that assume "managed" means "hands off."

Sources

aksrayazure-ainetworking
← All articles
Azure

AKS Multi-NIC public preview: multiple pod network interfaces with Azure CNI

AKS public preview: Multi-NIC lets pods attach multiple network interfaces for isolation and egress control, shifting networking operations onto platform teams.

Oct 10, 2026·3mazure-aksmulti-nic
Azure

Anyscale on Azure GA — Managed Ray on AKS clusters

Anyscale on Azure is GA: run managed Ray on AKS in your subscription. Treat Ray and agent runtimes as security and cost objects for platform teams.

Oct 8, 2026·3maksray
Azure

AKS GitOps: Argo CD extension GA; AKS AI isolation, agent workloads, Copilot resiliency agent

Azure released the GitOps with Argo CD AKS extension GA and announced AKS AI isolation, agent workloads, a Copilot resiliency agent preview, and AI cost controls.

Oct 7, 2026·3maksgitops