OpenTelemetry's graduation is no longer an academic milestone — CNCF's Observability Day (announced Sept 24) just declared AI workload telemetry and telemetry cost the operational problems everyone must solve now. The message is blunt: instrumentation that worked for microservices won't survive LLM inference at production scale unless you design for cardinality, ingestion cost, and data hygiene up front.
CNCF framed four priorities for the observability community: AI workloads, telemetry cost, scale, and data quality. AI inference generates high-cardinality, high-frequency telemetry — prompt variants, model versions, user/session IDs, embeddings — and that multiplies both ingest volume and the downstream complexity of querying and correlating signals. If you still treat traces and metrics as unlimited, your bill and your incident mean-time-to-resolution will prove otherwise.
The practical surface area here is concrete: expect OTLP pipelines to be the new battleground. You need processors that do selective sampling, pre-aggregation, and context-aware redaction before you ever hit storage. Tail-based sampling for traces, attribute hashing or bucketing for high-cardinality labels, and protocol choices (OTLP over gRPC or OTLP/HTTP with batching and compression) are operational levers, not nice-to-haves. And yes, downstream storage matters — histogram-first metric stores, compressed trace archives, and object-store cold paths are table stakes.
Data quality is the other shoe. Graduation means semantic conventions have to be enforced. For cloud-native APM that means agreed fields for model_id, model_commit, prompt_hash, prompt_size, and consistent latency reporting (p50/p95/p99 or histogram buckets) and error semantics. Without this, queries across teams are brittle and costly. OpenTelemetry's standardization gives us the knobs; the community event is a push to actually use them rather than creating bespoke, incompatible attribute names per service.
Cilium is a CNCF project focused on eBPF-based networking and observability; the current ecosystem focus is less about control plane feature releases and more about operational plumbing for telemetry. For platform teams that means prioritizing collector processors, ingestion-cost controls, and storage adapters that optimize for high-cardinality, histogram-heavy workloads.
Opinion: CNCF is right to put telemetry cost front-and-center. Platform teams historically outsourced observability decisions to libraries and individual devs; that model breaks when one AI inference pipeline can multiply telemetry by 10–100x overnight. Treating telemetry like a product — with SLAs, ingestion budgets, and schema governance — isn't optional; it's the only way to avoid surprises when agents and models are deployed at scale.
What changes in practice? A few non-negotiables: enforce semantic conventions at the OpenTelemetry Collector, implement context-aware sampling and aggregation, shift long-term retention to compressed object stores, and build ingestion quotas with graceful degradation paths. Tooling will follow: expect more OTLP processors, collector-as-a-service features for cost control, and SQL-like query engines that optimize for histogram and bucketed metric workloads.
Observability Day is the signal that the community understands the next phase: not just more telemetry, but manageable, meaningful telemetry. If your platform still bills telemetry as "instrumentation tax" with no product thinking, you will get surprised — and expensively so. The next 12 months will be about turning conventions into enforced policies and pipelines into cost-aware first-class infrastructure. Watch which collector processors and storage adapters get traction; they will determine who can run AI at scale without bankrupting their monitoring budget.