Skip to main content
The Kubernetes integration requires CXDOT Collector 1.4.0 or greater. Kubernetes is an open source system for automating the deployment, scaling, and management of containerized applications. Use the Kubernetes CXDOT Collector integration with the CXDOT Collector to collect workload, node, cluster state, and control-plane metrics and Kubernetes logs. This CXDOT Collector integration supports Kubernetes 1.31 or greater. It supports managed and self-managed distributions that implement the upstream Kubernetes APIs. Provider restrictions on control-plane endpoints and audit logs can limit the telemetry available from managed clusters.

Supported telemetry types

This CXDOT Collector integration supports these telemetry types: Kubernetes Events are collected through the logs signal.

Prerequisites

This CXDOT Collector integration has the following prerequisites:
  • Deploy the CXDOT Collector to your cluster with the CXDOT Collector Helm chart. For more information, see Install the Collector on Kubernetes.
  • Grant the collector service accounts the Kubernetes API permissions required by the enabled capabilities. The Helm chart creates the required roles and bindings, so the identity that installs the chart must be permitted to create them.
  • Make the Kubernetes API and the endpoints for each enabled Kubernetes component reachable from the collector tier that collects them. Each node collector connects directly to its node’s kubelet on port 10250 for kubelet and cAdvisor metrics.
  • To collect scheduler and controller manager metrics from discovered pods, schedule node collectors on control-plane nodes. These nodes commonly have taints that the node collectors must tolerate. API server metrics don’t need a node collector on a control-plane node.
  • Leave clusterCollector.enabled enabled in the CXDOT Collector Helm chart when you want API server metrics, cluster resource state metrics, Kubernetes Events as logs, or leader-elected static targets for control-plane components. The chart enables the cluster collector by default; cluster-scoped capabilities don’t run when it’s disabled.
  • To collect k8s.state.hpa.status.metric.current, make the metrics API that each HorizontalPodAutoscaler targets available in your cluster: metrics.k8s.io, typically served by metrics-server, for CPU and memory targets. The autoscaler controller records a current metric reading only when that API answers, and the CXDOT Collector reports the metric only where that reading exists.
  • Audit log collection is supported only when the API server writes its audit log to /var/log/kubernetes/audit.log. Managed Kubernetes providers typically don’t expose this file, and audit log webhook backends aren’t supported.

Configure

This CXDOT Collector integration is enabled by default when you install the CXDOT Collector with the CXDOT Collector Helm chart. To change what it collects, follow these steps: You can enable or disable collection, identify the cluster, and choose how the CXDOT Collector authenticates to the Kubernetes API. The following table summarizes each component: Custom kubelet and cluster metric selections replace their default metric selections. When you customize either set, restate every default metric that you want to retain. Enabling Secret metadata collection grants the cluster Collector permission to list and watch Secret objects across the cluster. For the complete list of fields, accepted values, and defaults, consult the configuration reference on this page.
  1. Optional: Enable metrics from additional Kubernetes components. The CXDOT Collector discovers supported component pods by their standard labels and ports. It discovers the API server from the endpoints of the kubernetes Service in the default namespace, which managed and self-managed clusters both provide, and one leader-elected cluster Collector replica scrapes them. For more information, see autodiscovery. For example, add the following to the values.yaml for your CXDOT Collector Helm chart:
    Metrics from discovered pods arrive with that pod’s Kubernetes metadata attached. For more information, see enrichment.
  2. Optional: Add tolerations for tainted nodes whose node-local components you want to monitor. A node collector must run on a control-plane node to collect its scheduler, controller manager, and audit log. For example, tolerate the standard control-plane taint:
  3. Optional: Configure static targets for Kubernetes components that the CXDOT Collector can’t discover, such as a scheduler that runs outside your cluster’s pods. Managed Kubernetes providers might not expose scheduler or controller manager metrics endpoints; this CXDOT Collector integration can’t collect those metrics unless the provider exposes a reachable endpoint. For example, add the following to the values.yaml for your CXDOT Collector Helm chart:
    Static targets turn off discovery for that component: the CXDOT Collector scrapes exactly the endpoints you list. One leader-elected cluster Collector replica collects static API server, scheduler, controller manager, and CoreDNS targets per cluster. Static kube-proxy targets run on the node tier. For more information, see architecture.

API server attributes

API server metrics from a managed control plane identify the endpoint, not a pod. Managed providers such as Amazon EKS, Google Kubernetes Engine, and Azure Kubernetes Service run the API server outside your cluster, so your cluster has no pod or node to attribute the metrics to. On a self-managed cluster, the CXDOT Collector matches each API server endpoint to its pod by address and port, and attaches that pod’s metadata. The following table lists the attributes that API server metrics carry on each type of cluster: To compare API server metrics across managed clusters, group them by server.address rather than by a node or pod attribute.

Validate

To validate this CXDOT Collector integration, follow these steps:
  1. In the Live Telemetry Analyzer, add the following filters:
    • __name__=cxdot.integration.target.health
    • cxdot.integration.name=kubernetes
    Confirm that time series appear for the targets you enabled.
  2. In Metrics Explorer, run the following query:
    Confirm that each reachable target reports 1.
  3. In Metrics Explorer, run the following query:
    Confirm that the query returns the expected time series for each node.
  4. Optional: If you enabled Kubernetes Events, in Logs Explorer, search for "cxdot.integration.name":=kubernetes "k8s.event.reason":*. Confirm that events from your cluster appear.
For more information about verifying ingested metrics, see Verify metrics. For more information about diagnosing a failing integration, see Troubleshooting.

Configuration reference

Configure one Kubernetes integration instance with the following settings. In Helm values, place these settings under config.integrations.kubernetes. In a Collector configuration file, place them under cxdot.integrations.kubernetes.

Optional settings

  • enabled Type: boolean. Optional. Default: true. Whether to enable this configuration block. If true, the Collector runs the integration or capability. If false, the Collector doesn’t run it.
  • cluster_name Type: string. Optional. Cluster identifier injected as the k8s.cluster.name resource attribute on every signal. NB: deployments that set a global cluster identifier apply it last, overriding this one — the CXDOT Collector Helm chart does, and rejects a cluster_name here that disagrees with its own value rather than letting the two diverge silently.
  • auth Type: string. Optional. Default: serviceAccount. How the integration authenticates to the Kubernetes API. serviceAccount uses the in-cluster pod token; kubeConfig reads $KUBECONFIG / ~/.kube/config (kubectl-style); none is unauthenticated. Allowed values: serviceAccount, kubeConfig, none.
  • kubelet Type: object. Optional. Settings for collecting node, pod, container, and volume metrics from each node’s kubelet stats endpoint.
  • kubelet.enabled Type: boolean. Optional. Default: true. Whether to enable this configuration block. If true, the Collector runs the integration or capability. If false, the Collector doesn’t run it.
  • kubelet.collection_interval Type: duration. Optional. Default: 20s. How often the Collector collects metrics from each node’s kubelet.
  • kubelet.timeout Type: duration. Optional. Default: 20s. Maximum time the Collector waits for a kubelet stats response during one collection.
  • kubelet.metrics Type: object. Optional. Default: {"k8s.container.cpu_limit_utilization":{"enabled":true},"k8s.container.memory_limit_utilization":{"enabled":true},"k8s.node.system_container.cpu.usage":{"enabled":true},"k8s.node.system_container.memory.usage":{"enabled":true},"k8s.pod.cpu_limit_utilization":{"enabled":true},"k8s.pod.cpu_request_utilization":{"enabled":true},"k8s.pod.memory_limit_utilization":{"enabled":true},"k8s.pod.memory_request_utilization":{"enabled":true},"k8s.pod.volume.usage":{"enabled":true}}. Pass-through to the kubeletstats receiver’s metrics: block — per-metric enabled: toggles for metrics that are off by default upstream, keyed by metric name. Validated against the receiver’s own config schema. Setting this replaces the whole default map below rather than merging into it, so restate any default you still want.
  • kubelet_metrics Type: object. Optional. Settings for collecting the kubelet’s own Prometheus metrics from each node.
  • kubelet_metrics.enabled Type: boolean. Optional. Default: true. Whether to enable this configuration block. If true, the Collector runs the integration or capability. If false, the Collector doesn’t run it.
  • kubelet_metrics.collection_interval Type: duration. Optional. Default: 15s. How often the Collector scrapes each node’s kubelet metrics endpoint.
  • kubelet_metrics.timeout Type: duration. Optional. Default: 10s. Maximum time the Collector waits for one scrape to complete.
  • cadvisor Type: object. Optional. Settings for collecting container resource metrics from each node’s cAdvisor endpoint.
  • cadvisor.enabled Type: boolean. Optional. Default: true. Whether to enable this configuration block. If true, the Collector runs the integration or capability. If false, the Collector doesn’t run it.
  • cadvisor.collection_interval Type: duration. Optional. Default: 15s. How often the Collector scrapes each node’s cAdvisor endpoint.
  • cadvisor.timeout Type: duration. Optional. Default: 10s. Maximum time the Collector waits for one scrape to complete.
  • cadvisor.honor_timestamps Type: boolean. Optional. Default: true. Whether to keep the timestamps cAdvisor attaches to its samples, which date from its last internal housekeeping pass and can trail the scrape by tens of seconds. Set false to stamp samples at scrape time instead — for ingest paths whose sample-recency window is tight enough to reject samples that are already old when scraped.
  • container_logs Type: object. Optional. Settings for collecting container logs from each node.
  • container_logs.enabled Type: boolean. Optional. Default: true. Whether to enable this configuration block. If true, the Collector runs the integration or capability. If false, the Collector doesn’t run it.
  • audit_logs Type: object. Optional. Settings for collecting Kubernetes API server audit logs.
  • audit_logs.enabled Type: boolean. Optional. Default: true. Whether to enable this configuration block. If true, the Collector runs the integration or capability. If false, the Collector doesn’t run it.
  • cluster Type: object. Optional. Settings for collecting cluster-level object state — deployments, pods, nodes, and other resources — from the Kubernetes API.
  • cluster.enabled Type: boolean. Optional. Default: true. Whether to enable this configuration block. If true, the Collector runs the integration or capability. If false, the Collector doesn’t run it.
  • cluster.namespaces Type: array of string. Optional. Default: []. Restrict cluster resource collection to these namespaces. Omitted or empty collects every namespace.
  • cluster.collection_interval Type: duration. Optional. Default: 10s. How often the Collector reads cluster object state.
  • cluster.timeout Type: duration. Optional. Default: 10s. Maximum time the Collector waits for one read to complete.
  • cluster.allow_secrets_read Type: boolean. Optional. Default: false. Allow the cluster collector to read Secrets cluster-wide (list/watch), which kube-state-metrics requires to emit the Secret metadata metrics (kubernetes_state.secret.count and .type). Off by default: because Kubernetes LIST/WATCH return full Secret objects including data, the grant exposes every Secret’s contents to the collector’s ServiceAccount. Leave it off unless you need these metrics.
  • cluster.metrics Type: object. Optional. Default: {"k8s.container.status.reason":{"enabled":true},"k8s.pod.status_reason":{"enabled":true},"k8s.service.endpoint.count":{"enabled":true}}. Pass-through to the k8s_cluster receiver’s metrics: block — per-metric enabled: toggles for metrics that are off by default upstream, keyed by metric name. Validated against the receiver’s own config schema. Setting this replaces the whole default map below rather than merging into it, so restate any default you still want.
  • events Type: object. Optional. Settings for collecting Kubernetes events as log records.
  • events.enabled Type: boolean. Optional. Default: true. Whether to enable this configuration block. If true, the Collector runs the integration or capability. If false, the Collector doesn’t run it.
  • apiserver Type: object. Optional. Settings for scraping this Kubernetes control-plane component’s metrics endpoint.
  • apiserver.enabled Type: boolean. Optional. Default: true. Whether to enable this configuration block. If true, the Collector runs the integration or capability. If false, the Collector doesn’t run it.
  • apiserver.collection_interval Type: duration. Optional. Default: 30s. How often the Collector scrapes the component’s metrics endpoint.
  • apiserver.timeout Type: duration. Optional. Default: 10s. Maximum time the Collector waits for one scrape to complete.
  • apiserver.static_targets Type: array of object. Optional. Default: []. Endpoints to scrape instead of discovering this component. A non-empty list turns off discovery for the component, and the Collector scrapes exactly the listed endpoints. Omit the list or leave it empty to discover the component’s endpoints automatically.
  • apiserver.static_targets[].endpoint Type: string. Required. host:port of the scrape target, for example, 127.0.0.1:6443.
  • scheduler Type: object. Optional. Settings for scraping this Kubernetes control-plane component’s metrics endpoint.
  • scheduler.collection_interval Type: duration. Optional. Default: 30s. How often the Collector scrapes the component’s metrics endpoint.
  • scheduler.enabled Type: boolean. Optional. Default: true. Whether to enable this configuration block. If true, the Collector runs the integration or capability. If false, the Collector doesn’t run it.
  • scheduler.static_targets Type: array of object. Optional. Default: []. Endpoints to scrape instead of discovering this component. A non-empty list turns off discovery for the component, and the Collector scrapes exactly the listed endpoints. Omit the list or leave it empty to discover the component’s endpoints automatically.
  • scheduler.static_targets[].endpoint Type: string. Required. host:port of the scrape target, for example, 127.0.0.1:6443.
  • scheduler.timeout Type: duration. Optional. Default: 10s. Maximum time the Collector waits for one scrape to complete.
  • controller_manager Type: object. Optional. Settings for scraping this Kubernetes control-plane component’s metrics endpoint.
  • controller_manager.collection_interval Type: duration. Optional. Default: 30s. How often the Collector scrapes the component’s metrics endpoint.
  • controller_manager.enabled Type: boolean. Optional. Default: true. Whether to enable this configuration block. If true, the Collector runs the integration or capability. If false, the Collector doesn’t run it.
  • controller_manager.static_targets Type: array of object. Optional. Default: []. Endpoints to scrape instead of discovering this component. A non-empty list turns off discovery for the component, and the Collector scrapes exactly the listed endpoints. Omit the list or leave it empty to discover the component’s endpoints automatically.
  • controller_manager.static_targets[].endpoint Type: string. Required. host:port of the scrape target, for example, 127.0.0.1:6443.
  • controller_manager.timeout Type: duration. Optional. Default: 10s. Maximum time the Collector waits for one scrape to complete.
  • kube_proxy Type: object. Optional. Settings for scraping the kube-proxy metrics endpoint on each node.
  • kube_proxy.collection_interval Type: duration. Optional. Default: 30s. How often the Collector scrapes the component’s metrics endpoint.
  • kube_proxy.enabled Type: boolean. Optional. Default: true. Whether to enable this configuration block. If true, the Collector runs the integration or capability. If false, the Collector doesn’t run it.
  • kube_proxy.static_targets Type: array of object. Optional. Default: []. Endpoints to scrape instead of discovering this component. A non-empty list turns off discovery for the component, and the Collector scrapes exactly the listed endpoints. Omit the list or leave it empty to discover the component’s endpoints automatically.
  • kube_proxy.static_targets[].endpoint Type: string. Required. host:port of the scrape target, for example, 127.0.0.1:6443.
  • kube_proxy.timeout Type: duration. Optional. Default: 10s. Maximum time the Collector waits for one scrape to complete.
  • coredns Type: object. Optional. Settings for scraping CoreDNS metrics from CoreDNS pods on each node, plus optional cluster-tier static targets.
  • coredns.collection_interval Type: duration. Optional. Default: 30s. How often the Collector scrapes the component’s metrics endpoint.
  • coredns.enabled Type: boolean. Optional. Default: true. Whether to enable this configuration block. If true, the Collector runs the integration or capability. If false, the Collector doesn’t run it.
  • coredns.static_targets Type: array of object. Optional. Default: []. Endpoints to scrape instead of discovering this component. A non-empty list turns off discovery for the component, and the Collector scrapes exactly the listed endpoints. Omit the list or leave it empty to discover the component’s endpoints automatically.
  • coredns.static_targets[].endpoint Type: string. Required. host:port of the scrape target, for example, 127.0.0.1:6443.
  • coredns.timeout Type: duration. Optional. Default: 10s. Maximum time the Collector waits for one scrape to complete.