When a Kubernetes cluster scales past two hundred nodes, Prometheus instances frequently hit memory cliffs, experiencing sudden OOMKilled crashes during daily TSDB compactions or scrape spikes. Production monitoring collapses precisely when engineers need it most: during cascade failures, control plane latency surges, or mass pod evictions. Setting up monitoring by stringing together loose YAML manifests or executing arbitrary Helm charts with default settings guarantees unbudgeted compute costs, silent metric dropouts, and blind spots across critical cluster components.
Achieving resilient, enterprise-grade Prometheus monitoring Kubernetes clusters demands a firm grasp of how the Prometheus Operator coordinates declarative Custom Resource Definitions (CRDs) alongside dynamic API discovery mechanisms. It requires distinguishing container-level cAdvisor telemetry from kube-state-metrics object metadata, provisioning strict PersistentVolume storage bounds, and controlling explosive label cardinality before it corrupts the time-series database engine.
This engineering reference provides complete, battle-tested patterns for running Prometheus on Kubernetes in 2026. From hardening kube-prometheus-stack configurations and writing deterministic ServiceMonitor manifests to analyzing control plane metrics and deploying horizontally scalable storage layers, this guide establishes a stable, high-performance observability infrastructure.
Kubernetes Observability Architecture Under the Hood
Running Prometheus monitoring Kubernetes infrastructure at scale requires understanding the intersection of pull-based telemetry and Kubernetes API primitives. Unlike traditional static infrastructure where scrape targets possess deterministic IP addresses and hostnames, Kubernetes environments feature dynamic, ephemeral pods whose lifecycles fluctuate continuously. Prometheus bridges this gap by decoupling metric generation from metric collection through dynamic API-driven service discovery.
Prometheus operates on an inverted data collection model: instead of workloads pushing telemetry across network perimeters to an ingestion gateway, the Prometheus server initiates outbound HTTP/S scrape calls to target endpoints at configured intervals. When pairing Prometheus and Kubernetes, this mechanism relies directly on continuous API watch loops managed by discovery engines.
+-----------------------------------------------------------------------------------+
| Kubernetes Control Plane |
| +--------------------+ +--------------------+ +-------------------------+ |
| | kube-apiserver | | kube-scheduler | | kube-controller-manager | |
| | (:6443/metrics) | | (:10259/metrics) | | (:10257/metrics) | |
| +---------+----------+ +---------+----------+ +------------+------------+ |
+------------|-------------------------|----------------------------|---------------+
| API Watch Loops | |
v | |
+-------------------------------------------------------------------|---------------+
| Prometheus Server Core | | |
| +---------------------------------+ | | |
| | Dynamic Service Discovery |<+----------------------------+ |
| | (Endpoints, Services, Pods, CRDs) |
| +----------------+----------------+ |
| | Resolved Target Map |
| v |
| +----------------+----------------+ Scrape Protocol (HTTP GET /metrics) |
| | Scrape Engine & Relabel Rules |====================+ |
| +----------------+----------------+ | |
| | Appended Samples | |
| v v |
| +----------------+----------------+ +--------------------------+ |
| | Local TSDB (Head Chunk Mem/WAL) | | Workload Pods & Daemons | |
| +----------------+----------------+ | cAdvisor (:10250) | |
| | Compaction | Node Exporter (:9100) | |
| v | App Services (:8080) | |
| +----------------+----------------+ +--------------------------+ |
| | Persistent Block Storage (PVC) | |
| +---------------------------------+ |
+-----------------------------------------------------------------------------------+
The internal lifecycle of an individual telemetry point involves multiple asynchronous stages:
- Discovery Phase: Prometheus queries the
kube-apiserverfor registered resources, inspecting Endpoints, Services, Pods, and Ingresses. When using the Operator pattern, it interrogates CRDs such asServiceMonitorandPodMonitor. - Relabeling Phase: Prometheus executes
relabel_configs. Labels prefaced with__meta_kubernetes_*are examined, mutated, or filtered out. This phase is critical: it enables or disables metric collection before any network I/O takes place. - Scrape Execution: At each
scrape_interval, the scraping loop makes an HTTPGETrequest against the resolved target endpoint (typically/metrics). Timeouts are strictly enforced viascrape_timeout. - Metric Relabeling Phase: Prometheus evaluates
metric_relabel_configson the returned payload. Unwanted metrics, internal instrumentation garbage, and excessively granular time-series are pruned prior to memory allocation. - TSDB Append: Ingested samples land in memory within the active Head chunk and are concurrently appended to the on-disk Write-Ahead Log (WAL) to guarantee zero data loss during pod evictions or node terminations.
Architecture Rule: Keep
scrape_timeoutstrictly lower than yourscrape_interval. Setting identical durations causes scrape task overlap during network latency spikes, resulting in connection pool exhaustion, memory leakage, and artificial metric gaps.
Below is a production-grade native configuration demonstrating how Prometheus interacts directly with the Kubernetes API to discover endpoints securely while injecting essential cluster metadata:
# Prometheus native scrape configuration for Kubernetes Endpoints
scrape_configs:
- job_name: 'kubernetes-endpoints'
scrape_interval: 15s
scrape_timeout: 10s
honor_timestamps: true
kubernetes_sd_configs:
- role: endpoints
namespaces:
names:
- production
- staging
scheme: http
tls_config:
insecure_skip_verify: false
relabel_configs:
# Select only services explicitly marked for Prometheus scraping
- source_labels: [__meta_kubernetes_service_annotation_prometheus_io_scrape]
action: keep
regex: true
# Extract custom metric paths if specified
- source_labels: [__meta_kubernetes_service_annotation_prometheus_io_path]
action: replace
target_label: __metrics_path__
regex: (.+)
# Dynamically map custom scrape ports
- source_labels: [__address__, __meta_kubernetes_service_annotation_prometheus_io_port]
action: replace
target_label: __address__
regex: ([^:]+)(?:\d+)?(\d+)
replacement: $1:$2
# Normalize metadata into clean, queryable metric labels
- source_labels: [__meta_kubernetes_namespace]
action: replace
target_label: kubernetes_namespace
- source_labels: [__meta_kubernetes_service_name]
action: replace
target_label: kubernetes_service
- source_labels: [__meta_kubernetes_pod_name]
action: replace
target_label: kubernetes_pod
Architectural Taxonomy: cAdvisor, Kube-State-Metrics, and Node Exporter
A common operational flaw in prometheus k8s architectures is conflating host metrics, container execution telemetry, and control-plane object statuses. A high-signal monitoring platform separates instrumentation into three distinct layers, each served by a dedicated metric pipeline.
These three instrumentation pillars form an integrated observability fabric:
- cAdvisor (Container Advisor): Embedded directly within the
kubeletbinary running on every node. It intercepts runtime events and reads the host’s Linux cgroups (both v1 and v2) to report raw resource saturation such as CPU ticks, memory page allocations, swap usage, and container network interface traffic. cAdvisor measures physical resource consumption at the container boundary. - kube-state-metrics (KSM): A cluster-wide deployment that listens to the
kube-apiserverevent stream. It does not measure CPU usage or latency. Instead, it translates Kubernetes objects (Deployments, Pods, StatefulSets, PersistentVolumeClaims, Node statuses) into time-series metrics. KSM answers structural questions: Is a deployment failing its replica guarantees? Did a container terminate due to an OOMKilled event? Are node conditions reportingDiskPressure? - Node Exporter (prometheus-node-exporter): Deployed as a
DaemonSetacross every physical or virtual node in the cluster. It ignores container constructs entirely, scraping low-level Linux kernel metrics via procfs and sysfs subsystems. It surfaces disk I/O wait times, thermal throttling, filesystem saturation, network socket exhaustion, and system-level context switching.
| Metric Pipeline Dimension | cAdvisor | kube-state-metrics | Node Exporter |
|---|---|---|---|
| Data Source | Linux cgroups & Kubelet runtime API | Kubernetes API Server etcd state | Linux /proc, /sys, systemd |
| Deployment Model | Daemon (embedded inside Kubelet) | Clustered deployment (Active/Standby) | DaemonSet (1 per node) |
| Scrape Port / Endpoint | TCP 10250 (/metrics/cadvisor) |
TCP 8080 (/metrics) |
TCP 9100 (/metrics) |
| Key Metric Examples | container_cpu_usage_seconds_totalcontainer_memory_working_set_bytes |
kube_pod_status_phasekube_deployment_status_replicas_unavailable |
node_cpu_seconds_totalnode_filesystem_avail_bytes |
| Primary Failure Mode | Kubelet API throttles under high load | Excessive API watch pressure on etcd | High host CPU usage from disk sweeps |
| Cardinality Footprint | Very High (ephemeral container IDs) | Moderate (proportional to object count) | Low (fixed to host count and devices) |
To ensure high cluster fidelity without causing collection failures or performance penalties, configure the following architectural settings across these exporters:
- Enable the
--metric-labels-allowlistflag insidekube-state-metricsto avoid converting arbitrary pod labels into high-cardinality Prometheus labels. - Disable high-frequency, non-critical cAdvisor metrics (such as UDP metrics, per-process tasks, and disk usage counters) directly within the Kubelet configuration using the
--disabled-metricsflag. - Run Node Exporter with host network and IPC access enabled, mounting host paths (
/procand/sys) as read-only volumes to guarantee isolation from container runtime file locks. - Use separate scrape pools for cAdvisor with lower collection frequencies (30s to 60s) compared to custom application workloads, significantly dampening TSDB storage growth.
How to Deploy Prometheus on Kubernetes Using kube-prometheus-stack
Attempting to deploy Prometheus on Kubernetes using hand-crafted, raw YAML manifests introduces configuration drift, error-prone RBAC handling, and breaks dynamic CRD validation on modern Kubernetes releases (v1.25+ through v1.32+). The industry standard approach is the kube-prometheus-stack, a Helm-driven collection centered around the CoreOS Prometheus Operator.
The prometheus stack bundles the Prometheus Operator, an orchestrated Prometheus server, Alertmanager, Grafana, Node Exporter, and kube-state-metrics into a coherent lifecycle framework. To deploy this stack cleanly in production, avoid default chart parameters. Storage backends, memory allocations, and retention boundaries must be declared explicitly.
- Add and Update the Prometheus Community Helm Repository:
Ensure local chart indices are synchronized with upstream releases.helm repo add prometheus-community https://prometheus-community.github.io/helm-charts helm repo update - Provision a Dedicated Isolated Namespace:
Keep cluster monitoring isolated from shared application namespaces.kubectl create namespace monitoring - Author a Hardened Production values.yaml Manifest:
The file below includes strict resource constraints, storage classes, and scrape optimizations.
# production-values.yaml
nameOverride: "k8s-mon"
prometheusOperator:
createCustomResource: true
resources:
limits:
cpu: 500m
memory: 512Mi
requests:
cpu: 100m
memory: 128Mi
nodeSelector:
topology.kubernetes.io/zone: "us-east-1a"
prometheus:
prometheusSpec:
image:
registry: quay.io
repository: prometheus/prometheus
tag: v2.55.1
replicas: 2
retention: 15d
retentionSize: 85Gi
scrapeInterval: "30s"
evaluationInterval: "30s"
serviceMonitorSelectorNilUsesHelmValues: false
podMonitorSelectorNilUsesHelmValues: false
ruleSelectorNilUsesHelmValues: false
walCompression: true
enableAdminAPI: false
resources:
requests:
cpu: "2000m"
memory: "8Gi"
limits:
cpu: "4000m"
memory: "16Gi"
storageSpec:
volumeClaimTemplate:
spec:
storageClassName: "gp3-ebs"
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 100Gi
securityContext:
runAsNonRoot: true
runAsUser: 1000
fsGroup: 2000
seccompProfile:
type: RuntimeDefault
alertmanager:
alertmanagerSpec:
replicas: 3
storage:
volumeClaimTemplate:
spec:
storageClassName: "gp3-ebs"
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 10Gi
resources:
requests:
cpu: "100m"
memory: "256Mi"
limits:
cpu: "500m"
memory: "1Gi"
grafana:
enabled: true
adminPassword: "SuperSecretStrongClusterPassword2026!"
persistence:
enabled: true
storageClassName: "gp3-ebs"
size: 20Gi
resources:
requests:
cpu: 200m
memory: 512Mi
limits:
cpu: 1000m
memory: 2Gi
- Deploy the Stack to the Target Cluster:
Execute the Helm deployment command passing the hardened configuration file.helm install prometheus-stack prometheus-community/kube-prometheus-stack \ --namespace monitoring \ --values production-values.yaml \ --wait \ --timeout 15m - Verify Pod Readiness Across the Pipeline:
Verify that StatefulSets and DaemonSets have reached steady state.kubectl get pods -n monitoring -o wide
Notice the two settings serviceMonitorSelectorNilUsesHelmValues: false and podMonitorSelectorNilUsesHelmValues: false. By default, the Helm chart configures Prometheus to discover only ServiceMonitors created by its own chart. Setting this property to false instructs the Prometheus Operator to discover all ServiceMonitors cluster-wide, regardless of which team or pipeline deployed them.
Instrumenting Custom Workloads with Prometheus Service Monitor CRDs
A frequent support issue encountered when engineers build out Prometheus monitoring Kubernetes clusters is the discovery failure of custom endpoints. The team adds instrumentation to their services, deploys a ServiceMonitor, and watches in frustration as target pools remain completely empty in the Prometheus UI. This occurs because engineers treat a ServiceMonitor as an active scraper rather than a declarative label selector.
The prometheus service monitor resource tells the Operator how to dynamically configure Prometheus to discover target pods via a matching Kubernetes Service object. There are three mandatory label references that must match precisely for discovery to succeed:
- Service Selector Match: The
spec.selector.matchLabelsinside theServiceMonitormust match the labels present in themetadata.labelsof the Kubernetes Service. - Port Identifier Match: The
spec.endpoints[].portvalue must align with thenameof the port defined inService.spec.ports[], not the raw numeric port. - Namespace Selection: The
spec.namespaceSelectormust either list the target namespace or declareany: trueto cross namespace perimeters.
The following diagram outlines the label-matching chain required to link pods to Prometheus targets:
+-----------------------------------------------------------+
| Workload Deployment Spec |
| metadata.labels: |
| app: orders-api |
+-----------------------------+-----------------------------+
| Selected by spec.selector
v
+-----------------------------------------------------------+
| Kubernetes Service Object |
| metadata.labels: |
| app: orders-api <---
| release: backend |
| spec.ports: | Must match labels!
| - name: http-metrics |
| port: 8080 |
+-----------------------------+----+------------------------+
^
| spec.selector.matchLabels:
| app: orders-api
| release: backend
+-----------------------------+-----------------------------+
| Prometheus ServiceMonitor CRD |
| spec: |
| endpoints: |
| - port: http-metrics <--- Must match port name! |
| path: /metrics |
+-----------------------------------------------------------+
Here is an end-to-end manifest implementing an instrumented Go microservice, its matching Service, and the declarative ServiceMonitor:
apiVersion: apps/v1
kind: Deployment
metadata:
name: orders-api
namespace: production
labels:
app: orders-api
spec:
replicas: 3
selector:
matchLabels:
app: orders-api
template:
metadata:
labels:
app: orders-api
spec:
containers:
- name: service
image: internal-registry.net/workloads/orders-api:v2.4.0
ports:
- name: http-metrics
containerPort: 8080
resources:
requests:
cpu: 250m
memory: 256Mi
limits:
cpu: 1000m
memory: 512Mi
---
apiVersion: v1
kind: Service
metadata:
name: orders-api-svc
namespace: production
labels:
app: orders-api
team: checkout
spec:
type: ClusterIP
selector:
app: orders-api
ports:
- name: http-metrics
port: 8080
targetPort: http-metrics
protocol: TCP
---
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: orders-api-monitor
namespace: production
labels:
team: checkout
spec:
selector:
matchLabels:
app: orders-api
team: checkout
namespaceSelector:
matchNames:
- production
endpoints:
- port: http-metrics
path: /metrics
interval: 15s
scrapeTimeout: 10s
metricRelabelings:
# Drop excessive internal runtime debug metrics
- sourceLabels: [__name__]
regex: "(go_gc_.*|go_threads)"
action: drop
# Enforce application label taxonomy
- targetLabel: environment
replacement: production
Troubleshooting Workflow: If targets do not display under
Status -> Targetsin the Prometheus web interface, execute these checks in sequence:
- Verify Endpoints generation:
kubectl get endpoints orders-api-svc -n production. If no IP addresses are listed, yourService.spec.selectordoes not match yourDeployment.spec.template.metadata.labels.- Inspect Operator RBAC: Ensure the Prometheus Operator service account possesses cluster-wide permissions to get, list, and watch
EndpointsandServicesacross application namespaces.- Verify Operator Discovery Logs: Run
kubectl logs -n monitoring -l app.kubernetes.io/name=prometheus-operatorto confirm if any CEL (Common Expression Language) or schema parsing rejections occurred.
Critical PromQL Alerting Rules for Cluster Health and Saturation
Collecting millions of metric samples per second provides zero operational value without well-calibrated, high-signal alerts. A common failure mode in Kubernetes operations is alert fatigue: on-call engineers receive alerts for harmless single-pod restarts while structural control plane latency surges pass unflagged.
Production PromQL alerting rules should adhere to the Golden Signals of site reliability engineering: latency, traffic, errors, and saturation. Alerts must measure the rate of degradation over sustained time windows rather than transient spikes.
| Alert Name | Severity | Evaluation Window | Target Failure Mode |
|---|---|---|---|
KubeContainerCrashLooping |
Critical | 5 minutes | Pod in CrashLoopBackOff failing to stay healthy |
KubeContainerThrottlingHigh |
Warning | 15 minutes | Aggressive CFS quota limits slowing response times |
KubeletPlegDurationHigh |
Critical | 10 minutes | Node runtime unresponsive; pod sync operations stuck |
APIServerLatencyHigh |
Critical | 5 minutes | Control plane saturation, etcd write stalls |
PersistentVolumeFillingFast |
Warning | 1 hour | Disk exhaustion projected within 24 hours |
The following PrometheusRule custom resource packages these critical alerting rules into declarative cluster manifests:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: kubernetes-cluster-health-rules
namespace: monitoring
labels:
role: alert-rules
spec:
groups:
- name: cluster.workload.health
rules:
# Alert on continuous container restarts indicating crash loops
- alert: KubeContainerCrashLooping
expr: >-
rate(kube_pod_container_status_restarts_total[5m]) * 60 > 2
for: 5m
labels:
severity: critical
annotations:
summary: "Container crash looping in {{ $labels.namespace }}/{{ $labels.pod }}"
description: "Container {{ $labels.container }} in pod {{ $labels.pod }} has restarted {{ $value | printf "%.2f" }} times per minute over the last 5 minutes."
# Detect severe CPU throttling (>25% throttled periods)
- alert: KubeContainerThrottlingHigh
expr: >-
sum(increase(container_cpu_cfs_throttled_periods_total[5m])) by (container, pod, namespace)
/
sum(increase(container_cpu_cfs_periods_total[5m])) by (container, pod, namespace)
> 0.25
for: 15m
labels:
severity: warning
annotations:
summary: "High CPU throttling for {{ $labels.namespace }}/{{ $labels.pod }}"
description: "Container {{ $labels.container }} is experiencing {{ $value | humanizePercentage }} CPU throttling, indicating CFS quota starvation."
- name: cluster.infrastructure.health
rules:
# Kubelet Pod Lifecycle Event Generator (PLEG) latency
- alert: KubeletPlegDurationHigh
expr: >-
node_quantile:kubelet_pleg_relist_duration_seconds:histogram_quantile{quantile="0.99"} > 10
for: 5m
labels:
severity: critical
annotations:
summary: "Kubelet PLEG latency critical on {{ $labels.node }}"
description: "PLEG relist duration is {{ $value }}s on node {{ $labels.node }}, indicating container runtime unresponsiveness."
# API Server p99 request latency degradation
- alert: APIServerLatencyHigh
expr: >-
histogram_quantile(0.99, sum(rate(apiserver_request_duration_seconds_bucket{verb!~"WATCH|CONNECT"}[5m])) by (le, verb, resource))
> 2.5
for: 5m
labels:
severity: critical
annotations:
summary: "API Server p99 latency above 2.5s on resource {{ $labels.resource }}"
description: "Kubernetes API Server is responding slowly for verb {{ $labels.verb }} on {{ $labels.resource }}."
# Predictive PVC volume exhaustion within 24 hours
- alert: PersistentVolumeFillingFast
expr: >-
(
kubelet_volume_stats_available_bytes
/
kubelet_volume_stats_capacity_bytes
) < 0.15
and
predict_linear(kubelet_volume_stats_available_bytes[6h], 24 * 3600) < 0
for: 30m
labels:
severity: warning
annotations:
summary: "PVC {{ $labels.persistentvolumeclaim }} running out of disk space"
description: "Storage capacity is below 15% and linear extrapolation predicts volume exhaustion within 24 hours."
TSDB Survival Guide: Taming High Cardinality and Horizontal Scaling
As a Kubernetes cluster expands past thousands of pods, Prometheus servers inevitably encounter high-cardinality bottlenecks. High cardinality refers to a scenario where metrics contain labels with millions of unique, constantly changing values: such as user IDs, UUIDs, IP addresses, or Git commit hashes. Because the TSDB engine creates an internal inverted index entry for every distinct combination of label key-value pairs, uncontrolled cardinality leads directly to exponential memory usage, massive block compaction pauses, and pod evictions.
To safeguard TSDB stability, apply metric relabeling rules directly at scrape boundaries to discard volatile parameters before they touch RAM:
# Metric relabeling configuration applied inside a ServiceMonitor or scrapeConfig
metricRelabelings:
# Strip out ephemeral, highly volatile labels before TSDB ingestion
- action: labeldrop
regex: "(trace_id|span_id|user_id|commit_hash|session_id)"
# Drop high-frequency, low-signal cAdvisor metrics
- sourceLabels: [__name__]
regex: "(container_tasks_state|container_memory_failures_total|container_sockets)"
action: drop
For clusters producing tens of millions of active time-series, vertical scaling of a single Prometheus StatefulSet reaches physical host memory limits. The ecosystem provides three primary architectures for scaling out storage and query performance horizontally.
| Scaling Architecture | Operational Model | Storage Strategy | Best-Fit Scenario |
|---|---|---|---|
| Standalone Prometheus | Single StatefulSet paired with local PVCs | Local disk filesystem (ext4/xfs) | Single cluster, under 2M active series, <30 day retention |
| Thanos | Sidecar pattern attached to local Prometheus | Object Storage (AWS S3, GCP GCS, Azure Blob) | Multi-cluster federation, historical retention, low operational footprint |
| Grafana Mimir | Centralized, microservices-based push model | Object storage with native block format | Massive multi-tenant clusters, global enterprise scale (>50M series) |
| VictoriaMetrics | Single binary or cluster with VMCluster CRD | Optimized local block or remote cloud object storage | High compression efficiency, constrained CPU/RAM budgets |
Cardinality Diagnostic Command: When memory surges suddenly, identify the offending metrics by executing a cardinality audit via Prometheus TSDB administrative endpoints using
kubectl exec:# Identify the top 10 metrics with the highest series count kubectl exec -it prometheus-k8s-mon-0 -n monitoring -c prometheus -- \ promtool tsdb analyze /prometheus/data
The standard enterprise scaling pattern uses Prometheus as a lightweight, stateless data collector on edge clusters, delegating long-term retention and global query aggregation to Thanos or Grafana Mimir. In this setup, Prometheus instances keep only 2 to 6 hours of telemetry in local WAL storage before shipping compressed immutable blocks directly into resilient cloud object storage.
Frequently Asked Questions
What is the primary difference between cAdvisor and kube-state-metrics?
cAdvisor measures resource consumption like container CPU, memory, and network usage directly from the container runtime. In contrast, kube-state-metrics listens to the Kubernetes API server to generate metrics about the health and status of deployments, pods, nodes, and PVCs.
Why is the kube-prometheus-stack preferred over manual manifests?
kube-prometheus-stack automates operations using the Prometheus Operator, providing declarative Custom Resource Definitions like ServiceMonitors. This prevents manual configuration drift, manages complex RBAC automatically, and includes pre-configured Alertmanager instances and production Grafana dashboards out of the box.
How does a Prometheus ServiceMonitor discover application endpoints?
A ServiceMonitor selects Kubernetes Services based on defined label selectors and namespace targets. Prometheus Operator detects these CRDs, parses matching endpoints, and automatically injects them into the Prometheus scrape configuration without requiring manual reloads or static IP definitions.
How do you resolve Prometheus OOMKilled errors in large clusters?
Mitigate Prometheus TSDB OOMKilled crashes by pruning high-cardinality labels using metricRelabelings, increasing memory limits, reducing scrape frequency for non-critical targets, and offloading long-term historical storage to horizontally scalable systems like Thanos or Grafana Mimir.
Running stable Prometheus monitoring Kubernetes systems at enterprise scale requires moving past standard Helm configurations and embracing a deterministic, well-bounded architecture. By relying on declarative Custom Resource Definitions like ServiceMonitor and PrometheusRule, platforms establish auditable monitoring configurations that maintain reliability through continuous cluster upgrades.
Protecting the Prometheus TSDB requires treating label cardinality with the same engineering rigor applied to database schema design. Pruning unneeded labels at the scrape edge, pairing targeted PromQL alerts with real user impact, and offloading historical persistence to systems like Thanos or Grafana Mimir transforms an unstable monitoring setup into an elite, highly reliable observability foundation.