A Prometheus scrape config (scrape_configs) in prometheus.yml defines how Prometheus discovers, scrapes, authenticates with, relabels, and ingests metrics from operational endpoints into its TSDB. Operating without explicit timeouts, memory limits, and target relabeling rules is one of the most common causes of out-of-memory (OOM) crashes and silent telemetry outages in enterprise clusters.
When Prometheus initiates a scrape loop, it pulls metrics over HTTP/HTTPS from hundreds or thousands of distributed endpoints simultaneously. If an exporter slows down or returns millions of high-cardinality series unexpectedly, Prometheus can exhaust socket pools, hit context deadlines, and crash under memory pressure. Misconfiguring pre-scrape target discovery versus post-scrape metric ingestion creates massive storage churn and obscures real production incidents.
This engineering blueprint deconstructs the internal scrape engine, contrasts global settings with per-job directives, delivers copy-paste configurations for static and dynamic environments (such as Kubernetes pods and AWS EC2), demystifies relabeling pipelines, and establishes production guardrails to ensure your monitoring layer remains resilient at scale.
Anatomy of prometheus.yml: Global Directives vs Scrape Configurations
Every Prometheus deployment centers on the root configuration file, conventionally named prometheus.yml. The file is split into functional hierarchies: top-level settings under global establish default intervals and timeouts across the entire instance, while the scrape_configs array defines individual operational jobs with granular controls. Understanding how per-job directives inherit and override these global defaults is fundamental to effective prometheus configuration.
Configuration Hierarchy and Fallback Logic
When Prometheus initiates a scraping cycle for a specific job, it looks for job-level keys first. If a directive like scrape_interval or scrape_timeout is omitted in the job stanza, the daemon automatically falls back to the values defined under global. However, specifying a job-level override completely supersedes the global directive for all targets in that job.
# /etc/prometheus/prometheus.yml
global:
scrape_interval: 15s # Scrape targets every 15 seconds by default.
scrape_timeout: 10s # Abort scrape if the target does not reply within 10 seconds.
evaluation_interval: 15s # Evaluate alerting and recording rules every 15 seconds.
external_labels:
cluster: 'us-east-prod'
datacenter: 'iad-01'
scrape_configs:
# Baseline system metrics: inherits global 15s scrape_interval
- job_name: 'node-exporter'
static_configs:
- targets: ['10.0.1.10:9100', '10.0.1.11:9100']
# Slow metrics target: requires longer interval and explicit timeout override
- job_name: 'deep-telemetry'
scrape_interval: 60s # Overrides global 15s
scrape_timeout: 45s # Overrides global 10s; MUST be <= scrape_interval
metrics_path: '/custom-metrics'
scheme: 'https'
tls_config:
insecure_skip_verify: false
ca_file: '/etc/prometheus/certs/ca.crt'
static_configs:
- targets: ['analytics-worker.internal.net:8443']
Rule of Scrape Deadlines: The
scrape_timeoutmust always be less than or equal to thescrape_interval. Settingscrape_timeout: 30swithscrape_interval: 15scauses Prometheus to reject the file during parsing with the error:scrape timeout greater than scrape interval.
Global Directives vs Per-Job Directives
To design an efficient prometheus scrape config, you must clearly distinguish between settings applied globally and parameters scoped to a target pool. The following table highlights the boundaries and behavior of key configuration keys:
| Configuration Directive | Scope Level | Default Behavior | Production Best Practice |
|---|---|---|---|
scrape_interval |
Global / Job | 1m (if not set in global) | Set to 15s or 30s globally; increase to 60s for slow custom exporters. |
scrape_timeout |
Global / Job | 10s (if not set in global) | Keep 2-5 seconds below interval to avoid overlapping scrapes and queue lockups. |
evaluation_interval |
Global | 1m | Match scrape_interval to keep recording rules aligned with incoming sample ticks. |
external_labels |
Global | Empty | Use exclusively to label Prometheus server identity (cluster, region) in federated/Cortex/Thanos setups. |
metrics_path |
Job only | /metrics |
Leave default unless monitoring specialized daemons exposing non-standard paths. |
scheme |
Job only | http |
Use https along with mTLS for production environments traversing unsegmented networks. |
A well-structured prometheus yml maintains a lean global footprint, relying on specialized job stanzas to accommodate services with divergent latency profiles and varying network security requirements.
Core Mechanics of Prometheus Scraping and Pull Architecture
At its core, prometheus scraping is an active, deterministic HTTP client operation driven by internal scrapeloop worker routines. Rather than waiting for services to push metrics, Prometheus schedules recurring HTTP requests against target addresses, streams the text-based payload, parses samples on the fly, and ingests them into the time-series database (TSDB) head chunk.
The Scrape Loop Lifecycle
Every discovered target runs inside its own lightweight Go goroutine managed by the Prometheus scrape manager. The lifecycle follows strict synchronous phases during each polling tick:
+-----------------------------------------------------------------------------------------+
| Prometheus Scrapeloop Lifecycle |
+-----------------------------------------------------------------------------------------+
[ Service Discovery ]
|
v
[ Target Relabeling ] (relabel_configs: Drop unwanted targets, build __address__)
|
v
+---------------------------------------------------------------------------------+
| Loop Interval Tick (scrape_interval: 15s, Deadline context: scrape_timeout) |
| |
| 1. HTTP GET http(s)://<target><port>/metrics |
| 2. Read Content-Type (text/plain, OpenMetrics, or Protobuf) |
| 3. Stream Response Body via Parser |
| 4. Enforce Guardrails (body_size_limit, sample_limit) |
+---------------------------------------------------------------------------------+
|
v
[ Metric Relabeling ] (metric_relabel_configs: Drop metrics, rewrite series labels)
|
v
[ TSDB Storage Appender ]
- Commit Samples to Head Block
- Update scrape synthetic metrics (up, scrape_duration_seconds)
Metrics Parsing and Synthetic Metadata
During an active prometheus scrape, Prometheus issues an HTTP GET with an Accept header specifying support for OpenMetrics and Prometheus text formats: application/openmetrics-text;version=1.0.0,text/plain;version=0.0.4;q=0.5. As the target emits lines matching the exposition format, Prometheus validates their syntax:
# HELP http_requests_total Total HTTP requests handled.
# TYPE http_requests_total counter
http_requests_total{method="POST",handler="/api/v1/checkout",status="200"} 10492 1774312800000
In parallel with the target’s exposed metrics, Prometheus synthesizes administrative telemetry representing the outcome of the scrape itself. These synthetic metrics are automatically injected into the TSDB for every scrape attempt:
up: Returns1if the target was reachable and responded with HTTP status 200 within the timeout window; returns0if unreachable, timed out, or returned an HTTP error.scrape_duration_seconds: Measures the complete elapsed wall-clock time from the initial socket connection to the final byte read of the response body.scrape_samples_scraped: The raw count of samples produced by the target payload before any post-scrape filters are applied.scrape_samples_post_metric_relabeling: The net count of samples committed to the TSDB after metric relabeling rules execute.scrape_series_added: The number of previously unseen time series registered in the TSDB inverted index during this specific scrape.
Under the Hood: If a target takes 9.8 seconds to respond when
scrape_timeoutis set to 10 seconds, the scrape succeeds, butscrape_duration_secondswill reflect the bottleneck. If network jitter pushes the response time to 10.01 seconds, Prometheus instantly cancels the underlying HTTP client context, marks the target asup == 0, drops all metrics from that scrape iteration, and records a context deadline exceeded error.
Step-by-Step Prometheus Tutorial: Static Targets vs Dynamic Discovery
Configuring endpoints requires selecting the correct target provider. While static configurations work well for immutable bare-metal devices, dynamic platforms like cloud providers and container schedulers demand automated service discovery (SD). This prometheus tutorial details the setup for both environments.
Configuring Static Targets
Use static target configurations for infrastructure with predictable IP addresses or persistent DNS records, such as physical network routers, persistent database appliances, and fixed virtualization hosts.
# Static configuration for bare-metal fleet
scrape_configs:
- job_name: 'baremetal-infrastructure'
scrape_interval: 15s
scrape_timeout: 10s
static_configs:
- targets:
- 'db-master-01.internal:9100'
- 'db-replica-01.internal:9100'
- 'db-replica-02.internal:9100'
labels:
tier: 'database'
environment: 'production'
- targets:
- 'edge-gw-01.internal:9100'
- 'edge-gw-02.internal:9100'
labels:
tier: 'gateway'
environment: 'production'
Implementing Dynamic Service Discovery
In elastic environments, hardcoding host lists is unmanageable. Prometheus integrates native discovery providers that query orchestration APIs to update target lists dynamically without restarting the server.
Pattern 1: Kubernetes Pod Discovery with Annotation Filters
The following prometheus scrape configuration scans the Kubernetes API for pods, dropping any pod that lacks the prometheus.io/scrape: "true" annotation:
scrape_configs:
- job_name: 'kubernetes-pods'
kubernetes_sd_configs:
- role: pod
relabel_configs:
# 1. Drop pods that lack the scrape annotation
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_scrape]
action: keep
regex: true
# 2. Extract custom metrics path if configured
- source_labels: [__meta_kubernetes_pod_annotation_prometheus_io_path]
action: replace
target_label: __metrics_path__
regex: (.+)
# 3. Handle custom scrape ports
- source_labels: [__address__, __meta_kubernetes_pod_annotation_prometheus_io_port]
action: replace
regex: ([^:]+)(?:\d+)?(\d+)
replacement: $1:$2
target_label: __address__
# 4. Map Kubernetes metadata to clean operational labels
- source_labels: [__meta_kubernetes_namespace]
action: replace
target_label: k8s_namespace
- source_labels: [__meta_kubernetes_pod_name]
action: replace
target_label: k8s_pod
Pattern 2: AWS EC2 Discovery Filtered by Tags
For cloud compute instances running in Amazon Web Services, Prometheus can query the EC2 API across specific regions and availability zones:
scrape_configs:
- job_name: 'aws-ec2-nodes'
ec2_sd_configs:
- region: 'us-east-1'
port: 9100
filters:
- name: 'instance-state-name'
values: ['running']
- name: 'tag:Monitoring'
values: ['enabled']
relabel_configs:
# Use the private IPv4 address for internal network scraping
- source_labels: [__meta_ec2_private_ip]
action: replace
target_label: __address__
replacement: '${1}:9100'
# Propagate instance tag values into searchable labels
- source_labels: [__meta_ec2_tag_Environment]
action: replace
target_label: environment
- source_labels: [__meta_ec2_instance_id]
action: replace
target_label: instance_id
Step-by-Step Discovery Validation Runbook
- Define Scrape Job: Add the service discovery stanza to your Prometheus YAML file.
- Validate Syntax: Run
promtool check config prometheus.ymlon your control machine. - Trigger Reload: Dispatch an HTTP POST request to the
/-/reloadendpoint. - Inspect Discovered Targets: Open the Prometheus Web UI at
http://<prometheus-host>9090/service-discovery. - Verify Metadata Mapping: Confirm that discovered metadata labels (prefixed with
__meta_) successfully resolve into the target’s labels without errors.
Relabeling Engine: relabel_configs vs metric_relabel_configs
A critical source of confusion in prometheus configuration is the operational distinction between relabel_configs and metric_relabel_configs. Applying relabeling rules at the wrong stage can lead to missing targets or unchecked metric ingestion, driving up TSDB costs and memory consumption.
Target Relabeling vs Metric Relabeling
The core difference comes down to execution timing in the pipeline: target relabeling executes before the HTTP request is made, while metric relabeling executes after the metrics payload is received but before it is written to disk.
| Operational Dimension | relabel_configs (Target Stage) | metric_relabel_configs (Sample Stage) |
|---|---|---|
| Execution Point | Pre-scrape (target discovery phase) | Post-scrape (ingestion phase) |
| Input Data | Target metadata labels (e.g. __meta_*, __address__) |
Parsed metric series names and label sets |
| Primary Objective | Target filtering, address rewriting, setting default job labels | Dropping noisy/unused metrics, scrubbing cardinality |
| Bandwidth Impact | Saves network bandwidth (dropped targets are never scraped) | Consumes scrape bandwidth (metrics are fetched over HTTP first) |
| TSDB Impact | Prevents target initialization | Prevents unwanted time series from writing to TSDB |
Under the Hood: The Internal Variable Pipeline
During the target relabeling phase, Prometheus maintains internal variables that govern connection logic:
__address__: The host and port used for the HTTP request (e.g.10.0.0.12:9100).__scheme__: The protocol scheme (defaults tohttp, can be rewritten tohttps).__metrics_path__: The HTTP request URI (defaults to/metrics).__param_<name>: URL query parameters passed along with the scrape request.
Any label prefixed with two underscores (__) is considered internal and is automatically stripped from the metric before storage in the TSDB, unless an explicit rule maps it to a standard label name.
Practical Blueprint: Filtering High-Cardinality Metrics
The following example uses both pipelines. First, it drops unhealthy targets during discovery. Second, it drops unwanted, high-frequency metrics at the ingestion stage to protect the TSDB:
scrape_configs:
- job_name: 'api-gateway'
static_configs:
- targets: ['api-gw-01.prod:8080', 'api-gw-02.prod:8080']
labels:
env: 'production'
# Phase 1: Pre-Scrape Target Relabeling
relabel_configs:
# Ignore canary target addresses during maintenance windows
- source_labels: [__address__]
regex: '.*-canary.*'
action: drop
# Phase 2: Post-Scrape Metric Relabeling
metric_relabel_configs:
# Drop high-cardinality debugging histograms
- source_labels: [__name__]
regex: '(http_request_duration_seconds_bucket|jvm_gc_memory_allocated_bytes_total)'
action: drop
# Strip user-agent labels to prevent unbounded series explosion
- regex: 'user_agent|client_ip'
action: labeldrop
# Rewrite legacy label key names to align with corporate standards
- source_labels: [old_service_tag]
target_label: service
action: replace
Using metric_relabel_configs to drop series before TSDB storage is one of the most effective ways to lower RAM usage and index size in high-throughput Prometheus clusters.
Production Hardening: Guardrails, Limits, and Authentication
Running an unconstrained prometheus scrape config in production can lead to reliability issues. If an application update inadvertently exports millions of unique metric series, Prometheus can run out of memory, crash, and enter an unrecoverable crash loop while replaying write-ahead logs (WAL). Protecting your instance requires robust guardrails, resource ceilings, and secure transport layers.
Memory Guardrails and Protective Limits
Prometheus includes built-in safeguards to cap resource utilization during scrapes. These limits should be applied to all high-throughput jobs:
scrape_configs:
- job_name: 'saas-microservice'
scrape_interval: 15s
scrape_timeout: 10s
# --- INGESTION HARDENING LIMITS ---
# Maximum number of scraped samples accepted per scrape.
# If exceeded, Prometheus drops ALL metrics for that scrape and marks target failed.
sample_limit: 50000
# Maximum number of targets allowed to be discovered and scraped in this job.
target_limit: 150
# Maximum uncompressed response size allowed. Prevents OOM from massive HTTP payloads.
body_size_limit: 15MB
# Maximum label pairs per series (prevents excessive metadata attacks).
label_limit: 40
# Maximum length of label names and label values in characters.
label_name_length_limit: 128
label_value_length_limit: 512
static_configs:
- targets: ['app-svc-01.internal:8080']
Hardening Authentication and Transport Security
Metrics endpoints often expose sensitive operational metadata, including internal hostnames, software versions, and query patterns. Transport encryption and mutual authentication should be enforced across zero-trust networks:
scrape_configs:
- job_name: 'secure-workloads'
scheme: 'https'
metrics_path: '/federated-metrics'
# Bearer token authentication (e.g. Kubernetes service accounts)
bearer_token_file: '/var/run/secrets/kubernetes.io/serviceaccount/token'
# Optional HTTP Basic Authentication fallback
# basic_auth:
# username: 'prometheus-collector'
# password_file: '/etc/prometheus/secrets/collector-pass.txt'
tls_config:
# Root Certificate Authority used to sign the target certificates
ca_file: '/etc/prometheus/certs/internal-ca.crt'
# Client certificates for Mutual TLS (mTLS)
cert_file: '/etc/prometheus/certs/prom-client.crt'
key_file: '/etc/prometheus/certs/prom-client.key'
# Strict hostname validation
server_name: 'telemetry.platform.internal'
insecure_skip_verify: false
static_configs:
- targets: ['telemetry.platform.internal:9443']
Production Scrape Hardening Checklist
- [ ] Ensure
scrape_timeoutis explicitly defined and strictly lower thanscrape_interval. - [ ] Set a conservative
sample_limiton all third-party and community exporters. - [ ] Configure
body_size_limitto stop large, misconfigured metric payloads before they are parsed. - [ ] Verify that
insecure_skip_verifyis set tofalsein production environments. - [ ] Store authentication secrets in separate files using
bearer_token_fileorpassword_filerather than hardcoding credentials directly in the YAML configuration.
Validating and Inspecting Targets in Prometheus Web UI
Applying a broken scrape configuration can interrupt metric collection across an entire cluster. Establishing a reliable validation pipeline using the Prometheus CLI and the built-in prometheus web ui ensures changes can be deployed and monitored safely.
Pre-Deployment Syntax Validation with promtool
Never reload or restart a Prometheus instance without first verifying the configuration using the bundled promtool utility. This tool checks YAML syntax, validates relabeling action semantics, and catches missing files:
# Validate configuration syntax and reference rules
promtool check config /etc/prometheus/prometheus.yml
# Expected success output:
# SUCCESS: /etc/prometheus/prometheus.yml is valid and all rule files are good!
If a property is misconfigured, such as setting a scrape timeout longer than the scrape interval, promtool identifies the line number and the exact failure reason:
FAILED: /etc/prometheus/prometheus.yml:
scrape_configs[0].scrape_timeout: "30s" is greater than scrape_interval: "15s"
Zero-Downtime Configuration Reloads
Prometheus does not require a full service restart to apply updated scrape configurations. You can reload configuration on the fly using either of these approaches:
# Method 1: Send a process hang-up signal (SIGHUP)
kill -HUP $(pgrep prometheus)
# Method 2: Trigger the lifecycle HTTP reload endpoint (requires --web.enable-lifecycle flag)
curl -X POST http://localhost:9090/-/reload
Troubleshooting and Debugging in the Prometheus UI
Once reloaded, use the prometheus ui to verify target health and debug configuration errors:
Prometheus Web UI -> Status -> Targets
Job: 'node-exporter' (2/2 UP)
Endpoint State Labels Last Scrape Scrape Duration Error
------------------------------------------------------------------------------------------------------
http://10.0.1.10:9100 UP env="prod", instance="node1" 1.2s ago 12.4ms ""
http://10.0.1.11:9100 DOWN env="prod", instance="node2" 4.1s ago 0ms "dial tcp 10.0.1.11:9100: connect: connection refused"
Diagnosing Common Target Failure States
- dial tcp.. connect: connection refused: The target exporter process is not running, bound to the wrong interface, or blocked by local host firewall rules.
- context deadline exceeded: The target took longer to respond than the configured
scrape_timeout. Investigate heavy network latency, CPU saturation on the target, or expensive runtime collection metrics. - server returned HTTP status 401 Unauthorized: Authentication credentials failed. Verify paths set in
bearer_token_fileor check the username/password in your basic auth configuration. - x509: certificate signed by unknown authority: The target presented an invalid or self-signed TLS certificate. Ensure the signing authority is specified in the
ca_fileblock undertls_config. - samples exceed limit: The target exported more samples than allowed by
sample_limit. Increase the threshold if appropriate, or usemetric_relabel_configsto filter out noisy metrics.
Pro Tip: Query
rate(scrape_duration_seconds_sum[5m]) / rate(scrape_duration_seconds_count[5m])directly in the Prometheus graph console to spot targets that are approaching their scrape timeout window before they trigger alerts.
Frequently Asked Questions
What is the primary role of scrape_configs in prometheus.yml?
The scrape_configs section in prometheus.yml defines the endpoints, polling intervals, HTTP paths, and authentication methods Prometheus uses to collect metrics. Each entry represents a distinct monitoring job, controlling target discovery, scrape timeouts, relabeling rules, and ingestion limits before storing data into the TSDB.
How do you check target health in the Prometheus UI?
To check target health in the Prometheus UI, open the web console and navigate to Status then Targets. This screen displays all configured endpoints grouped by job name, showing their current state (UP or DOWN), last scrape timestamp, scrape duration, error messages, and evaluated labels.
What is the difference between relabel_configs and metric_relabel_configs?
relabel_configs operates prior to the scrape, using target metadata to dynamically set labels or decide whether to drop an entire target. metric_relabel_configs operates immediately after the scrape, filtering out specific ingested metrics or altering series labels before writing the raw samples to TSDB storage.
How do you safely validate a new Prometheus scrape config before reloading?
Validate syntax by executing the CLI command ‘promtool check config prometheus.yml’. If the check passes without syntax or schema errors, apply changes with zero downtime by triggering an HTTP POST request to the ‘/-/reload’ endpoint or sending a SIGHUP signal directly to the Prometheus process.
A reliable Prometheus monitoring infrastructure relies on carefully structured scrape configurations. By establishing sound global defaults, matching scrape intervals to system latency profiles, and isolating slow exporters into dedicated jobs, you can prevent collection bottlenecks across your environments. Using relabel_configs during the discovery phase ensures Prometheus scrapes only healthy, intended infrastructure, while metric_relabel_configs keeps metric bloat and TSDB cardinality under control.
Protect your deployment by enforcing guardrails like sample_limit, body_size_limit, and mTLS verification across every target pool. Integrate automated syntax validation with promtool check config into your CI/CD pipelines, and use the runtime debug visibility provided by the Prometheus Web UI to ensure stable, reliable telemetry at scale.