Skip to main content

Overview

AWX exposes detailed metrics in Prometheus format via the /api/v2/metrics endpoint. These metrics provide visibility into system performance, job execution, resource utilization, and subsystem health.

Metrics Endpoint

Access Metrics

Endpoint Parameters

Enable Anonymous Access

Anonymous metrics access exposes system information. Only enable this in secure, trusted networks or behind authentication at the load balancer level.

Subsystem Metrics

The subsystem metrics system provides a flexible framework for collecting and aggregating metrics across AWX components.

Architecture

How Subsystem Metrics Work

  1. Collection: Components track metrics in memory using Metrics objects
  2. Aggregation: Metrics accumulate locally to minimize Redis overhead
  3. Persistence: Periodically save to Redis via pipe_execute()
  4. Broadcast: Metrics from each node are stored separately in Redis
  5. Exposure: API endpoint reads all node metrics and formats for Prometheus

Metric Types

Incrementing Metrics

IntM: Integer counter
FloatM: Floating-point counter

Set Metrics (Override)

SetIntM: Integer value (replaces previous)
SetFloatM: Float value (replaces previous)

Histogram Metrics

HistogramM: Observations in buckets
Generates Prometheus histogram:

Using Metrics in Code

Basic Pattern

Thread Safety

Each thread must create its own Metrics object. In-memory operations are not thread-safe, but pipe_execute() is thread-safe at the Redis level.

Configuration

SUBSYSTEM_METRICS_INTERVAL_SAVE_TO_REDIS

Default: 2 seconds
Description: Minimum interval between Redis saves

SUBSYSTEM_METRICS_INTERVAL_SEND_METRICS

Default: 3 seconds
Description: Interval for broadcasting metrics to other nodes

SUBSYSTEM_METRICS_TASK_MANAGER_RECORD_INTERVAL

Default: 15 seconds
Description: Task manager metrics recording interval
Set this to match or exceed your Prometheus scrape interval to avoid unnecessary overhead.

SUBSYSTEM_METRICS_BATCH_INSERT_BUCKETS

Default: [10, 50, 150, 350, 650, 2000]
Description: Histogram buckets for batch insert metrics

Key Metrics Reference

Job Execution Metrics

Callback Receiver Metrics

Task Manager Metrics

Database Metrics

Instance Metrics

Prometheus Integration

Prometheus Configuration

Multi-Node Cluster

AWX automatically includes node labels in metrics. Scraping any control node returns metrics from all nodes in the cluster.

Service Discovery (Kubernetes)

Grafana Dashboards

Example Dashboard Panels

Job Throughput

Capacity Utilization

Job Queue Depth

Event Processing Rate

Task Manager Performance

P95 Job Completion Time

Dashboard Import

Create a comprehensive Grafana dashboard:
  1. Create metrics user:
  1. Configure Prometheus datasource in Grafana
  2. Import or create dashboard with panels for:
    • Job execution rates
    • Capacity utilization
    • Queue depths
    • Event processing
    • Database performance
    • Instance health

Alerting

Prometheus Alert Rules

Performance Monitoring

Job Performance

Database Performance

Direct Redis Access

Inspect raw metrics in Redis:

Troubleshooting

Metrics Endpoint Returns 403

Cause: Insufficient permissions Solution:

Missing Metrics from Some Nodes

Cause: Node not broadcasting metrics Solution:

High Memory Usage from Metrics

Cause: Too frequent Redis updates Solution:

Stale Metrics

Cause: Scrape interval mismatch Solution:
  • Ensure Prometheus scrape interval ≤ SUBSYSTEM_METRICS_TASK_MANAGER_RECORD_INTERVAL
  • Reduce SUBSYSTEM_METRICS_INTERVAL_SAVE_TO_REDIS for fresher data

Best Practices

  1. Match scrape intervals: Align Prometheus scrape with AWX metric recording
  2. Monitor continuously: Set up alerts for critical metrics
  3. Baseline performance: Establish normal operating ranges
  4. Correlate metrics: Connect job performance with system resources
  5. Archive data: Retain long-term metrics for capacity planning
  6. Secure access: Use dedicated service accounts for metric collection
  7. Document thresholds: Define what constitutes “normal” for your workload
  8. Test under load: Validate metrics accuracy during peak usage