Overview
AWX exposes detailed metrics in Prometheus format via the/api/v2/metrics endpoint. These metrics provide visibility into system performance, job execution, resource utilization, and subsystem health.
Metrics Endpoint
Access Metrics
Endpoint Parameters
Enable Anonymous Access
Subsystem Metrics
The subsystem metrics system provides a flexible framework for collecting and aggregating metrics across AWX components.Architecture
How Subsystem Metrics Work
- Collection: Components track metrics in memory using
Metricsobjects - Aggregation: Metrics accumulate locally to minimize Redis overhead
- Persistence: Periodically save to Redis via
pipe_execute() - Broadcast: Metrics from each node are stored separately in Redis
- Exposure: API endpoint reads all node metrics and formats for Prometheus
Metric Types
Incrementing Metrics
IntM: Integer counterSet Metrics (Override)
SetIntM: Integer value (replaces previous)Histogram Metrics
HistogramM: Observations in bucketsUsing Metrics in Code
Basic Pattern
Thread Safety
Configuration
SUBSYSTEM_METRICS_INTERVAL_SAVE_TO_REDIS
Default:2 secondsDescription: Minimum interval between Redis saves
SUBSYSTEM_METRICS_INTERVAL_SEND_METRICS
Default:3 secondsDescription: Interval for broadcasting metrics to other nodes
SUBSYSTEM_METRICS_TASK_MANAGER_RECORD_INTERVAL
Default:15 secondsDescription: Task manager metrics recording interval
Set this to match or exceed your Prometheus scrape interval to avoid unnecessary overhead.
SUBSYSTEM_METRICS_BATCH_INSERT_BUCKETS
Default:[10, 50, 150, 350, 650, 2000]Description: Histogram buckets for batch insert metrics
Key Metrics Reference
Job Execution Metrics
Callback Receiver Metrics
Task Manager Metrics
Database Metrics
Instance Metrics
Prometheus Integration
Prometheus Configuration
Multi-Node Cluster
AWX automatically includes node labels in metrics. Scraping any control node returns metrics from all nodes in the cluster.
Service Discovery (Kubernetes)
Grafana Dashboards
Example Dashboard Panels
Job Throughput
Capacity Utilization
Job Queue Depth
Event Processing Rate
Task Manager Performance
P95 Job Completion Time
Dashboard Import
Create a comprehensive Grafana dashboard:- Create metrics user:
- Configure Prometheus datasource in Grafana
-
Import or create dashboard with panels for:
- Job execution rates
- Capacity utilization
- Queue depths
- Event processing
- Database performance
- Instance health
Alerting
Prometheus Alert Rules
Performance Monitoring
Job Performance
Capacity Trends
Database Performance
Direct Redis Access
Inspect raw metrics in Redis:Troubleshooting
Metrics Endpoint Returns 403
Cause: Insufficient permissions Solution:Missing Metrics from Some Nodes
Cause: Node not broadcasting metrics Solution:High Memory Usage from Metrics
Cause: Too frequent Redis updates Solution:Stale Metrics
Cause: Scrape interval mismatch Solution:- Ensure Prometheus scrape interval ≤
SUBSYSTEM_METRICS_TASK_MANAGER_RECORD_INTERVAL - Reduce
SUBSYSTEM_METRICS_INTERVAL_SAVE_TO_REDISfor fresher data
Best Practices
- Match scrape intervals: Align Prometheus scrape with AWX metric recording
- Monitor continuously: Set up alerts for critical metrics
- Baseline performance: Establish normal operating ranges
- Correlate metrics: Connect job performance with system resources
- Archive data: Retain long-term metrics for capacity planning
- Secure access: Use dedicated service accounts for metric collection
- Document thresholds: Define what constitutes “normal” for your workload
- Test under load: Validate metrics accuracy during peak usage
Related Resources
- Capacity Planning - Optimize resource usage
- Configuration - Configure metric collection
- Prometheus Documentation - Prometheus setup
- Grafana Documentation - Dashboard creation