Skip to main content

Overview

The AWX capacity system determines how many jobs can run on an instance based on available memory and CPU resources. Capacity management ensures efficient resource utilization while preventing system overload.

Capacity Fundamentals

Capacity is calculated based on:
  • Memory capacity (mem_capacity): Available system memory
  • CPU capacity (cpu_capacity): Available CPU cores
  • Forks: Number of simultaneous connections Ansible maintains

How Capacity Works

  1. Each instance has a calculated capacity based on hardware resources
  2. Jobs consume capacity based on their “impact” (primarily fork count)
  3. The task manager assigns jobs to instances with sufficient capacity
  4. When capacity is exhausted, jobs wait until resources free up
Capacity is not a zero-sum system. If only one instance is available, AWX allows jobs to run even if they exceed capacity, ensuring jobs don’t become permanently blocked.

Capacity Algorithms

Memory-Relative Capacity (Default)

Calculates capacity based on available memory, allowing CPU overcommit:
Example: 4GB system
Key Points:
  • Reserves 2GB for AWX services
  • Default: 100MB per fork (SYSTEM_TASK_FORKS_MEM)
  • Best for I/O-bound workloads
  • Protects against out-of-memory conditions
Configuration:

CPU-Relative Capacity

Calculates capacity based on CPU cores:
Example: 4-core system
Key Points:
  • Default: 4 forks per core (SYSTEM_TASK_FORKS_CPU)
  • Best for CPU-bound workloads
  • Reduces contention for compute resources
Configuration:

Capacity Adjustment

Balance between memory and CPU capacity using capacity_adjustment:
Values:
  • 0.0: Use minimum (most conservative)
  • 0.5: 50/50 balance
  • 1.0: Use maximum (most aggressive)
Example: CPU=16, Memory=20, adjustment=0.5
Set via API:

Job Impact

Job Types and Impact

Jobs have two impact types:

Control Impact

Fixed: AWX_CONTROL_NODE_TASK_IMPACT (default: 1)
Applied to: The instance controlling the job

Execution Impact

Variable: Based on job type
The +1 accounts for the Ansible parent process that coordinates execution.

Impact Examples

Example 1: Hybrid Node (Control + Execution)

Settings: AWX_CONTROL_NODE_TASK_IMPACT=1, forks=5, hosts=3

Example 2: Container Group Job

Settings: AWX_CONTROL_NODE_TASK_IMPACT=1

Example 3: Project Update

Settings: AWX_CONTROL_NODE_TASK_IMPACT=1

Control Node Task Impact

The AWX_CONTROL_NODE_TASK_IMPACT setting controls how much capacity controlling jobs consumes.

When to Adjust

Increase (AWX_CONTROL_NODE_TASK_IMPACT = 2 or higher):
  • Control plane CPU/memory usage is high
  • Many concurrent container group jobs
  • Job event processing is slow
  • Need to throttle concurrent jobs
Decrease (AWX_CONTROL_NODE_TASK_IMPACT = 0.5 or lower):
  • Control plane is underutilized
  • Most jobs run on execution nodes
  • Want more concurrent job control
Configuration:
Container groups have effectively infinite capacity. Without proper control plane throttling, you can overwhelm your control nodes with too many concurrent jobs.

Instance Groups

Instance Group Capacity

Instance groups aggregate capacity from member instances. Configure group-wide limits:

max_concurrent_jobs

Maximum concurrent jobs across entire group:

max_forks

Maximum total forks across entire group:

Container Group Capacity Planning

Calculate max_concurrent_jobs

Based on pod resource requests:
Example: 8GB node, 100MB pod

Calculate max_forks

Based on Ansible memory usage (100MB per fork):
Example: 8GB node
With max_forks=81:
  • 81 jobs with 1 fork each, OR
  • 40 jobs with 2 forks each, OR
  • 2 jobs with 40 forks each
Configure Container Group:

Capacity Monitoring

Check Instance Capacity

Monitor Running Jobs

Instance Group Status

Capacity Optimization

Recommendations by Workload

I/O-Bound Workloads

(Network operations, cloud APIs, service calls)
On instances:

CPU-Bound Workloads

(Computation, template rendering, encryption)
On instances:

Mixed Workloads

On instances:

Instance Sizing Guidelines

Dedicated Instance Groups

Create dedicated groups for specific workloads:

Troubleshooting

Jobs Stuck in Pending

Cause: Insufficient capacity
Solutions:
  1. Add more instances to the group
  2. Increase instance capacity (add memory/CPU)
  3. Adjust capacity_adjustment to 1.0
  4. Reduce job fork counts
  5. Add fallback instance groups

Capacity Calculations Seem Wrong

High Memory Usage

Symptoms: Jobs failing with OOM, system slowness Solutions:
  1. Reduce SYSTEM_TASK_FORKS_MEM (more conservative)
  2. Use CPU capacity (capacity_adjustment=0.0)
  3. Reduce JOB_EVENT_WORKERS
  4. Limit concurrent jobs on instance group
  5. Add more memory to instances

High CPU Usage

Symptoms: Slow job execution, high load average Solutions:
  1. Reduce SYSTEM_TASK_FORKS_CPU
  2. Use memory capacity (capacity_adjustment=1.0)
  3. Add more CPU cores
  4. Reduce job forks in templates
  5. Use dedicated execution nodes

Best Practices

  1. Monitor continuously: Track capacity metrics and adjust as needed
  2. Start conservative: Begin with lower capacity and increase gradually
  3. Separate workloads: Use dedicated instance groups for different job types
  4. Test under load: Simulate production load before going live
  5. Document changes: Keep records of capacity adjustments and their effects
  6. Plan for growth: Size instances with 20-30% headroom
  7. Use execution nodes: Offload job execution from control plane
  8. Throttle container groups: Set appropriate max_concurrent_jobs limits