Skip to main content

Overview

pyinfra is built around a unique two-phase execution model that separates operation discovery from execution. This architecture enables predictable, idempotent infrastructure management with efficient parallel execution.
The two-phase model is fundamental to how pyinfra works. Understanding this concept is key to mastering the framework.

The Two-Phase Model

pyinfra’s execution is divided into two distinct phases:

Phase 1: Prepare (Discovery/Diff)

During the Prepare phase, pyinfra:
  1. Discovers operations: Executes your deploy scripts to discover all operations
  2. Collects facts: Gathers current state from target hosts
  3. Determines changes: Compares desired state against current state
  4. Builds operation DAG: Creates a dependency graph of operations
  5. Generates command list: Produces the exact commands needed
During this phase, operations are called but not executed. The @operation decorator intercepts the call and returns an OperationMeta object instead:
The Prepare phase is what enables pyinfra’s “dry-run” mode (pyinfra ... --dry). Since operations are discovered but not executed, you can see exactly what would change.

Phase 2: Execute

During the Execute phase, pyinfra:
  1. Sorts operations: Uses the DAG to determine optimal execution order
  2. Executes in parallel: Runs operations across hosts concurrently
  3. Handles errors: Manages failures and retries
  4. Collects results: Tracks success/failure for each operation
  5. Completes operations: Marks operations as complete with results

State Management

The State class (defined in src/pyinfra/api/state.py) is the central coordinator for a pyinfra deployment:

State Data Structures

pyinfra uses several key data structures to manage operations:

StateOperationMeta

Shared metadata about an operation across all hosts:

StateOperationHostData

Host-specific operation data:

OperationMeta

The return value from operations that tracks execution status:

Operation Execution Flow

Here’s the complete flow from operation call to execution:

Parallel Execution

pyinfra uses gevent for concurrent execution across multiple hosts:
By default, pyinfra calculates optimal parallelism:
The maximum parallelism is limited by the system’s file descriptor limit, since each SSH connection requires file descriptors.

Operation DAG (Dependency Graph)

Operations are ordered using a Directed Acyclic Graph (DAG) to handle dependencies:
The DAG ensures operations run in the correct order while maximizing parallelism. Operations with no dependencies can run simultaneously across different hosts.

Context Management

pyinfra uses context variables to track the current host and state:
This allows operations and facts to access the current execution context without explicit parameter passing.

Error Handling and Retries

pyinfra includes sophisticated error handling with retry support:

Callback System

pyinfra provides callback hooks for monitoring execution:

Command Types

pyinfra supports multiple command types:

StringCommand

Shell commands to execute:

FunctionCommand

Python functions to run:

FileUploadCommand / FileDownloadCommand

File transfer operations:

Best Practices

Idempotency

Design operations to be safely re-runnable. The two-phase model helps by checking state first.

Fact Usage

Use facts to check current state instead of assuming. Let pyinfra determine what needs to change.

Operation Order

Rely on the DAG for dependencies. Define operations in logical order and let pyinfra optimize execution.

Error Handling

Use _ignore_errors, _continue_on_error, and _retries for robust deployments.

Performance Characteristics

Prepare Phase

  • Time complexity: O(hosts × operations × facts)
  • Parallelism: Facts are collected in parallel across hosts
  • Network: One round-trip per unique fact per host

Execute Phase

  • Time complexity: O(operations × commands_per_operation)
  • Parallelism: Operations run in parallel across hosts (respecting DAG)
  • Network: One round-trip per command (pipelined where possible)

Key Takeaways

  1. Two phases: Prepare discovers and diffs, Execute runs commands
  2. Operation generators: Commands are generated lazily via iterators
  3. DAG ordering: Operations are topologically sorted for optimal execution
  4. Parallel execution: gevent enables concurrent operations across hosts
  5. State tracking: Comprehensive state management through State, Host, and OperationMeta classes
  6. Context-aware: Operations access current host/state via context variables

Operations

Learn how to define operations using the @operation decorator

Facts

Understand how facts collect state information

State

Deep dive into the State class and state management

Connectors

Explore how connectors interface with target systems