Skip to main content

Overview

Policies define how agents select actions given states. Neurenix provides multiple policy types for different learning scenarios, supporting both discrete and continuous action spaces.

Base Policy Class

All policies inherit from the base Policy class:
Source: neurenix/rl/policy.py:16

Key Methods

Random Policy

Selects actions uniformly at random from the action space:
Source: neurenix/rl/policy.py:84

Continuous Action Spaces

Greedy Policy

Selects the action with the highest value according to a value function:
Source: neurenix/rl/policy.py:124

Epsilon-Greedy Policy

Balances exploration and exploitation with epsilon parameter:
Source: neurenix/rl/policy.py:174

Exploration Schedule

The epsilon value decays over time:
This ensures the agent:
  • Explores broadly early in training (high epsilon)
  • Exploits learned knowledge later (low epsilon)

Softmax Policy

Selects actions according to a Boltzmann distribution:
Source: neurenix/rl/policy.py:240

Temperature Parameter

  • High temperature (T >> 1): More uniform distribution (more exploration)
  • Low temperature (T → 0): More peaked distribution (more exploitation)

Gaussian Policy

For continuous action spaces, samples from Gaussian distribution:
Source: neurenix/rl/policy.py:300

Action Clipping

Actions are automatically clipped to valid range:

Policy Comparison

Using Policies with Agents

Source: neurenix/rl/agent.py:18

Custom Policies

Implement domain-specific action selection:

Best Practices

Exploration Schedule

Action Space Normalization

Policy Evaluation

Next Steps

Algorithms

Learn about RL algorithms

Value Functions

Understand value estimation