Skip to main content

Framework Architecture

Neurenix is built on a flexible, modular architecture that enables seamless switching between different hardware backends at runtime. The framework consists of three core systems:

Hot-Swappable Backends

Switch between CPU, GPU, TPU, and other accelerators without code changes

Genesis System

Intelligent hardware detection and automatic device selection

Device Manager

Centralized device orchestration and memory management

Core Components

Device Manager

The DeviceManager is a singleton that orchestrates all device operations and provides hot-swappable backend functionality.
The DeviceManager uses the singleton pattern, ensuring a single instance manages all device operations across your application.

Genesis: Intelligent Device Selection

Genesis automatically detects available hardware and selects the optimal device for your workload.
Genesis prioritizes TPUs for inference workloads and CUDA/ROCm devices for training, automatically falling back to CPU when specialized hardware is unavailable.

Workload-Specific Selection

Supported Backends

Neurenix supports an extensive range of hardware backends:

Device Benchmarking

Genesis can benchmark your hardware to optimize device selection:

Memory Management

The DeviceManager tracks memory usage across all devices:

Device Synchronization

Synchronize GPU operations to ensure computation completes:

Architecture Benefits

Portability

Write once, run on any hardware backend without modification

Performance

Automatic selection of optimal hardware for each workload

Flexibility

Hot-swap between devices at runtime for testing and optimization

Simplicity

High-level API abstracts hardware complexity

Best Practices

Recommendation: Let Genesis handle device selection for production workloads. Manual device selection is best reserved for debugging and specific optimization scenarios.
  1. Use Genesis for automatic selection - It considers memory, performance, and workload type
  2. Synchronize before timing - GPU operations are asynchronous
  3. Monitor memory usage - Especially important for large models on GPU
  4. Benchmark your hardware - Run genesis.benchmark_devices() once to optimize future selections