Skip to main content
The serve command deploys a trained Neurenix model as an API server, making it accessible for real-time inference.

Usage

Options

API Types

REST API

Default HTTP REST API with JSON payloads:
Endpoints:
  • POST /predict - Make predictions
  • GET /info - Get model information
  • GET /health - Health check endpoint

WebSocket API

Real-time bidirectional communication:
Connection:
  • WebSocket ws://host:port/ws

gRPC API

High-performance RPC:
Connection:
  • gRPC endpoint at host:port

Examples

Basic server

Custom host and port

GPU inference

Multiple workers

WebSocket server

gRPC server

Enable authentication

Enable CORS

Custom configuration

server_config.json:

Production deployment

Making Requests

REST API

Predict endpoint

Response:

Info endpoint

Response:

Health endpoint

Response:

WebSocket API

Python Client

Error Handling

Model not found

Port already in use

Solution: Use a different port:

Performance Tuning

Batch Size

Increase batch size for higher throughput:

Workers

Increase workers for concurrent requests:

GPU Acceleration

Best Practices

1. Use production-ready configuration

2. Monitor server health

3. Use reverse proxy for production

4. Set resource limits

5. Enable logging

Create a config file with logging:

6. Implement graceful shutdown

The server handles SIGINT and SIGTERM for graceful shutdown:

Deployment Scenarios

Local Development

Docker Container

Kubernetes

Cloud Deployment

Stopping the Server

Press Ctrl+C to stop the server gracefully:

See Also