Skip to main content
AsyncTrustifai wraps the synchronous Trustifai engine in a thread-safe async interface. Each worker thread gets its own Trustifai instance via threading.local(), so concurrent evaluations never race on shared state. The evaluate_dataset function orchestrates concurrency, rate limiting, retries, and result ordering — letting you focus on your data rather than async plumbing.

Installation

evaluate_dataset uses tqdm for progress reporting. Install it alongside TrustifAI if you want the progress bar:

Basic usage

1

Build an AsyncTrustifai engine

Create one engine instance and share it across all evaluations in a session:
2

Prepare MetricContext objects

Each row in your dataset becomes a MetricContext. Documents can be plain strings, LangChain Document objects, LlamaIndex NodeWithScore objects, or dicts — see Integrations for details.
3

Run evaluate_dataset

Call evaluate_dataset inside an async context (a script’s asyncio.run, a FastAPI route, or a Jupyter cell):

Complete example

The following script reproduces the full example from examples/evaluation_script.py:

evaluate_dataset parameters

Rate limiting

evaluate_dataset uses a token-bucket RateLimiter that proactively spaces requests before they reach the API, preventing 429 errors before they occur. The semaphore caps how many evaluations run simultaneously; the rate limiter caps how fast they start.
Set requests_per_minute to roughly 80% of your actual API quota to leave headroom for retries. For free-tier Gemini or Mistral keys, 8–12 RPM is a safe starting point. OpenAI Tier-1 keys typically allow 500 RPM, so 400 is a good ceiling.

Retry backoff schedule

When a rate-limit error is detected, evaluate_dataset retries with exponential backoff plus ±25% jitter to prevent thundering-herd effects: Non-rate-limit errors (authentication failures, malformed requests) are re-raised immediately and are not retried.

Working with BatchResult

evaluate_dataset returns a BatchResult dataclass with the following attributes and properties:

Handling failures

Failures are isolated — a single failed evaluation never aborts the batch (unless fail_fast=True). Inspect .failed after the batch completes:

Pandas integration

batch.results is a plain list of dicts, making it straightforward to load into a DataFrame for further analysis:

Jupyter usage

Jupyter already runs an event loop, so asyncio.run() raises a RuntimeError. Use nest_asyncio to patch the running loop instead:

Configuration

Set concurrency defaults and LLM credentials in config_file.yaml.

Integrations

Feed LangChain, LlamaIndex, or plain string documents into MetricContext.