Skip to main content
BaseMetric is the abstract foundation for every metric in TrustifAI, including all four built-in offline metrics. You can subclass it to create custom trust signals that receive the same service dependencies and integrate transparently with get_trust_score and the async batch pipeline. The only method you must implement is calculate.

Constructor

ExternalService
required
The shared service layer for LLM calls, embedding calls, and document text extraction. TrustifAI injects this automatically — you do not construct it directly.
Config
required
The parsed configuration object loaded from your YAML config file. Exposes threshold values, model names, and pipeline settings. TrustifAI injects this automatically.

Inherited attributes

All subclasses have access to these attributes after calling super().__init__():

Abstract methods

calculate

The synchronous evaluation entry point. You must implement this method. Receives a fully populated MetricContext (with embeddings already computed) and must return a MetricResult.

a_calculate (optional override)

The async variant. The default implementation simply delegates to calculate via the synchronous path. Override this method if your metric can make non-blocking LLM or embedding calls natively — for example, using await self.service.llm_call_async(...).

Custom metric example

The following example implements a query-answer relevance metric using cosine similarity between the query and answer embeddings:

Registering the custom metric

Use Trustifai.register_metric to add your class to the global metric registry, then configure its weight in config_file.yaml. Call register_metric before instantiating any engine.
The YAML configuration for your custom metric must appear in both metrics (for thresholds) and score_weights (for its contribution weight):
Use self.threshold_evaluator with one of its built-in evaluate_* methods whenever the score range aligns with an existing metric category. This ensures your custom metric respects the same configurable thresholds as the built-in metrics.
If your custom metric makes blocking I/O calls (LLM or embedding APIs), the default a_calculate will block the async event loop when used with evaluate_dataset. Override a_calculate with a native async implementation to maintain concurrency.