Skip to main content
TrustifAI’s metric system is a plugin registry. Every built-in metric — evidence coverage, epistemic consistency, semantic drift, and source diversity — is registered against a string key and instantiated at evaluation time. You can add your own metrics by following the same three-step pattern: inherit, register, configure.

The three-step process

1

Inherit from BaseMetric and implement calculate()

Your class must inherit from BaseMetric and implement calculate(context: MetricContext) -> MetricResult. The context argument carries the query, answer, documents, and pre-computed embeddings for a single evaluation.
2

Register the metric class

Call Trustifai.register_metric with a unique string key. This key must match the type field you will add to config_file.yaml.
Registration is a class-level operation — call it once, before you instantiate any Trustifai engine.
3

Add the metric to config_file.yaml

Add entries to both the metrics list (to set thresholds and mark it enabled) and the score_weights list (to assign its contribution to the final Trust Score). Weights across all enabled metrics must sum to at most 1.0.

Full example: TemporalConsistencyMetric

The following example detects temporal hallucinations — cases where the answer references dates or times that are not present in the retrieved documents. It is the canonical custom metric example from the TrustifAI README.

Metric implementation

Registration and usage

Updated config_file.yaml

MetricResult fields

Every calculate() implementation must return a MetricResult. The to_dict() method serializes it into the format consumed by the trust score aggregator.

BaseMetric helpers

When you inherit from BaseMetric, your class automatically gets access to: Use self.service.llm_call(prompt, system_prompt) if your metric needs an LLM inference step, and self.service.embedding_call(text) for additional embeddings beyond what the engine pre-computes.

Async metrics

BaseMetric provides a default async implementation that calls calculate() synchronously:
Override a_calculate if your metric can benefit from native async I/O (for example, if it makes multiple LLM calls that can be parallelized):
When adding a new metric, make sure the total of all score_weights still sums to at most 1.0. TrustifAI raises a ValueError at startup if the sum exceeds this limit. Reduce existing weights proportionally to accommodate the new metric’s weight.

Configuration

Learn the full config_file.yaml schema including metric thresholds and weights.

BaseMetric API

Full API reference for BaseMetric, MetricResult, and MetricContext.