Skip to main content
Offline metrics evaluate a response that has already been generated — they do not require hooking into the LLM generation process. You provide a query, an answer, and the retrieved documents, and TrustifAI scores the answer across four orthogonal dimensions of trustworthiness. Each metric is independent, configurable, and can be individually enabled or disabled based on your use case.

Metric overview


Evidence Coverage

Evidence Coverage checks whether every claim in the response is actually supported by the retrieved documents. It is the most heavily weighted metric by default (0.40) because unsupported claims are the most direct form of hallucination. How it works: The answer is passed to an LLM alongside the retrieved documents and the original query. The LLM performs sentence-level entailment checking — it breaks the answer into individual sentences and labels each one as supported: true or supported: false relative to the document content. The final score is the fraction of supported sentences.
This LLM-based NLI (Natural Language Inference) approach is more robust than embedding similarity alone because it reasons about meaning, not just vector proximity.

Threshold labels

Configuration

Result details


Semantic Drift

Semantic Drift measures how closely the response stays within the semantic envelope of the retrieved documents. A high-scoring response is topically close to the source material; a low-scoring one has drifted into territory not covered by the documents. How it works: TrustifAI splits each document into individual sentences and embeds them all. It then computes the cosine similarity between the answer embedding (and query embedding) and each sentence embedding, taking the best match. The final score reflects the peak semantic alignment between the answer and the document corpus.
Semantic Drift uses embedding similarity, not LLM reasoning. It catches cases where the response is topically off-base, but will not catch factual errors within the same topic. Use it alongside Evidence Coverage for full hallucination coverage.

Threshold labels

Configuration

Result details


Epistemic Consistency

Epistemic Consistency measures how stable the LLM’s responses are when asked the same question multiple times with different random seeds. Hallucinated answers tend to vary wildly between runs — a model that is genuinely confident produces semantically similar responses even under stochastic conditions. How it works: TrustifAI generates k additional responses to the same query at elevated temperatures (randomly sampled from [0.7, 0.8, 0.9, 1.0]). Each sample is embedded and compared to the original answer using cosine similarity. The final score is the mean cosine similarity across all samples — equivalent to 1 - σ in the semantic space.
The number of samples k is controlled by k_samples in your config. Setting k_samples: 0 skips generation and assumes full consistency (score = 1.0), which is useful for cost-sensitive pipelines.
Epistemic Consistency makes k additional LLM calls per evaluation. This increases both latency and API cost proportionally. Start with k_samples: 3 and increase only if you need higher confidence in the stability estimate.

Threshold labels

Configuration

Result details


Source Diversity

Source Diversity measures whether the response draws on multiple independent sources or leans entirely on a single document. Answers synthesized from multiple distinct sources are generally more trustworthy than those derived from a single reference. How it works: TrustifAI resolves a unique source ID for each retrieved document — using source, file_path, or url metadata fields if present, or falling back to a SHA-256 content hash. It then counts the number of distinct source IDs and computes a normalized score that combines a diversity ratio with a count-based exponential reward:
The exponential term rewards having more than one source non-linearly — the jump from 1 to 2 sources matters more than the jump from 5 to 6. If only one document is retrieved and it is the only semantically relevant one, TrustifAI considers this “justified” and assigns a score of 0.8 rather than penalizing the response.

Threshold labels

Configuration

Result details


Accessing individual metric scores

get_trust_score() returns all active metric scores in the details dictionary. Access individual metric results from the return value:
For advanced use cases, you can also instantiate individual metric classes directly. See the Offline Metrics API reference for class signatures.