Metric overview
Evidence Coverage
Evidence Coverage checks whether every claim in the response is actually supported by the retrieved documents. It is the most heavily weighted metric by default (0.40) because unsupported claims are the most direct form of hallucination. How it works: The answer is passed to an LLM alongside the retrieved documents and the original query. The LLM performs sentence-level entailment checking — it breaks the answer into individual sentences and labels each one assupported: true or supported: false relative to the document content. The final score is the fraction of supported sentences.
Threshold labels
Configuration
Result details
Semantic Drift
Semantic Drift measures how closely the response stays within the semantic envelope of the retrieved documents. A high-scoring response is topically close to the source material; a low-scoring one has drifted into territory not covered by the documents. How it works: TrustifAI splits each document into individual sentences and embeds them all. It then computes the cosine similarity between the answer embedding (and query embedding) and each sentence embedding, taking the best match. The final score reflects the peak semantic alignment between the answer and the document corpus.Semantic Drift uses embedding similarity, not LLM reasoning. It catches cases where the response is topically off-base, but will not catch factual errors within the same topic. Use it alongside Evidence Coverage for full hallucination coverage.
Threshold labels
Configuration
Result details
Epistemic Consistency
Epistemic Consistency measures how stable the LLM’s responses are when asked the same question multiple times with different random seeds. Hallucinated answers tend to vary wildly between runs — a model that is genuinely confident produces semantically similar responses even under stochastic conditions. How it works: TrustifAI generatesk additional responses to the same query at elevated temperatures (randomly sampled from [0.7, 0.8, 0.9, 1.0]). Each sample is embedded and compared to the original answer using cosine similarity. The final score is the mean cosine similarity across all samples — equivalent to 1 - σ in the semantic space.
k is controlled by k_samples in your config. Setting k_samples: 0 skips generation and assumes full consistency (score = 1.0), which is useful for cost-sensitive pipelines.
Threshold labels
Configuration
Result details
Source Diversity
Source Diversity measures whether the response draws on multiple independent sources or leans entirely on a single document. Answers synthesized from multiple distinct sources are generally more trustworthy than those derived from a single reference. How it works: TrustifAI resolves a unique source ID for each retrieved document — usingsource, file_path, or url metadata fields if present, or falling back to a SHA-256 content hash. It then counts the number of distinct source IDs and computes a normalized score that combines a diversity ratio with a count-based exponential reward:
0.8 rather than penalizing the response.
Threshold labels
Configuration
Result details
Accessing individual metric scores
get_trust_score() returns all active metric scores in the details dictionary. Access individual metric results from the return value: