Skip to main content
TrustifAI evaluates LLM and RAG responses across multiple independent dimensions and combines them into a single Trust Score — a number between 0 and 1 that answers the question: can you rely on this response? Rather than treating quality as a single axis, TrustifAI treats it as a multi-dimensional signal, which makes it far harder to game and far more informative when something goes wrong.

Multi-dimensional evaluation model

The Trust Score is built from up to five component metrics, computed in parallel:
  • 4 offline metrics — calculated after a response has already been generated, against the retrieved documents.
  • 1 online metric — calculated in real time during generation, using token log probabilities.
Each metric produces a score in the [0, 1] range. These component scores are then combined into the final Trust Score using a configurable weighted sum.

Weighted aggregation formula

The aggregation formula is a simple weighted linear combination:
Weights are read from your config_file.yaml and automatically normalized so they always sum to 1.0. This means you can freely adjust relative priorities without worrying about the math.

Default weights

These defaults reflect a deliberate priority ordering: factual grounding matters most, followed by topical alignment, then consistency across generations, and finally breadth of sourcing.

Configuring weights in YAML

Weights are proportional, not absolute. Setting evidence_coverage to 0.8 and semantic_drift to 0.2 (with others at 0) produces the same scores as 0.4 and 0.1. TrustifAI normalizes before computing.
Any metric with a weight: 0 (or omitted entirely) is automatically excluded from both the computation and the reasoning graph, so you can disable metrics without deleting their config block.

Decision labels

After computing the final score, TrustifAI maps it to one of three human-readable decision labels using configurable thresholds: These thresholds are set under metrics[type=trust_score] in your config:

Calling get_trust_score()

1

Build a MetricContext

Wrap your query, answer, and retrieved documents in a MetricContext object. TrustifAI accepts LangChain Document objects, LlamaIndex nodes, plain strings, dicts, or lists — it normalizes them automatically.
2

Initialize the engine

Point Trustifai at your config file. The engine reads weights, thresholds, and LLM/embedding settings from YAML.
3

Call get_trust_score()

Pass the context to get_trust_score(). All active metrics run in parallel and the results are aggregated automatically.

Interpreting the result dict

get_trust_score() returns a dictionary with four keys:
If no documents are provided, get_trust_score() immediately returns a score of 0.0 with label Unreliable without making any API calls. Always pass at least one retrieved document.

Async usage

For high-concurrency server deployments, use the native async variant to avoid blocking the event loop:
The async path runs metric calculations concurrently with asyncio.gather and uses an async embedding pipeline, making it significantly faster under load.