Skip to main content
Online metrics are computed during LLM generation, not after. Rather than evaluating a finished response against retrieved documents, they tap into the generation process itself — specifically the log probabilities that the LLM assigns to each output token. This makes them fast (no extra API calls) and complementary to offline metrics, giving you a real-time view into how certain the model was while producing its response.

The Confidence Score

The Confidence Score is TrustifAI’s single online metric. It quantifies how sure the LLM was about its own output by analyzing the probability distribution across generated tokens. How it works: TrustifAI captures the per-token log probability (logprob) values returned by the LLM alongside the generated text. From these it computes:
  1. Geometric mean probability — exp(mean(logprobs)) gives the normalized per-token probability for the entire sequence. This captures the model’s average certainty across all tokens.
  2. Variance penalty — exp(-variance(logprobs)) penalizes sequences where the model oscillated between high- and low-confidence tokens. Consistent uncertainty (uniform logprobs) is rated less harshly than erratic uncertainty.
The final score combines both:
The result is a value in [0, 1] where higher means the model was more consistently certain about its output.
The Confidence Score is only available for LLMs that expose token log probabilities. This includes OpenAI models (e.g., gpt-4o, gpt-4, gpt-4-turbo) when logprobs: true is set in your config. Models served through providers that strip logprob data will return score: 0.0 with label N/A.

Threshold labels

These thresholds have defaults of 0.90 and 0.70. You can override them by adding them to any metric’s params section in your YAML — for example, alongside the trust_score thresholds:

Using generate() to get Confidence Scores

The generate() method wraps an LLM call and automatically computes the Confidence Score from the returned logprobs.
1

Initialize the engine

The engine reads your LLM config, including the model name and logprob settings, from the YAML file.
2

Call generate()

Pass your prompt (and an optional system prompt). TrustifAI automatically requests logprobs from the LLM.
3

Read the response and confidence metadata

The return value contains both the generated text and the full confidence breakdown.

Interpreting the result dict

generate() returns a dictionary with two top-level keys:

Enabling logprobs in config

Make sure your LLM config requests logprobs. TrustifAI sets this automatically when generate() is called, but having it in the config ensures consistency:
The Confidence Score is only as reliable as the LLM’s calibration. A well-calibrated model assigns high log probabilities to tokens it is genuinely likely to get right, and lower log probabilities when it is uncertain. Many models — especially smaller, fine-tuned, or instruction-tuned ones — are poorly calibrated and may express high confidence even when hallucinating. Treat the Confidence Score as a useful signal, not a guarantee, and always combine it with offline metrics for full trustworthiness evaluation.

Combining online and offline metrics

The Confidence Score operates independently of the offline metrics. A typical production workflow combines both:
This two-step pattern gives you the fullest picture: online confidence tells you how certain the model was during generation, and the offline Trust Score tells you how well the finished response holds up against your retrieved documents.