Skip to main content

Overview

The StatisticsCalculator class provides comprehensive statistical analysis for tokenized text. It calculates token counts, character counts, word counts, cost estimates, context utilization, and provides model comparison capabilities.
This calculator works with data from 48 AI models and provides accurate cost estimates based on current pricing.

Constructor

Creates a new StatisticsCalculator instance.
The calculator is stateless and can be reused for multiple calculations.

Methods

calculateStatistics()

Calculates comprehensive statistics for the given text and model.
string
required
The original input text
Object
required
Result object from TokenizationService.tokenizeText()
string
required
Model identifier (e.g., “gpt-4o”, “claude-3.5-sonnet”)
Object
Comprehensive statistics object
Return value structure:
number
Total number of tokens
number
Total number of characters
number
Total number of words
number
Estimated cost in USD for input tokens
number
Percentage of context window used (0-100)
number
Average tokens per word ratio
number
Cost per 1M input tokens in USD
number
Cost per 1M output tokens in USD

countWords()

Counts words in text using intelligent word boundary detection.
string
required
Text to analyze
number
Number of words (0 for empty text)
Algorithm:
  1. Trims whitespace from text
  2. Splits on whitespace characters (\s+)
  3. Filters out empty strings
  4. Returns count

calculateCost()

Calculates estimated cost based on token count and model pricing.
number
required
Number of tokens
Object
required
Model information object from MODELS_DATA
number
Estimated cost in USD
Cost calculation formula:
Cost estimates are based on input token pricing. Output tokens typically cost more.

calculateContextUtilization()

Calculates the percentage of the model’s context window being used.
number
required
Number of tokens in the text
number
required
Maximum context window size for the model
number
Percentage from 0 to 100 (capped at 100)

exceedsContextLimit()

Checks if token count exceeds the model’s context limit.
number
required
Number of tokens
string
required
Model identifier
boolean
True if exceeds limit, false otherwise

getContextWarning()

Returns a warning message if context usage is high or exceeded.
number
required
Number of tokens
string
required
Model identifier
string|null
Warning message or null if no warning needed
Warning thresholds:
Text exceeds the model’s maximum context window.

formatStatistics()

Formats statistics for display with proper localization and units.
Object
required
Raw statistics object from calculateStatistics()
Object
Formatted statistics with string values
Use formatted statistics for displaying in UI. They include proper thousand separators, currency symbols, and percentage signs.

compareModels()

Compares tokenization statistics across multiple models.
string
required
Text to analyze
string[]
required
Array of model IDs to compare
TokenizationService
required
Tokenization service instance
Promise<Array>
Array of comparison objects sorted by cost (cheapest first)
Comparison object structure:
string
Model identifier
string
Model provider (e.g., “OpenAI”, “Anthropic”)
Object
Raw statistics object
Object
Formatted statistics for display
Comparison results are automatically sorted by cost estimate, making it easy to find the most economical model for your text.

getEfficiencyMetrics()

Calculates efficiency metrics for tokenization analysis.
Object
required
Statistics object from calculateStatistics()
Object
Efficiency metrics object
Return value structure:
number
Cost per thousand tokens (lower is better)
number
Tokens per character (lower = better compression)
number
Tokens per word (lower = more efficient encoding)

Usage Examples

Statistics Interpretation

The total number of tokens the text is divided into. This directly impacts:
  • API costs (priced per token)
  • Processing time
  • Context window usage
Typical ranges:
  • Short prompt: 10-100 tokens
  • Medium text: 100-1,000 tokens
  • Long document: 1,000-10,000+ tokens
Total number of characters including spaces and punctuation.Rule of thumb: English text averages ~4 characters per token.
Number of words (whitespace-separated).Rule of thumb: English text averages ~0.75 tokens per word.
Estimated API cost for processing the text.Note: Based on input pricing. Output tokens cost more.Cost ranges (GPT-4o):
  • 1K tokens: ~$0.0025
  • 10K tokens: ~$0.025
  • 100K tokens: ~$0.25
Percentage of the model’s context window being used.Guidelines:
  • Less than 50%: Comfortable usage
  • 50-75%: Moderate usage
  • 75-90%: High usage
  • 90-100%: Near limit
  • Greater than 100%: Exceeds limit (will fail)
Average number of tokens per word.Typical values:
  • English: 1.3-1.5
  • Code: 1.5-2.0
  • Non-English: varies by language
Lower values indicate more efficient tokenization.

Cost Optimization Tips

Choose Efficient Models

Compare models to find the best token-to-cost ratio for your use case

Minimize Prompt Length

Remove unnecessary context and instructions to reduce token count

Use Smaller Models

Consider mini variants (e.g., gpt-4o-mini) for simpler tasks

Batch Requests

Process multiple items in one request to reduce per-request overhead

See Also

TokenAnalyzer

Main application orchestrator

TokenizationService

Tokenization engine

UIController

UI management

Supported Models

View all model pricing