Skip to main content

Overview

The TokenizationService class manages all tokenization operations in Tokenizador. It integrates with the tiktoken library to provide accurate token IDs and counts, with intelligent fallback mechanisms when tiktoken is unavailable.
This service supports 48 AI models from OpenAI, Anthropic, Google, Meta, and other providers.

Constructor

Creates a new TokenizationService instance and begins initialization.
Properties initialized:
  • encoder: null (set after initialization)
  • isInitialized: false (set to true when ready)
  • initPromise: Promise for initialization tracking
  • isRealTiktoken: Indicates if real tiktoken or fallback is being used

Methods

initializeTokenizer()

Initializes the tiktoken encoder asynchronously.
Promise<void>
Resolves when tokenizer is initialized (or fallback is ready)
Initialization process:
  1. Waits up to 10 seconds for tiktoken library to load
  2. Checks multiple locations: global context, window object
  3. Initializes cl100k_base encoding (GPT-4 compatible)
  4. Performs test tokenization to verify functionality
  5. Sets isRealTiktoken flag based on tiktoken availability
If tiktoken fails to load, the service automatically uses fallback tokenization. Token IDs will be marked as approximate.

waitForInitialization()

Waits for the tokenizer to complete initialization.
Promise<void>
Resolves when initialization is complete

tokenizeText()

Tokenizes text using the appropriate method for the specified model.
string
required
The text to tokenize
string
required
Model identifier (e.g., “gpt-4o”, “claude-3.5-sonnet”)
Promise<Object>
Object containing tokens array and count
Return value structure:
Array<Object>
Array of token objects with text, type, ID, and metadata
number
Total number of tokens
Token object structure:
string
The actual text of the token
string
Token type: “palabra”, “subword”, “number”, “punctuation”, “special”, “espacio_en_blanco”
string
Unique identifier for the token (e.g., “token_0”)
number
Numeric token ID from tiktoken (or approximation)
number
Zero-based position in the token sequence
boolean
True if token ID is approximate (fallback mode)
For models using cl100k_base encoding (GPT-4, Claude, etc.), token IDs are exact. For other models, counts are adjusted using model-specific ratios.

createTokensFromEncoding()

Creates visual token objects from tiktoken encoding.
string
required
Original input text
number[]
required
Array of token IDs from tiktoken.encode()
string
required
Model identifier
Array<Object>
Array of token objects for visualization
Process:
  1. Iterates through each encoded token ID
  2. Decodes individual tokens to get exact text
  3. Determines token type based on content
  4. Creates token object with metadata
  5. Marks tokens as approximate if using fallback

fallbackTokenization()

Provides tokenization when tiktoken is unavailable.
string
required
Text to tokenize
string
required
Model identifier
Object
Object with tokens array and count
Fallback strategy:
  • Splits text into words and whitespace segments
  • Uses heuristics to approximate token boundaries
  • Generates deterministic IDs based on content
  • Marks all tokens as isApproximate: true
Fallback tokenization provides approximate results. Token IDs will not match actual tiktoken IDs but counts are reasonably accurate.

splitWordIntoTokens()

Splits a word into smaller tokens simulating tiktoken behavior.
string
required
Word to split into tokens
number
required
Starting token index
Array<Object>
Array of token objects
Algorithm:
  • Words ≤3 characters: single token
  • Longer words: split based on ~2.8 characters per token ratio
  • First chunk marked as “palabra”, subsequent as “subword”
  • Dynamic chunk sizing based on remaining characters

determineTokenType()

Determines the type of a token based on its content.
string
required
Token text to classify
string
Token type: “number”, “punctuation”, “special”, or “palabra”
Classification rules:
Token contains only digits: ^\d+$

createDeterministicId()

Creates a deterministic numeric ID for fallback tokens.
string
required
Token text
number
required
Token index
number
Deterministic ID in range 10000-109999
Algorithm:
  1. Generates simple hash from character codes
  2. Combines with index for uniqueness
  3. Normalizes to 5-digit range (10000-109999)

getTokenizerName()

Returns a human-readable name for a tokenizer encoding.
string
required
Encoding identifier (e.g., “cl100k_base”)
string
Display name for the tokenizer

getAlgorithmName()

Returns a description of the tokenization algorithm for a model.
string
required
Model identifier
string
Algorithm description

Token Types

The service classifies tokens into these categories:

palabra

Standard word token

subword

Part of a longer word

palabra_con_espacio

Word with leading space

number

Numeric token

punctuation

Punctuation marks

special

Special characters

espacio_en_blanco

Whitespace

unknown

Decode failure

Usage Examples

Model Support

The service supports multiple encoding strategies:
OpenAI: GPT-4, GPT-4 Turbo, GPT-3.5 Turbo
Anthropic: Claude 3 Opus, Claude 3.5 Sonnet
Meta: Llama 3.1 (with ratio adjustment)
Uses exact tiktoken encoding with model-specific token ratios.
OpenAI: GPT-4o, GPT-4o MiniUses newer tokenizer with improved efficiency.
Google: Gemini (SentencePiece approximation)
Mistral: Mistral models (ratio-based approximation)
Cohere: Command models (ratio-based approximation)
Uses fallback with model-specific token ratios.

Error Handling

The service gracefully falls back to approximate tokenization if tiktoken fails to load. Your application continues working with slightly reduced accuracy.

See Also

TokenAnalyzer

Main application orchestrator

StatisticsCalculator

Calculate costs and statistics

Supported Models

View all 48 supported models

Architecture

Understand the system design