Skip to main content

Overview

The models configuration module (models-config.js) contains comprehensive data about AI language models, their token encodings, pricing, context limits, and company branding information. This configuration powers the model selector and cost calculations throughout the application.

Configuration Objects

MODEL_ENCODINGS

A mapping of model identifiers to their tokenization encoding schemes. Most models use cl100k_base or o200k_base encodings.
object
required
Object mapping model IDs to encoding identifiers
Models without native encoding information use cl100k_base as an approximation, marked with comments in the source code.

Supported Encodings

OpenAI’s latest encoding (2024+)Used by:
  • GPT-4o
  • GPT-4o Mini
More efficient token usage than cl100k_base.

COMPANIES

Company branding information including colors and emoji logos for visual representation in the UI.
object
required
Object mapping company names to branding data

Supported Companies

  • OpenAI - #00a67e 🤖
  • Anthropic - #d97757 🧠
  • Mistral AI - #ff6b35 💨
  • Cohere - #39a0ed 🔗
  • DeepSeek - #2c5aa0 🔍
  • 01.AI - #1a73e8 🤖
  • AI21 Labs - #6c5ce7 🧪
  • xAI - #000000 ❌
  • Google - #4285f4 🔍
  • Meta - #1877f2 📘
  • Microsoft - #00bcf2 💻
  • Amazon - #ff9900 📦
  • NVIDIA - #76b900 💚
  • Alibaba - #ff6a00 🛒
  • Reka - #ff4757 🦄
  • Perplexity - #20bf6b ❓
  • IBM - #054ada 💼
  • Nous Research - #8e44ad 🔬
  • Snowflake - #29b5e8 ❄️

MODELS_DATA

Complete configuration data for all supported AI models including pricing, context limits, and technical specifications.
object
required
Object mapping model IDs to complete model configuration

Model Data Examples

OpenAI’s most capable multimodal model with 128K context window and efficient o200k_base encoding.

Usage Examples

Retrieving Model Configuration

Calculating Token Costs

Building Model Selector UI

Validating Context Length

Comparing Model Costs

Model Categories

The configuration includes 48 models across multiple categories:

OpenAI Models

5 models including GPT-4o, GPT-4 Turbo, and GPT-3.5 Turbo

Anthropic Models

4 Claude models from Haiku to Opus

Google Models

2 Gemini 1.5 models with massive context windows

Open Source Models

37 models from Meta, Mistral, Alibaba, and others

Token Ratio Explained

The tokenRatio field adjusts for differences in how models count tokens:
OpenAI models and most approximations use 1.0 as the baseline.Models: GPT-4o, GPT-4, GPT-4 Turbo, GPT-3.5 Turbo
Models that typically count more tokens for the same text.Examples:
  • Claude models: 1.1 (10% more tokens)
  • Gemini models: 1.05 (5% more tokens)
  • Amazon Titan: 1.04
  • Snowflake Arctic: 1.06
Models that typically count fewer tokens for the same text.Examples:
  • Llama models: 0.95 (5% fewer tokens)
  • Alibaba Qwen: 0.92 (8% fewer tokens)
  • DeepSeek: 0.93
  • AI21 Jamba: 0.94
Token ratios are approximations based on empirical testing. Actual token counts may vary depending on text characteristics.

Best Practices

1

Always Check Model Availability

2

Apply Token Ratio for Accurate Estimates

3

Consider Context Limits

Check that your content fits within the model’s context window before making API calls.
4

Use Company Branding Consistently

Always reference the COMPANIES object for visual consistency across the UI.

Tokenization Service

Learn how tokenization works with these encodings

Statistics Calculator

Implementation details for cost calculations

UI Controller

UI component that uses this configuration

Understanding Tokenization

Deep dive into tokenization encodings