Skip to main content

Overview

Meikipop includes several built-in OCR providers optimized for Japanese text recognition. Each provider offers different trade-offs between accuracy, speed, cost, and resource requirements.

Provider comparison

Built-in providers

Dummy OCR

The dummy provider is designed as a template for creating custom providers. It returns fixed mock data for testing.
src/ocr/providers/dummy/provider.py
Implementation highlights:
  • Returns hardcoded Japanese text with both horizontal and vertical examples
  • Demonstrates proper coordinate normalization
  • Shows character-level and word-level Word objects
  • Fully commented for educational purposes
Use cases:
  • Developing and testing UI without a real OCR backend
  • Template for creating custom providers
  • Understanding the data transformation process
Example output:

meikiocr (local)

The meikiocr provider uses a high-performance local model specifically optimized for Japanese video game text.
src/ocr/providers/meikiocr/provider.py
Implementation highlights:
  • Uses the meikiocr Python library
  • Converts PIL images to NumPy RGB arrays
  • Returns character-level boxes for precise lookups
  • Groups individual lines into paragraphs using postprocessing
  • Filters out non-Japanese text
Configuration:
Processing pipeline:
1

Initialize

Creates a MeikiOCR client that handles model downloading and session management internally.
2

Convert image

Converts PIL Image to NumPy RGB array for library compatibility.
3

Run OCR

Calls run_ocr() with confidence thresholds to get character-level results.
4

Transform results

Converts [x1, y1, x2, y2] pixel coordinates to normalized BoundingBox objects.
5

Group paragraphs

Uses group_lines_into_paragraphs() to combine related lines.
Key methods:
Requirements:
  • Install: pip install meikiocr
  • GPU recommended for best performance
  • Models downloaded automatically on first run

Google Lens v2 (remote)

This provider sends screenshots to Google’s servers. Do not use with sensitive or private information.
src/ocr/providers/glensv2/provider.py
Implementation highlights:
  • Uses Google Lens API via protobuf protocol
  • Maintains persistent HTTP session for performance
  • Supports low-bandwidth mode (50% resolution, 16-color quantization)
  • Returns normalized coordinates directly (no conversion needed)
  • Filters for Japanese text using regex
Image processing:
Text direction detection:
Requirements:
  • Active internet connection
  • Accepts Google’s data processing terms
Performance:
  • Network latency: ~200-500ms typical
  • Request timeout: 10 seconds
  • Logs detailed timing information

owocr (WebSocket)

The owocr provider connects to a running owocr daemon via WebSocket, allowing flexible deployment options.
src/ocr/providers/owocr/provider.py
Implementation highlights:
  • Maintains persistent WebSocket connection
  • Automatic reconnection on connection loss
  • Uses direct IP (127.0.0.1) to avoid localhost resolution delays
  • Two-part response protocol (acknowledgment + JSON results)
  • Returns normalized coordinates directly
Connection handling:
Communication protocol:
1

Send image

Converts PIL Image to BMP format and sends as binary.
2

Receive acknowledgment

Waits for “True” confirmation (5 second timeout).
3

Receive results

Waits for JSON response with OCR results (30 second timeout).
4

Transform data

Converts owocr’s format to meikipop’s Paragraph objects.
Retry logic:
Requirements:
  • Running owocr daemon
  • Command: owocr -r websocket -w websocket -of json -e glens
  • WebSocket connection to localhost:7331

Chrome Screen AI (local)

This provider uses Chrome’s Screen AI component for local, offline OCR processing.
src/ocr/providers/screenai/provider.py
Implementation highlights:
  • Uses Chrome’s native Screen AI library via ctypes
  • Singleton pattern for library initialization (once per app lifetime)
  • Suppresses verbose native library output
  • Returns character-level (symbol) boxes
  • Automatically downsizes large images (>4MP)
Library initialization:
Image preparation:
Output suppression:
Text direction detection:
Requirements:
  • Download Screen AI components from: https://chrome-infra-packages.appspot.com/p/chromium/third_party/screen-ai
  • Extract to: ~/.config/screen_ai/resources/
  • Platform: Windows (DLL) or Linux (SO)

Common patterns

Postprocessing: Grouping lines into paragraphs

Most providers use the shared group_lines_into_paragraphs() utility:
This function:
  • Combines adjacent lines into logical paragraphs
  • Respects text direction (vertical vs. horizontal)
  • Improves text readability and context

Japanese text filtering

Several providers filter for Japanese text:
Character ranges:
  • \u3040-\u309F: Hiragana
  • \u30A0-\u30FF: Katakana
  • \u4E00-\u9FAF: Kanji

Selecting a provider

Choose based on your requirements: For offline gaming:
For maximum accuracy:
For development:
For custom deployment:
For Chrome integration:

Next steps

Create custom provider

Build your own OCR provider using these as examples

OCR provider interface

Understand the interface contract and data models