Skip to main content

Overview

The OcrProvider interface defines the contract that all OCR providers must implement to work with meikipop. This abstraction allows you to swap different OCR backends without modifying the core application logic. All interface definitions are located in src/ocr/interface.py.

OcrProvider abstract class

Your custom provider must inherit from OcrProvider and implement its abstract methods:
src/ocr/interface.py

Required properties

str
required
A unique, user-friendly string for your provider (e.g., "My Cool OCR"). This name appears in the settings and tray icon menus.

Required methods

method
required
The core method where all OCR processing happens.Parameters:
  • image (PIL.Image.Image): The screen region to scan
Returns:
  • List[Paragraph]: If OCR succeeds (return empty list [] if no text found)
  • None: If a critical error occurred
The scan method receives a PIL Image object and must return data in meikipop’s standard format. Your main task is converting your OCR engine’s output into this format.

Data models

Your scan method must return data using these three immutable dataclasses:

BoundingBox

Represents the location and size of text with normalized coordinates.
src/ocr/interface.py
float
required
Horizontal center position, normalized to 0.0-1.0 range (0.0 is left edge)
float
required
Vertical center position, normalized to 0.0-1.0 range (0.0 is top edge)
float
required
Width of the bounding box, normalized to 0.0-1.0 range
float
required
Height of the bounding box, normalized to 0.0-1.0 range
All coordinates and dimensions must be normalized to a 0.0-1.0 float range, relative to the input image’s dimensions. (0.0, 0.0) represents the top-left corner.

Converting pixel coordinates to normalized format

If your OCR engine returns absolute pixel coordinates, you need to convert them:

Word

Represents a single recognized text element.
src/ocr/interface.py
str
required
The recognized text. Can be a full word ("日本語") or a single character ("日"). Single-character boxes often lead to more precise lookups.
str
required
The character that follows the word. Usually an empty string "" for Japanese text.
BoundingBox
required
The bounding box for this specific word or character.
Meikipop’s hit-scanning works well with both word-level and character-level boxes. Providing single-character boxes often leads to more precise dictionary lookups.

Paragraph

Represents a block of text composed of words.
src/ocr/interface.py
str
required
The complete, reconstructed text of the paragraph.
List[Word]
required
A list of Word objects that form this paragraph.
BoundingBox
required
The bounding box encompassing the entire paragraph.
bool
required
Must be True if the text is written top-to-bottom (vertical Japanese text). If your OCR engine doesn’t provide this information, you can infer it from the bounding box aspect ratio: height > width.
For Japanese text, correctly setting is_vertical is crucial for proper text rendering and lookups.

Example implementation

Here’s how a typical scan method transforms OCR data:
src/ocr/providers/dummy/provider.py

Next steps

Create a custom provider

Step-by-step guide to building your own OCR provider

Available providers

Explore the built-in OCR providers in meikipop