Skip to main content

Documentation Index

Fetch the complete documentation index at: https://mintlify.com/theislampill/orthemology/llms.txt

Use this file to discover all available pages before exploring further.

Every coined term in the Orthemology project — orthemma, ortheme, metaortheme, metaorthemma, and orthing — is a candidate, not an adopted term. Adoption requires the coined term to outperform established alternatives on a controlled benchmark under predeclared decision criteria. That benchmark has not been run. No term has been adopted. Every document in this project remains fully readable using plain-language substitutes in place of any coined term.
No coined term has been adopted. No benchmark run has occurred. The pilot-0 v2 packet (TERM-P0-V2) status is READY_FOR_HUMAN_MATCHING_REVIEW (Decision 0018); NOT RUN; NO TERM ADOPTED. No utility result, adoption decision, or term retirement exists. Canonical status: experiments/experiment-status.yaml (packet TERM-P0-V2).

Why terminology is gated

The benchmark exists to avoid a specific failure mode: adopting terms for elegance or familiarity rather than demonstrable utility. “Can be said in ordinary words” does not settle the question — “Gettier case” is fully paraphrasable and permanently useful. The benchmark therefore measures behavioral and communicative deltas under matched conditions: does an auditor equipped with the coined vocabulary do measurably better than one using ordinary language expressing the same distinctions? Until a term clears the benchmark at adequate power, it is treated as a candidate label. Plain-language substitutes carry the same meaning and are always available:
Coined termPlain-language substitute
orthemmaoccurrence / case / version-identified instance
orthemestate-type / label / classification category
metaorthemegoverning rule / evidence standard / policy
metaorthemmainstantiated governing configuration / bound governing record
orthingthe classification episode / the auditing act
orthable is excluded from the operational terminology core. It is exploratory in a companion lane only; no philosophical-comprehension module is currently justified, and it does not appear in any arm of the benchmark instrument.

Pilot-0 v2 packet status

FieldValue
Packet IDTERM-P0-V2
Readiness stateREADY_FOR_HUMAN_MATCHING_REVIEW
Registration stateNOT_REGISTERED
Run existsfalse
Licensed nownothing — no utility result, no adoption, no retirement
READY_FOR_HUMAN_MATCHING_REVIEW means the packet’s blind human matching review is the owner-gated prerequisite for advancing to READY_TO_RUN. The review must be recorded as passing before any execution can proceed. Pilot 0 is a feasibility study only — no adoption or retirement conclusion of any kind is available from it.
The four arms of pilot-0 v2 are described below.
Participants receive an exposure-matched filler primer (the same length as the active primers, with no construct teaching). This arm establishes the baseline against which the teaching effect of both ordinary-language distinctions (Arm B) and coined vocabulary (Arm C) is measured.Estimand E3: P(probe pass | B) − P(probe pass | A) (teaching effect of distinctions alone).
Participants receive a primer that teaches the operational distinctions — occurrence identity and version; plural profiles; candidate sets; route-sufficient vs identity-complete; evidence property, scope, and expiry; false closure; correct-by-luck and pathway adequacy; governing-rule revision — expressed entirely in ordinary words, with no coined vocabulary.Estimand E1: P(probe pass | C) − P(probe pass | B) (vocabulary effect over and above the distinctions themselves).
Participants receive a primer that teaches the same distinctions as Arm B, but using the coined terms: orthemma, ortheme, metaortheme, metaorthemma, and orthing, defined once in the primer. The primer length is matched to Arm B’s primer.A win for Arm C over Arm B is the necessary (but not sufficient) condition for adoption. The exact definition of “wins” is: Arm C is non-inferior to Arm B on every primary (−5 pp margin) and superior by the minimum important effect on at least one primary or on compression-at-equal-correctness, with no secondary regressing beyond the harm ceiling.
Participants receive a primer using machine-generated sham labels — a 1:1 lexical map of Arm C’s coined terms to invented non-words generated by scripts/gen_sham_primer.py. This arm is mandatory in Pilot 0 and separates three distinct potential benefits: (1) the benefit of these specific terms, (2) the benefit of any labels at all, and (3) the benefit of the distinctions themselves.Estimand E2 (label-specificity): P(probe pass | C) − P(probe pass | C′), computed only on items flagged eligible_for_c_vs_cprime: true. The false-closure item and negative controls are excluded by flag.A sham comprehension gap of more than 10 percentage points between C and C′ indicates that label specificity matters — informing Pilot 1’s design.

The benchmark instrument

The frozen instrument is terminology/pilot0-v2/items/ITEMS.json — 9 items (7 substantive, 2 negative controls), each rendered once per arm using a common scenario, a length-matched framing sentence, and an identical question stem. All construct teaching lives in the primers. Items are byte-identical across arms for negative controls. Item eligibility for the C-vs-C′ contrast is declared per-item with the eligible_for_c_vs_cprime flag and enforced by scripts/audit_terminology_matching.py in CI.
// Item structure (from items/ITEMS.json — each item)
{
  "item_id": "string",
  "scenario_common": "string",
  "probes": ["string"],
  "expected_key": {},
  "eligible_for_c_vs_cprime": true,
  "negative_control": false
}

Feasibility gates (Pilot 0 only)

Pilot 0’s decision rule is feasibility-only. The analysis script analysis/analyze_pilot0_v2.py outputs exactly one of four verdicts, derived mechanically from the numeric feasibility gates:
GateThreshold
Pairwise rater agreement≥ 0.70 in every (item, arm) cell
Sham comprehension gap |C − C′|≤ 10 pp
Per-item pooled pass rateWithin [0.10, 0.90] for at least 5 of 7 substantive items
Negative-control overhead≤ 30 tokens added vs Arm A, every arm
The four feasibility verdicts are:
  • ADVANCE_TO_PILOT1
  • REVISE_AND_RETEST_INSTRUMENT
  • DO_NOT_ADVANCE_THIS_ITEM_VERSION (this item/instrument version only — never a term retirement)
  • INCONCLUSIVE
No adoption or retirement conclusion is available from Pilot 0. Estimands E1–E3 are computed as descriptive values only, to inform Pilot 1’s design. Efficacy margins, minimum important effects, noninferiority margins, and harm thresholds are not Pilot 0 parameters — they bind Pilot 1 and the confirmatory stage.

Scoring rubrics and adjudication

Rating is performed by ≥3 human raters. Each rater scores one (item, answer) pair against the item’s probes and expected_key rubric (binary per probe plus overhead-token count for negative controls). Blinding: raters see answers with arm labels stripped and coined/sham/ordinary construct nouns masked by neutral placeholders ([TERM-1], [TERM-2], etc.). The masking map is generated per answer and recorded. Adjudication: probe-level disagreements are resolved by majority vote. Three-way splits escalate to a written adjudication note, which is appended to the deviation ledger.
A rater who has seen the Arm C primer is contaminated for later Arm A or B judgments. The protocol uses between-subject assignment for the vocabulary-exposure comparison, or a first-exposure-only primary analysis, to prevent carryover. Within-session drift (early vs late items) must be reported.

Pilot 1 and the confirmatory study

Pilot 0 is a feasibility and instrumentation study. The program beyond it has two further stages, both currently at DRAFT readiness:

Pilot 1 (TERM-P1-TEMPLATE)

An efficacy pilot instantiated only after Pilot 0 feasibility outcomes are known. The Pilot 1 template specifies: an item-variant inventory (≥3 variants per family, including the metaorthemma/binding family), a held-out-domain plan, simulation-based power (synthetic placeholders, clearly marked), rater-assignment planning, and a mixed-model analysis specification. Status: DRAFT.

Confirmatory study (TERM-CONFIRMATORY-TEMPLATE)

The only stage that can adopt or retire a term. Instantiated only after Pilot 1. Contains a freeze checklist with unfillable-until-pilot slots, a three-outcome decision surface (adopt / do not adopt yet / retire), and a metric-freeze rule. External preregistration is required before any confirmatory run. Status: DRAFT.

Decision 0018 and registration

Decision 0018 defined the closed readiness vocabulary and established that READY_FOR_HUMAN_MATCHING_REVIEW applies specifically to packets that declare a required pre-run human review gate. The TERM-P0-V2 packet is the only current packet in that state; it cannot advance to READY_TO_RUN until the blind human matching review is recorded as passing. Registration state for TERM-P0-V2 is NOT_REGISTERED. No document may describe this packet as “preregistered” without naming a real external registry record. External preregistration is required before any human run of the confirmatory study; it is an owner/external act not represented by any repository artifact.
# From experiments/experiment-status.yaml
- packet_id: TERM-P0-V2
  readiness_state: READY_FOR_HUMAN_MATCHING_REVIEW
  registration_state: NOT_REGISTERED
  run_exists: false
  licensed: nothing — Pilot 0 is feasibility-only; no adoption or retirement
             decision is available from it
  required_remaining_gates:
    - owner-gated blind human matching review (recorded pass promotes to READY_TO_RUN)
    - owner execution authorization
    - external preregistration (owner/external act)

Three-outcome decision rule

The benchmark applies a strict three-outcome rule at every stage where adoption or retirement is in scope (the confirmatory study only):
OutcomeCondition
ADOPTAdequately powered win at the preregistered minimum important effect
DO NOT ADOPT YETInsufficient evidence — including every underpowered null and every ordinary nonsignificant tie
RETIRE / REJECTOnly on an adequately powered equivalence or noninferiority result, or a harm-ceiling breach
A term is never declared useless on an underpowered miss. An ordinary nonsignificant tie means “do not adopt yet” — not retirement.

Build docs developers (and LLMs) love