Every coined term in the Orthemology project — orthemma, ortheme, metaortheme, metaorthemma, and orthing — is a candidate, not an adopted term. Adoption requires the coined term to outperform established alternatives on a controlled benchmark under predeclared decision criteria. That benchmark has not been run. No term has been adopted. Every document in this project remains fully readable using plain-language substitutes in place of any coined term.Documentation Index
Fetch the complete documentation index at: https://mintlify.com/theislampill/orthemology/llms.txt
Use this file to discover all available pages before exploring further.
Why terminology is gated
The benchmark exists to avoid a specific failure mode: adopting terms for elegance or familiarity rather than demonstrable utility. “Can be said in ordinary words” does not settle the question — “Gettier case” is fully paraphrasable and permanently useful. The benchmark therefore measures behavioral and communicative deltas under matched conditions: does an auditor equipped with the coined vocabulary do measurably better than one using ordinary language expressing the same distinctions? Until a term clears the benchmark at adequate power, it is treated as a candidate label. Plain-language substitutes carry the same meaning and are always available:| Coined term | Plain-language substitute |
|---|---|
| orthemma | occurrence / case / version-identified instance |
| ortheme | state-type / label / classification category |
| metaortheme | governing rule / evidence standard / policy |
| metaorthemma | instantiated governing configuration / bound governing record |
| orthing | the classification episode / the auditing act |
orthable is excluded from the operational terminology core. It is exploratory in a companion lane only; no philosophical-comprehension module is currently justified, and it does not appear in any arm of the benchmark instrument.Pilot-0 v2 packet status
| Field | Value |
|---|---|
| Packet ID | TERM-P0-V2 |
| Readiness state | READY_FOR_HUMAN_MATCHING_REVIEW |
| Registration state | NOT_REGISTERED |
| Run exists | false |
| Licensed now | nothing — no utility result, no adoption, no retirement |
READY_FOR_HUMAN_MATCHING_REVIEW means the packet’s blind human matching review is the owner-gated prerequisite for advancing to READY_TO_RUN. The review must be recorded as passing before any execution can proceed. Pilot 0 is a feasibility study only — no adoption or retirement conclusion of any kind is available from it.Arm A — Exposure-matched filler (control)
Arm A — Exposure-matched filler (control)
Participants receive an exposure-matched filler primer (the same length as the active primers, with no construct teaching). This arm establishes the baseline against which the teaching effect of both ordinary-language distinctions (Arm B) and coined vocabulary (Arm C) is measured.Estimand E3:
P(probe pass | B) − P(probe pass | A) (teaching effect of distinctions alone).Arm B — Ordinary-vocabulary primer
Arm B — Ordinary-vocabulary primer
Participants receive a primer that teaches the operational distinctions — occurrence identity and version; plural profiles; candidate sets; route-sufficient vs identity-complete; evidence property, scope, and expiry; false closure; correct-by-luck and pathway adequacy; governing-rule revision — expressed entirely in ordinary words, with no coined vocabulary.Estimand E1:
P(probe pass | C) − P(probe pass | B) (vocabulary effect over and above the distinctions themselves).Arm C — Coined-vocabulary primer
Arm C — Coined-vocabulary primer
Participants receive a primer that teaches the same distinctions as Arm B, but using the coined terms: orthemma, ortheme, metaortheme, metaorthemma, and orthing, defined once in the primer. The primer length is matched to Arm B’s primer.A win for Arm C over Arm B is the necessary (but not sufficient) condition for adoption. The exact definition of “wins” is: Arm C is non-inferior to Arm B on every primary (−5 pp margin) and superior by the minimum important effect on at least one primary or on compression-at-equal-correctness, with no secondary regressing beyond the harm ceiling.
Arm C′ — Sham-vocabulary primer (label-specificity control)
Arm C′ — Sham-vocabulary primer (label-specificity control)
Participants receive a primer using machine-generated sham labels — a 1:1 lexical map of Arm C’s coined terms to invented non-words generated by
scripts/gen_sham_primer.py. This arm is mandatory in Pilot 0 and separates three distinct potential benefits: (1) the benefit of these specific terms, (2) the benefit of any labels at all, and (3) the benefit of the distinctions themselves.Estimand E2 (label-specificity): P(probe pass | C) − P(probe pass | C′), computed only on items flagged eligible_for_c_vs_cprime: true. The false-closure item and negative controls are excluded by flag.A sham comprehension gap of more than 10 percentage points between C and C′ indicates that label specificity matters — informing Pilot 1’s design.The benchmark instrument
The frozen instrument isterminology/pilot0-v2/items/ITEMS.json — 9 items (7 substantive, 2 negative controls), each rendered once per arm using a common scenario, a length-matched framing sentence, and an identical question stem. All construct teaching lives in the primers.
Items are byte-identical across arms for negative controls. Item eligibility for the C-vs-C′ contrast is declared per-item with the eligible_for_c_vs_cprime flag and enforced by scripts/audit_terminology_matching.py in CI.
Feasibility gates (Pilot 0 only)
Pilot 0’s decision rule is feasibility-only. The analysis scriptanalysis/analyze_pilot0_v2.py outputs exactly one of four verdicts, derived mechanically from the numeric feasibility gates:
| Gate | Threshold |
|---|---|
| Pairwise rater agreement | ≥ 0.70 in every (item, arm) cell |
| Sham comprehension gap |C − C′| | ≤ 10 pp |
| Per-item pooled pass rate | Within [0.10, 0.90] for at least 5 of 7 substantive items |
| Negative-control overhead | ≤ 30 tokens added vs Arm A, every arm |
ADVANCE_TO_PILOT1REVISE_AND_RETEST_INSTRUMENTDO_NOT_ADVANCE_THIS_ITEM_VERSION(this item/instrument version only — never a term retirement)INCONCLUSIVE
Scoring rubrics and adjudication
Rating is performed by ≥3 human raters. Each rater scores one (item, answer) pair against the item’sprobes and expected_key rubric (binary per probe plus overhead-token count for negative controls).
Blinding: raters see answers with arm labels stripped and coined/sham/ordinary construct nouns masked by neutral placeholders ([TERM-1], [TERM-2], etc.). The masking map is generated per answer and recorded.
Adjudication: probe-level disagreements are resolved by majority vote. Three-way splits escalate to a written adjudication note, which is appended to the deviation ledger.
Pilot 1 and the confirmatory study
Pilot 0 is a feasibility and instrumentation study. The program beyond it has two further stages, both currently atDRAFT readiness:
Pilot 1 (TERM-P1-TEMPLATE)
An efficacy pilot instantiated only after Pilot 0 feasibility outcomes are known. The Pilot 1 template specifies: an item-variant inventory (≥3 variants per family, including the metaorthemma/binding family), a held-out-domain plan, simulation-based power (synthetic placeholders, clearly marked), rater-assignment planning, and a mixed-model analysis specification. Status: DRAFT.
Confirmatory study (TERM-CONFIRMATORY-TEMPLATE)
The only stage that can adopt or retire a term. Instantiated only after Pilot 1. Contains a freeze checklist with unfillable-until-pilot slots, a three-outcome decision surface (adopt / do not adopt yet / retire), and a metric-freeze rule. External preregistration is required before any confirmatory run. Status: DRAFT.
Decision 0018 and registration
Decision 0018 defined the closed readiness vocabulary and established thatREADY_FOR_HUMAN_MATCHING_REVIEW applies specifically to packets that declare a required pre-run human review gate. The TERM-P0-V2 packet is the only current packet in that state; it cannot advance to READY_TO_RUN until the blind human matching review is recorded as passing.
Registration state for TERM-P0-V2 is NOT_REGISTERED. No document may describe this packet as “preregistered” without naming a real external registry record. External preregistration is required before any human run of the confirmatory study; it is an owner/external act not represented by any repository artifact.
Three-outcome decision rule
The benchmark applies a strict three-outcome rule at every stage where adoption or retirement is in scope (the confirmatory study only):| Outcome | Condition |
|---|---|
| ADOPT | Adequately powered win at the preregistered minimum important effect |
| DO NOT ADOPT YET | Insufficient evidence — including every underpowered null and every ordinary nonsignificant tie |
| RETIRE / REJECT | Only on an adequately powered equivalence or noninferiority result, or a harm-ceiling breach |