Docs
Beginner's guide
One diagram to see what MEI rates, how we rate it, and how to read a conclusion. World models only. No cross-domain leaderboard.
World-model rating system at a glance
Flow from rated object to dual-axis conclusion, seven dimensions, three method layers, and run conditions.
01
What we rate
Model versionuse-domain
02
Dual-axis call
Levelevidence strength
03
Seven dims
Shown separately
Three method layers
L1
Replication audit
Independently re-run public benchmarks; publish values and conditions.
L2
Independent probes
Fill gaps on uncertainty calibration and failure boundaries.
L3
Rating judgment
Level × strength by use-domain, with a reasoning chain.
04
Run conditions are evidence
N · precision · GPU · subset / offload — on the same card
The rated object is always world-model version × use-domain. Domains lack a reliable common order, so we never publish a total ranking.
Conclusions use two axes: level (1 / 2 / 3 / NR) for maturity, evidence strength (high / medium / low) for how strong the evidence is. They must appear as a pair.
Seven dimensions stay separate—never one total score. Run conditions (N, precision, GPU, subset, …) ship on the same card. Beyond “does it look good,” we ask “dare you trust it.”