若水研究院

Docs

Beginner's guide

One diagram to see what MEI rates, how we rate it, and how to read a conclusion. World models only. No cross-domain leaderboard.

World-model rating system at a glance

Flow from rated object to dual-axis conclusion, seven dimensions, three method layers, and run conditions.

  1. 01

    What we rate

    Model versionuse-domain

  2. 02

    Dual-axis call

    Levelevidence strength

  3. 03

    Seven dims

    Shown separately

Three method layers

  • L1

    Replication audit

    Independently re-run public benchmarks; publish values and conditions.

  • L2

    Independent probes

    Fill gaps on uncertainty calibration and failure boundaries.

  • L3

    Rating judgment

    Level × strength by use-domain, with a reasoning chain.

Read left to right, then down: version × domain → level × strength → seven dimensions → L1–L3 sources → run conditions as the trust boundary.
01

The rated object is always world-model version × use-domain. Domains lack a reliable common order, so we never publish a total ranking.

02

Conclusions use two axes: level (1 / 2 / 3 / NR) for maturity, evidence strength (high / medium / low) for how strong the evidence is. They must appear as a pair.

03

Seven dimensions stay separate—never one total score. Run conditions (N, precision, GPU, subset, …) ship on the same card. Beyond “does it look good,” we ask “dare you trust it.”