若水研究院

Docs

Reading a model card

A Model Card is the smallest readable rating unit: one card = one world-model version × use-domain, always paired with evidence strength and run conditions.

Model card anatomy

Callouts for use-domain, level, evidence strength, seven dimensions, and run conditions.

  1. 01

    Use-domain

    The conclusion holds only inside this domain.

  2. 02

    Level

    1 / 2 / 3 / NR: maturity, not a leaderboard place.

  3. 03

    Evidence strength

    High / medium / low: how strong the evidence is—paired with level.

  4. 04

    Seven scores

    Shown separately; never collapsed to one total.

  5. 05

    Run conditions

    N, precision, GPU, …—the trust boundary on the card.

Schematic structure only—not a real rating. On production cards, level and evidence strength always share the screen.
01

Start with use-domain and model version—without them, numbers mean nothing. Then check that level and evidence strength appear together.

02

The seven-dimension radar or table is detail, not a total. Deployers should watch uncertainty calibration and failure boundary closely.

03

Run conditions answer under which hardware, precision, and sample regime the number was obtained. With a subset or N=1, do not over-read strength.