若水研究院

若水研究院Model Evidence Institute

上善若水,循证而衡。

Safer, auditable world models—for the people who deploy them.

An independent rating institute for world models only: replication audits of public scores, and deployment-risk judgment.

What we do

Not another world-model leaderboard. An auditable evidence layer for deployers.

Three layers stack: replication audits earn trust, deployment risk is the moat, longitudinal tracking compounds over time.

01

Replication audit

Independently re-run public benchmarks; publish our measured values, run conditions, and variance—we invent no metrics, we run and record honestly.

02

Deployment-risk judgment

Failure boundaries, uncertainty calibration, long-horizon drift. Buyers care less about “does it look good” than “when will it deceive me.”

03

Longitudinal tracking

A version × date evidence archive. After updates, prior conclusions are labeled expired—an asset others cannot catch up to a year later.

Value

Who the evidence is written for, and which decisions it changes.

Buyers put world models into cars, robots, and factories. Others score how good generation looks; we ask whether you dare trust it.

For industry

Selection is not chasing a board. It is a defensible adoption decision under safety constraints—where failure boundaries, uncertainty, and run conditions must share one screen.

  • Absolute use-domain ratings: every conclusion binds to world-model version × use-domain. No cross-domain total ranking to hide trade-offs.
  • Deployment-risk language: uncertainty calibration and failure boundaries appear as a pair—citable in PoC, compliance, certification, and internal review.
  • User-pays independence: revenue comes only from users. We never sell ratings to the vendors we evaluate—trust is institutional, not rhetorical.

For academia

World-model evaluation is already crowded. The contribution is not another board—it is making scores independently verifiable, citable, and open to further pressure from follow-on work.

  • Replication-audit protocol: independently re-run public benchmarks and publish our measured values, run conditions, and variance—conclusions stay traceable.
  • Evidence strength × maturity level: separate how strong the evidence is from what it claims—one language for research and engineering.
  • Close the safety gap: beyond generation quality, systematically ask about physical consistency, long-horizon drift, and when the model is not to be trusted.

How to read

Understand the system first, then enter the evidence.

A rating is not a place on a board. Start with one diagram, learn the model card and methodology—then every issue can stand.

Now

Evidence being published

Latest issue · Model of the Week

#1 · Feature

Cosmos Predict-1: forward boundaries and deployment risk for driving world models

NVIDIA Cosmos Predict-1·Prediction: driving & robotics

Tech (architecture & capability) → application (what works, where it breaks) → business. Includes a prediction-domain rating card with run conditions and evidence strength on-screen.

Free to read until September 11, 2026

Level 2Demonstrated, bounded

2026-08-12

Recent ratings

Browse ratings

Each card is an absolute rating for one world-model version × use-domain, paired with evidence strength and on-card run conditions—from different domains, not a cross-domain leaderboard.

Loading ratings…

Trust

A value claim must be institutionally checkable.

Credibility rests on open, reproducible, time-bound ratings—run conditions are evidence.

Open & reproducible

Methodology and rating conclusions stay permanently free. N, precision, GPU, and subset status ship on the same card for anyone to verify.

Read the methodology

No universal leaderboard

Ratings ship as model × use-domain. We never publish a cross-domain total ranking—industry trade-offs and academic comparisons both need the dimensions kept.

Browse the ratings library

Ratings expire

Every rating binds to a model version and test date; after updates, prior conclusions are labeled expired so stale evidence is never treated as present truth.

Read the rating constitution

Vision

Safer, auditable world models entering the real world.

We do not build world models, and we do not compete with those we rate. We aim to be neutral evaluation infrastructure for the world-model era—so deployers in cars, robots, and factories hold evidence that stays verifiable, time-bound, and correctable, instead of unauditable vendor self-reports.

Evaluation as infrastructure

Replication pipelines, evidence archives, and use-domain protocols become reusable assets—so each world-model rating settles into a citable public record.

Independence as credibility

Interest order is constitutional: public interest first; information walls between ratings and commerce; when principle conflicts with revenue, drop the revenue.

Openness as defense

Methodology and rating conclusions stay permanently free; corrections have deadlines; transparency reports ship even at zero revenue—authority comes from verifiability, not volume.

News

Founder

Naren Bao

Project Assistant Professor at The University of Tokyo, working on driving perception, world models, and safety-critical vision–language systems. Founder & CTO of AquaAge. She started Model Evidence Institute.

Personal site

Join

Deliver safer world-model evidence to the decision room.

Follow Model of the Week: one world model per issue—tech → application → business, ending with a rating card. Biweekly features plus mid-cycle replication briefs; membership unlocks the archive, deep reviews, and specials.

¥500 / month (tax included). Stripe checkout lands in a later milestone.

Free email subscription

Delivered via Mailchimp. Issue updates only. Unsubscribe anytime.