New·See the latest water task rankings
Water Benchmarks

Real Benchmarks, Real Results

Peer-shaped studies on how AI performs on real water-treatment use cases. All research is open-access — no paywalls, no gated reports.

Featured report · v1.0

Operational RO diagnosis

Water Benchmarks Technical Committee · May 2026

We evaluate 26 model runs on 31 real RO-plant cases across 5 failure families, scored 0–12 against private expert golds with a critical safety gate.

In this report

  1. 1.Abstract
  2. 2.The benchmark
  3. 3.Methodology & rubric
  4. 4.Results
  5. 5.The max-trap
  6. 6.Calibration
  7. 7.Lessons for the water sector
  8. 8.Limitations

Headline findings

The max-trap

Reasoning effort buys risk, not safety — the same model triggers a critical fail only at its highest effort levels.

The effort-pricing trap

Maxed reasoning ties the same competence as reasoning off, at several times the cost and latency.

NOWE fragility

One failure family collapses to a 38% pass-rate and concentrates most of the critical fails.

Calibration is orthogonal

The best-calibrated models are not the highest-ranked. A high score doesn't mean trustworthy confidence.

All cases, rubric and scripts are public for reproducibility.