Real Benchmarks, Real Results
Peer-shaped studies on how AI performs on real water-treatment use cases. All research is open-access — no paywalls, no gated reports.
Featured report · v1.0
Operational RO diagnosis
Water Benchmarks Technical Committee · May 2026
We evaluate 26 model runs on 31 real RO-plant cases across 5 failure families, scored 0–12 against private expert golds with a critical safety gate.
In this report
- 1.Abstract
- 2.The benchmark
- 3.Methodology & rubric
- 4.Results
- 5.The max-trap
- 6.Calibration
- 7.Lessons for the water sector
- 8.Limitations
Headline findings
The max-trap
Reasoning effort buys risk, not safety — the same model triggers a critical fail only at its highest effort levels.
The effort-pricing trap
Maxed reasoning ties the same competence as reasoning off, at several times the cost and latency.
NOWE fragility
One failure family collapses to a 38% pass-rate and concentrates most of the critical fails.
Calibration is orthogonal
The best-calibrated models are not the highest-ranked. A high score doesn't mean trustworthy confidence.
All cases, rubric and scripts are public for reproducibility.