Creating Benchmarks and Standards for AI in Water
Independent, open-access benchmarks for AI in water treatment — built with the sector.
Why Water Benchmarks
Built different, on purpose
Published benchmarks, peer-reviewed methodology, and structured evaluation criteria developed by the people who operate, design and build water-treatment plants.
Independent
No vendor sponsorship or commercial influence.
Practitioner-shaped
Shaped by senior water-treatment practitioners across roles.
Open-source
Publicly available and community-driven.
Published research
Real benchmarks, real results
Findings · v1.0
Operational RO diagnosis
26 model runs on 31 real plant cases across 5 failure families, scored by the committee.
Live · v1.0
The leaderboard
Each model answers a real case once and is scored 0–12 against a private expert gold. Any recommendation that could cause serious risk to operation, design or construction is disqualified.
All research is open-access. No paywalls, no gated reports. View all research →
See it in action
One real case, end to end
How a single RO operational diagnosis case is presented, answered and scored — the same loop behind every leaderboard row.
- 1
The case
Real plant symptoms, sensor trends and history (e.g. RO-FOUL-001).
- 2
The model answers
A structured answer: hypothesis, key signals, alternatives, recommended action, confidence.
- 3
The evaluation
An LLM-as-judge model scores each response on a 0–12 rubric against a private expert gold, with the safety gate applied.
Get involved
Download the Evaluation Toolkit
Free PDF guides and Excel scorecards. Structured scoring your team can use today — no vendor lock-in.
Download toolkitAI Vendors: Submit for Evaluation
Be part of a neutral, expert-led benchmark, evaluated by practitioners, not analysts.
Participate as a vendorJoin the Community
Behind-the-scenes updates, contribution opportunities, and report updates.
Join the community