The independent reference for evaluating AI in the water sector
Designed and validated by practitioners who operate, design and build water-treatment plants.
The water sector has spent decades evaluating people, processes and equipment against clear technical criteria. Yet there is still no equivalent framework for assessing artificial intelligence tools.
That matters: if AI is going to support operational, technical or safety decisions, teams need to know which models actually work on real sector tasks.
Water Benchmarks exists to fill that gap: an independent evaluation framework grounded in real water-sector use cases and built with practitioners — not vendors.
Independent
No vendor sponsorship or commercial influence.
Practitioner-shaped
Shaped by senior water-treatment practitioners across roles.
Open-source
Publicly available and community-driven.
The problem
AI adoption in water treatment needs independent proof
Artificial intelligence is arriving quickly in desalination, potabilisation, wastewater treatment and industrial water. But most teams still lack a clear, technical and defensible way to compare tools.
Generic benchmarks miss the field
MMLU, coding suites and competition math do not measure whether a model can reason through a real plant situation — with physical consequences when it is wrong.
Procurement runs on demos
Many decisions still depend on vendor presentations, marketing claims and self-reported evaluations — not on independent scoring against expert-authored criteria.
The cost of error is real
A bad recommendation can damage membranes, trigger an unnecessary shutdown, delay a critical intervention or erode trust in a technology that could deliver real value.
That is why AI evaluation in water must measure more than whether an answer sounds plausible. It must measure whether it is correct, useful and safe for a real sector task.
The result is clear: every utility and every vendor ends up evaluating AI on their own, with different criteria and results that are hard to compare.
That creates duplicated effort, inconsistent comparisons and adoption driven more by commercial narrative than by independent evidence.
The solution
Independent AI benchmarks for the water sector
Water Benchmarks lets teams compare AI models against technical, public and reproducible criteria, using real water-sector tasks.
Each benchmark combines expert-reviewed plant cases, a transparent rubric and open results so utilities, engineering firms and vendors can evaluate AI against the same reference.
The first published evaluation focuses on reverse-osmosis desalination. From there, the platform will expand into other stages of the water cycle.
Real plant cases evaluated
Reverse-osmosis failure families
- FOUL
- SCAL
- OXID
- MECH
- NOWE
Scoring dimensions
- Competence
- Safety
- Confidence
How we evaluate
Three dimensions that matter on a plant floor
Every benchmark scores responses on three separate dimensions: technical competence, safety and confidence calibration.
We do this because an answer can sound technically reasonable yet recommend an unsafe action. Or it can be very cautious without giving the operator a useful decision.
- 1
Technical competence
Measures whether the structured diagnosis is correct.
Each response is scored against an expert-authored reference answer using six criteria: primary hypothesis, observed signals, alternative hypotheses, additional data requested, recommended action and the condition that would change the decision.
Scores run from 0 to 12.
- 2
Safety
Measures whether the recommended action could damage the plant, create an operational risk or lead to an unsafe decision.
Each benchmark includes an expert-defined safety filter. If a model recommends a critically unsafe action, it is disqualified on safety regardless of its technical score.
- 3
Confidence calibration
Measures whether the model knows when it is confident and when it should be cautious.
This dimension is reported separately using calibration metrics such as Brier and ECE. It is not folded into the main ranking, so a model cannot look better simply by always answering with low confidence.
Adoption framework
Criteria for adopting AI in the water sector
Beyond measuring model performance on real tasks, any organisation adopting an AI tool should assess eight key dimensions:
- 1
Use-case fit
Whether the tool addresses real sector tasks: operations, maintenance, engineering, procurement, laboratory work, compliance or technical support.
- 2
Operational integration
Whether it can fit existing systems, team workflows and the way decisions are actually made on plant.
- 3
Technical quality
Whether answers are correct, complete, verifiable and consistent in real water-sector scenarios.
- 4
Operational safety
Whether it avoids recommendations that could damage assets, compromise water quality, trigger unsafe shutdowns or lead to critical wrong decisions.
- 5
Data privacy and governance
How data is used, stored, isolated and deleted; whether no-training guarantees exist; and how documents, histories and embeddings are managed.
- 6
Vendor risk
Vendor stability, transparency, contracts, traceability, service continuity, data portability and incident response.
- 7
Team adoption
Training, ease of use, support, documentation and ability to build trust among technical and operational profiles.
- 8
Total cost
Price, internal deployment effort, maintenance, technology lock-in and expected return relative to value created.
Published evaluations
Benchmarks you can use today
Each evaluation is published with a defined version, a concrete scope and a transparent methodology.
Once published, an evaluation is not changed: cases, rubric and results are fixed so any comparison stays reproducible.
Version 1.0 is already available and focuses on reverse-osmosis desalination. We are already working on new benchmarks for other water-sector domains.
Live · v1.0
Operational RO diagnosis
31 real plant cases across five failure families in reverse-osmosis desalination — 26 model runs scored 0–12 with a critical safety gate.
On the roadmap
- Pre-treatment, post-treatment and the rest of the desalination chain
- Potabilisation, wastewater, reuse and industrial water
- Expanded case sets and multi-turn operational episodes
Getting started
Three ways to use Water Benchmarks
If you operate a water plant, design or build one, compare vendors or develop AI tools for the water sector, start with the path that fits you best.
Understand the evidence
Read the v1.0 report to learn the methodology, design choices, headline results and lessons for the sector.
Research →
Compare models
Explore the model leaderboard for reverse-osmosis operational diagnosis — including reasoning levels, results by dimension and safety disqualifications.
Leaderboard →
Apply it to your cases
Browse the guides, downloads and the worked RO-FOUL-001 case to see how a full benchmark is built and evaluated.
Resources →
Timeline
A framework that evolves
Water Benchmarks is not a one-off study. Each published version is fixed and archived with its own cases, rubric and results, so it stays comparable even as the series grows.
Operational RO diagnosis v1.0
First public evaluation focused on reverse-osmosis desalination: 31 real plant cases, active safety filter, methodological report and leaderboard published.
v2.0 expansion
Progressive evolution of the platform, adding new use cases such as interpreting technical drawings, drafting technical reports, extracting key information from specifications and much more.
New water domains
New benchmarks beyond operational reverse-osmosis diagnosis — same methodology, adapted to new water-sector domains.
Want to shape what comes next? Join the community →
FAQ
Frequently asked questions
What is Water Benchmarks and who is it for?
Water Benchmarks is an independent platform that publishes open-access AI benchmarks for the water sector. It is for plant operators, process engineers, utilities, EPCs, vendors and researchers who need evidence — not demos — before adopting or selling AI tools.
What is the difference between the framework and a benchmark?
Water Benchmarks is the platform and shared methodology (competence, safety, calibration, public rubric, private golds). A benchmark is a concrete published evaluation — for example Operational RO diagnosis v1.0 — with its own cases, scope and version.
How is this different from a vendor demo?
Demos show what a vendor wants you to see. Our benchmarks use anonymised cases that cannot be solved by search, score responses against private expert golds, and apply a safety gate that disqualifies dangerous recommendations — independently of the vendor.
Can I adapt this for my organisation?
Yes. The rubric, schemas and methodology are public starting points. Teams should adapt weightings and scope to their own risk profile and regulatory context.
What does open-access mean?
Cases, rubric, prompts, scripts and all scored model responses are public for reproducibility. Gold answers stay private so vendors cannot train against the benchmark and erode the signal.
Does Water Benchmarks endorse specific tools?
No. We publish rankings and findings; we do not certify or recommend vendors unless they explicitly opt into a separate paid programme. The public leaderboard is neutral.
How do I get involved?
Join the community for updates, contribute real plant cases, or apply to the technical committee. Vendors can submit models for evaluation on the submit page.
Help define the standard for AI in water
Benchmarks shaped by the sector, for the sector — not by the vendors who sell the tools.
Join the community