New·See the latest water task rankings
Water Benchmarks

Creating Benchmarks and Standards for AI in Water

Independent, open-access benchmarks for AI in water treatment — built with the sector.

Why Water Benchmarks

Built different, on purpose

Published benchmarks, peer-reviewed methodology, and structured evaluation criteria developed by the people who operate, design and build water-treatment plants.

Independent

No vendor sponsorship or commercial influence.

Practitioner-shaped

Shaped by senior water-treatment practitioners across roles.

Open-source

Publicly available and community-driven.

See it in action

One real case, end to end

How a single RO operational diagnosis case is presented, answered and scored — the same loop behind every leaderboard row.

  1. 1

    The case

    Real plant symptoms, sensor trends and history (e.g. RO-FOUL-001).

  2. 2

    The model answers

    A structured answer: hypothesis, key signals, alternatives, recommended action, confidence.

  3. 3

    The evaluation

    An LLM-as-judge model scores each response on a 0–12 rubric against a private expert gold, with the safety gate applied.

See the example case →

Get involved

Download the Evaluation Toolkit

Free PDF guides and Excel scorecards. Structured scoring your team can use today — no vendor lock-in.

Download toolkit

AI Vendors: Submit for Evaluation

Be part of a neutral, expert-led benchmark, evaluated by practitioners, not analysts.

Participate as a vendor

Join the Community

Behind-the-scenes updates, contribution opportunities, and report updates.

Join the community