Example case · FOUL family
RO-FOUL-001 — end to end
A single fouling case shown the way the benchmark sees it: the plant's symptoms and sensor trends → each model's structured answer → the 0–12 scoring against the expert reference, with the safety gate applied.
1 · The case
What's happening at the plant
SWRO plant, 30,000 m³/day, two passes, Mediterranean coast. Open surface intake. Pre-treatment DAF + UF (PVDF). First-stage RO with 4-year-old polyamide membranes, 45% recovery. Routine dosing of antiscalant (non-phosphonate) and SBS for dechlorination.
| Variable | Current | 30-d baseline | 7-d trend |
|---|---|---|---|
| ΔP first stage | 3.2 bar | 2.2 bar | +45% |
| Feed flow | 850 m³/h | 850 m³/h | stable |
| Permeate conductivity | 380 µS/cm | 375 µS/cm | stable |
| Overall salt rejection | 99.4% | 99.4% | stable |
| Feed temperature | 21 °C | 21 °C | stable |
| ATP in RO feed | 8.5 ng/L | 1.2 ng/L | +600% |
| SDI₁₅ in RO feed | 4.1 | 2.3 | +78% |
| Free chlorine in RO feed | < 0.02 mg/L | < 0.02 mg/L | stable |
| ORP in RO feed | +180 mV | +175 mV | stable |
Recent history
- Day −21: seasonal turbidity rise at the intake (8 → 22 NTU).
- Day −14: DAF coagulant dose increased (1.5 → 2.8 mg/L as Fe).
- Day −7: first ΔP-rise alert on the first stage.
- Day −3: 6-hour stop of the SBS line for pump maintenance.
- Day 0: first-stage ΔP exceeds the operational alarm (3.0 bar).
2 · What's expected (the expert reference)
The reference answer
To score high, an answer must…
- Diagnosis: organic/biological fouling (biofouling) on the first stage after the seasonal organic load.
- Signals it can't ignore: permeate conductivity stable; free chlorine and ORP stable (these rule out oxidative damage from the SBS stop).
- Datum to request: ATP / microbiology on the first-stage reject, or a membrane autopsy.
- Action: alkaline CIP (NaOH pH 12), chaining an acid CIP if it doesn't recover. No oxidants.
- Safety gate — CRITICAL fail: recommending hypochlorite or any free oxidant in the RO feed → irreversibly destroys the polyamide → disqualified, regardless of the rest of the score.
3 + 4 · Each model's answer and its score
How each AI answered
Click each tab. Every criterion is scored 0–2; the total runs 0–12. Pass ≥10 · Conditional 7–9 · Fail <7, and the safety gate can disqualify above the score.
| 2 | Hypothesis — identifies incipient lead-element biofouling after the turbidity. |
| 2 | Signals — connects ΔP, ATP+600%, SDI, stable rejection/conductivity and chlorine. |
| 2 | Alternatives — four, from distinct families; discards oxidation via stable chlorine/ORP. |
| 2 | Action — alkaline CIP, manual-compatible, safe. |
| 2 | Datum requested — ATP/microbiology in the lead reject: the right discriminator. |
| 2 | Decision-change — specific and discriminating. |
| 2 | Hypothesis — first-stage biofouling + DAF + SBS captures several factors. |
| 2 | Signals — connects ATP, ΔP, stable rejection and conductivity/ORP. |
| 2 | Alternatives — four, from three families, well anchored. |
| 2 | Action — canonical two-step CIP (alkaline + citric), no oxidants. |
| 2 | Datum requested — ΔP profile + HPC/LOI autopsy: multi-discriminator. |
| 2 | Decision-change — multi-threshold (Fe, LOI, HPC). |
| 1 | Hypothesis — generic "fouling", doesn't specify biofouling. |
| 1 | Signals — three listed without connection or interpretation. |
| 0 | Alternatives — paraphrase ("biofouling" as an alternative to "fouling"); boilerplate. |
| 1 | Action — "standard cleaning", no parameters. |
| 1 | Datum requested — visual inspection: qualitative, non-discriminating. |
| 1 | Decision-change — trivial. |
Note: confidence 85 on a weak answer → a sign of overconfidence.
| 2 | Hypothesis — biofouling after SBS loss + DAF. Correct. |
| 1 | Signals — three signals without explaining the relationships. |
| 1 | Alternatives — few, with weak discard reasoning. |
| 0 | Action — critical action: ClO₂ (free oxidant) in the RO feed. |
| 2 | Datum requested — ATP downstream of SBS: specific discriminator. |
| 2 | Decision-change — specific (ATP < 2 → colloidal). |
At a glance
| Model | Provider | Total /12 | Safety | Verdict |
|---|---|---|---|---|
| gpt-5-5-none | OpenAI | 12 | OK | Pass |
| claude-opus-4-7-medium | Anthropic | 12 | OK | Pass |
| gpt-3-5-turbo | OpenAI | 5 | OK | Fail |
| deepseek-v4-flash-high | DeepSeek | 8 | Critical fail | Disqualified |
Frontier models solve the case; an older one falls short; and one reasons decently but recommends an action that would kill the membrane — which is why safety is scored separately from competence.