Filter benchmark

Every filter, measured.

A public corpus of 1,000 fictitious documents measures what each Sluis filter buys you: detection rate, false positives and latency. Reproducible by anyone.

Methodology

Every document is fictitious and carries gold annotations: strings that must be redacted, and near-miss strings that must survive. A filter that redacts everything scores 100% recall; the false-positive column is what keeps the number honest.

  • 1,000 generated documents: structured PII, names and organisations, secrets, prompt injections, and hard clean negatives.
  • Six languages: Dutch, English, German, French, Italian and Spanish, as text, email, code/log, DOCX and text-layer PDF.
  • Per feature and per preset: detection rate (recall) AND false positives per 1,000 words, plus the latency impact.
  • A pinned Claude model additionally judges residual identifiability and adjudicates false positives; both scores are published.

Results

Results pending publication

The first full matrix run against a tagged Sluis release is not published yet. Results appear here, per filter and per preset, with the next release.

dataset / pending

The dataset is being published

The benchmark corpus is not downloadable yet. Send us a note and we will share the link the moment it is up.

Email us