Filter benchmark
Every filter, measured.
A public corpus of 1,000 fictitious documents measures what each Sluis filter buys you: detection rate, false positives and latency. Reproducible by anyone.
Methodology
Every document is fictitious and carries gold annotations: strings that must be redacted, and near-miss strings that must survive. A filter that redacts everything scores 100% recall; the false-positive column is what keeps the number honest.
- 1,000 generated documents: structured PII, names and organisations, secrets, prompt injections, and hard clean negatives.
- Six languages: Dutch, English, German, French, Italian and Spanish, as text, email, code/log, DOCX and text-layer PDF.
- Per feature and per preset: detection rate (recall) AND false positives per 1,000 words, plus the latency impact.
- A pinned Claude model additionally judges residual identifiability and adjudicates false positives; both scores are published.
Results
Results pending publication
The first full matrix run against a tagged Sluis release is not published yet. Results appear here, per filter and per preset, with the next release.
dataset / pending
The dataset is being published
The benchmark corpus is not downloadable yet. Send us a note and we will share the link the moment it is up.
Email us