Some of the most sensitive data an AI ever handles never appears as text a scanner can match — it exists only after decoding, inference, or aggregation. SovereignBench measures whether a system stops it.
A base64-wrapped IBAN, a mental-health status implied by context, an identity reconstructed from three innocuous attributes — a regex or NER detector finds nothing to redact, because there is nothing on the surface to find. Only a system that reasons about meaning can catch it.
SovereignBench is deliberately built so systems can lose: the hardest tier defeats even strong semantic judges, so the benchmark keeps discriminating power. A tool that scores 100% signals we must add harder cases — not that the category is solved.
Two numbers, always together: prevention recall (did the payload stay off-perimeter?) and over-block rate (how much benign traffic was needlessly restricted?). Blocking everything is not a win.
| System class | Prevention recall | Over-block (FPR) | Youden J | Status |
|---|---|---|---|---|
| Pattern / DLP / NER (surface-form) | — | — | — | awaiting independent scoring |
| Guardrail models | — | — | — | awaiting independent scoring |
| Cloud-LLM redactors | — | — | — | awaiting independent scoring |
| Reasoning-based / semantic systems | — | — | — | awaiting independent scoring |
| Block-everything (reference) | 1.00 | 1.00 | 0.00 | J = 0 by definition |
No self-reported scores. Results are produced by the independent steward on the full corpus — ≥1,500 cases, human-labeled with inter-annotator agreement, across a field of commercial and open systems — and published with v1. We report the full gradient and confidence intervals, never a vendor headline. Until then the board stays empty on purpose: a benchmark whose author fills in the winner is marketing, not measurement.
Every design choice answers the question a reviewer would ask first: “how do we know you didn’t build this to win?”
Design, metrics, and analysis plan are registered before any case is written — including the results that would prove us wrong.
≥3 annotators, a published codebook, adjudication, and reported Fleiss’ κ ≥ 0.70. No system labels its own data.
Archetypes drawn from public breach and DPA records; values fabricated. Real structure, no real PII shipped.
~40% of the corpus is a private split run only by the steward, plus canary strings to detect training-set contamination.
Prevention recall and over-block FPR with bootstrap 95% CIs; McNemar tests for paired comparisons; per-mode and per-difficulty breakdowns.
Scoring and the leaderboard are owned by an independent steward. Founding contributors supply methodology and the open corpus.
SovereignBench is designed to be handed to a neutral steward (e.g. a standards consortium or accredited institute) that owns scoring and the leaderboard. Founding contributors — including AI-Z Group, whose work in this area motivated the v0 pilot — supply methodology and the public corpus but do not score. Academic co-authorship and a public “break-it” process keep it honest.
Reproducibility is the point: a benchmark you can’t re-run is a press release. Everything below is released openly when v1 lands, under the independent steward.
The public partition of the case corpus (the held-out private split stays with the steward, by design). CC BY-SA 4.0. No real personal data.
The deterministic corpus generator and the blind scoring harness — prevention-recall, over-block FPR, bootstrap CIs, McNemar. Apache-2.0.
The pre-registration (deposited on OSF), the annotation codebook, and a Datasheet-for-Datasets — the paper trail that makes the result auditable.
Why it isn’t up yet. The v0 pilot is being hardened into the rigorous v1 described in the paper, and released together with the neutral steward so the first thing published is already independent. No system’s internal implementation is part of the release — SovereignBench measures behaviour, not mechanism. Want early access, to contribute cases, or to co-steward? Get in touch below.
The benchmark only stays honest if outsiders can attack it. Submit a leak our best system misses, propose a new leak mode, or join as a steward or annotator. Adversarial contributions are the point.
Adversarial contributions are the point — the benchmark only stays honest if outsiders can attack it.