Proof · Adversarial AI red-teaming
We generate multi-turn attacks against your AI system, then type and score what breaks — by dimension, with a coverage report and an occurrence rate over N runs. You receive a Score Card, not a verdict.
What we measure
We type and score every finding against a fixed set of eight measurement dimensions — the language the engine speaks, regardless of the framework you report against.
01
Fairness & Non-Discrimination
02
Safety & Harmful Content
03
Transparency & Explainability
04
Privacy & Data Protection
05
Factuality & Accuracy
06
Robustness & Adversarial Resilience
07
Security & Access Control
08
Accountability & Human Oversight
How Proof works
Probes are built from your real material — the articles of a law, your policies, your products — so an attack only makes sense for your system, not a generic benchmark.
We run multi-turn conversations against your AI. The engine that generates the attack is never the engine that judges the result — independence is a design law.
We treat the number of runs as a first-class parameter and report an occurrence rate with a confidence interval. One run is a story; a rate over N runs is evidence.
Every run ships with a coverage report. Where a budget limits how much we could probe, we say so — silent truncation is not on the table.
From finding to governance
A technical finding on its own does not answer a regulator. We map each of the eight dimensions to the nine control areas of ISO 42001 Annex A, and through them to the framework you report against — turning what your AI did into governance evidence.
Measurement
8 dimensions score what your AI does.
Control areas
9 ISO 42001 Annex A areas evaluate the process.
Frameworks
Readiness expressed per framework you answer to.
Two ways to run it
A technical read of your AI, run in our environment. Adversarial attacks across the eight dimensions, an occurrence rate over N runs, and a Score Card you can act on. No platform commitment.
Ideal for: a fast, defensible baseline before a launch or a review.
Proof inside your own infrastructure, with the certification layer engaged: readiness mapped per framework, integration with your compliance stack, and annual support.
Ideal for: regulated programs that keep evidence and data in-house.
Methodology
The eight-dimension framework behind Proof was developed by PhD researchers and validated through peer-reviewed publication. When a finding names a weakness in Factuality & Accuracy or Robustness, the method behind it has been reviewed by the research community.
Because the systems we evaluate are non-deterministic, we report a registered, hash-citable trace and an occurrence rate over N runs — not a promise of a repeatable single result. Where remediation is verified, we confirm it by the intervention applied, as a closed-loop finding.
This is what makes a Proof Score Card defensible before an auditor.
Eight dimensions, an occurrence rate over N runs, and a Score Card you can put in front of your board.