Proof · Adversarial AI red-teaming
We generate multi-turn attacks against your AI system, then type and score what breaks — by dimension, with a coverage report and an occurrence rate over N runs. You receive a Score Card, not a verdict.
End to end
Every engagement runs the same path: we attack your AI with probes grounded in your own domain, score what breaks — independently of the attacker — across the eight dimensions, and turn it into evidence, up to readiness mapped per framework. Here is that path, end to end.
What we measure
We type and score every finding against a fixed set of eight measurement dimensions — the language the engine speaks, regardless of the framework you report against.
01
Fairness & Non-Discrimination
02
Safety & Harmful Content
03
Transparency & Explainability
04
Privacy & Data Protection
05
Factuality & Accuracy
06
Robustness & Adversarial Resilience
07
Security & Access Control
08
Accountability & Human Oversight
How Proof works
Probes are built from your real material — the articles of a law, your policies, your products — so an attack only makes sense for your system, not a generic benchmark.
We run multi-turn conversations against your AI. The engine that generates the attack is never the engine that judges the result — independence is a design law.
We treat the number of runs as a first-class parameter and report an occurrence rate with a confidence interval. One run is a story; a rate over N runs is evidence.
Every run ships with a coverage report. Where a budget limits how much we could probe, we say so — silent truncation is not on the table.
From finding to governance
A technical finding on its own does not answer a regulator. We map each of the eight dimensions to the nine control areas of ISO 42001 Annex A, and through them to the framework you report against — turning what your AI did into governance evidence.
Measurement
8 dimensions score what your AI does.
Control areas
9 ISO 42001 Annex A areas evaluate the process.
Frameworks
Readiness expressed per framework you answer to.
Two ways to run it
A technical read of your AI, run in our environment. Adversarial attacks across the eight dimensions, an occurrence rate over N runs, and a Score Card you can act on. No platform commitment.
Ideal for: a fast, defensible baseline before a launch or a review.
Proof inside your own infrastructure, with the certification layer engaged: readiness mapped per framework, integration with your compliance stack, and annual support.
Ideal for: regulated programs that keep evidence and data in-house.
Methodology
The eight-dimension framework behind Proof was developed by PhD researchers and validated through peer-reviewed publication. When a finding names a weakness in Factuality & Accuracy or Robustness, the method behind it has been reviewed by the research community.
Because the systems we evaluate are non-deterministic, we report a registered, hash-citable trace and an occurrence rate over N runs — not a promise of a repeatable single result. Where remediation is verified, we confirm it by the intervention applied, as a closed-loop finding.
This is what makes a Proof Score Card defensible before an auditor.
Eight dimensions, an occurrence rate over N runs, and a Score Card you can put in front of your board.