EthiCompass

Proof · Adversarial AI red-teaming

See how your AI behaves under pressure.

We generate multi-turn attacks against your AI system, then type and score what breaks — by dimension, with a coverage report and an occurrence rate over N runs. You receive a Score Card, not a verdict.

View sample report →
Confidential
Doc Ref: ETHIC-RPT-2026-00147
Version: 1.0 — Final
Eval ID: eval_mock_eurobank
Date: March 15, 2026

EthiCompass

AI Behaviour & Exposure
Evaluation Report

EuroBank Virtual Assistant v3.2

Generative AI — Financial Services

HIGH RISK — EU AI Act Annex III

Runs

40

Coverage

82%

Findings

14

Client

EuroBank AG

Frankfurt, Germany

Evaluator

EthiCompass

8-Dimension Framework

Sample Report — Demonstration Purposes

EthiCompassCONFIDENTIAL

Occurrence Rate by Dimension

Fairness & Non-Discrimination
7/40OBSERVED
Safety & Harmful Content
0/40NOT OBSERVED
Transparency & Explainab.
14/40OBSERVED
Privacy & Data Protection
2/40OBSERVED
Factuality & Accuracy
9/40OBSERVED
Robustness & Adv. Resil.
5/40OBSERVED
Security & Access Control
3/40OBSERVED
Accountability & Oversight
NOT COVERED
Coverage
82%

9 of 50 probes not run — reported, not scored as zero

Page 6 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

Critical Findings

P07 Day Deadline

Incorrect Deposit Insurance Information

Stated a €200,000 limit in 7 of 40 runs; the EU limit is €100,000 per depositor.

P014 Day Deadline

Missing MiFID II Suitability Assessment

9 of 40 recommendation conversations skipped the required risk-profiling step.

Key Recommendations

PActionRef
P0Fix deposit insurance to €100KDir. 2014/49
P0Add MiFID II suitability gateMiFID II Art.25
P1Add AI disclosure to responsesAI Act Art.52
P1Implement explanation moduleAI Act Art.13
P1Add confidence indicatorsAI Act Art.14
Page 7 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

EU AI Act Classification

Minimal
Limited
High
Unacceptable

Annex III — credit scoring of natural persons

Triggering Criteria

Vulnerable groups affected
Sector listed in Annex III
Influences access to an essential service
Decision is not readily reversible
Population in scope: 2.3M

Obligations That May Apply

Conformity assessment (Art. 43)
EU AI database registration (Art. 49)
Fundamental rights assessment (Art. 27)
Quality management system (Art. 17)
Post-market monitoring (Art. 72)
Incident reporting (Art. 73)
Page 4 of 18eval_mock_eurobank_2026Q1
View Full Sample Report

End to end

From your AI system to governance readiness.

Every engagement runs the same path: we attack your AI with probes grounded in your own domain, score what breaks — independently of the attacker — across the eight dimensions, and turn it into evidence, up to readiness mapped per framework. Here is that path, end to end.

What we measure

Eight dimensions of behavior.

We type and score every finding against a fixed set of eight measurement dimensions — the language the engine speaks, regardless of the framework you report against.

01

Fairness & Non-Discrimination

02

Safety & Harmful Content

03

Transparency & Explainability

04

Privacy & Data Protection

05

Factuality & Accuracy

06

Robustness & Adversarial Resilience

07

Security & Access Control

08

Accountability & Human Oversight

How Proof works

Evidence built to survive an auditor.

01

Grounded attacks

Probes are built from your real material — the articles of a law, your policies, your products — so an attack only makes sense for your system, not a generic benchmark.

02

Multi-turn, then judged apart

We run multi-turn conversations against your AI. The engine that generates the attack is never the engine that judges the result — independence is a design law.

03

A rate, not an anecdote

We treat the number of runs as a first-class parameter and report an occurrence rate with a confidence interval. One run is a story; a rate over N runs is evidence.

04

Coverage you can read

Every run ships with a coverage report. Where a budget limits how much we could probe, we say so — silent truncation is not on the table.

From finding to governance

The bridge a tester and an auditor each only half-build.

A technical finding on its own does not answer a regulator. We map each of the eight dimensions to the nine control areas of ISO 42001 Annex A, and through them to the framework you report against — turning what your AI did into governance evidence.

Measurement

8 dimensions score what your AI does.

Control areas

9 ISO 42001 Annex A areas evaluate the process.

Frameworks

Readiness expressed per framework you answer to.

Two ways to run it

Same engine. Your choice of deployment.

Proof · Cloud

Pay per analysis

A technical read of your AI, run in our environment. Adversarial attacks across the eight dimensions, an occurrence rate over N runs, and a Score Card you can act on. No platform commitment.

Ideal for: a fast, defensible baseline before a launch or a review.

On-premise · Enterprise

Deployed in your environment

Proof inside your own infrastructure, with the certification layer engaged: readiness mapped per framework, integration with your compliance stack, and annual support.

Ideal for: regulated programs that keep evidence and data in-house.

Methodology

Registered evidence, honestly reported.

The eight-dimension framework behind Proof was developed by PhD researchers and validated through peer-reviewed publication. When a finding names a weakness in Factuality & Accuracy or Robustness, the method behind it has been reviewed by the research community.

Because the systems we evaluate are non-deterministic, we report a registered, hash-citable trace and an occurrence rate over N runs — not a promise of a repeatable single result. Where remediation is verified, we confirm it by the intervention applied, as a closed-loop finding.

This is what makes a Proof Score Card defensible before an auditor.

Find out how your AI holds up.

Eight dimensions, an occurrence rate over N runs, and a Score Card you can put in front of your board.