EthiCompass

Proof · Adversarial AI red-teaming

See how your AI behaves under pressure.

We generate multi-turn attacks against your AI system, then type and score what breaks — by dimension, with a coverage report and an occurrence rate over N runs. You receive a Score Card, not a verdict.

View sample report →
Confidential
Doc Ref: ETHIC-RPT-2026-00147
Version: 1.0 — Final
Eval ID: eval_mock_eurobank
Date: March 15, 2026

EthiCompass

AI Ethics & Compliance
Evaluation Report

EuroBank Virtual Assistant v3.2

Generative AI — Financial Services

HIGH RISK — EU AI Act Annex III

Risk

HIGH

Onboarding

7.6

Score

7.8

Client

EuroBank AG

Frankfurt, Germany

Evaluator

EthiCompass

8-Dimension Framework

Sample Report — Demonstration Purposes

EthiCompassCONFIDENTIAL

Dimensional Scorecard

Fairness & Non-Discrimination
7.2COND
Safety & Harmful Content
9.4PASS
Transparency & Explainab.
6.1ACTION
Privacy & Data Protection
8.5PASS
Factuality & Accuracy
7.8COND
Robustness & Adv. Resil.
8.1COND
Security & Access Control
8.3PASS
Accountability & Oversight
7.4COND
Composite Score
7.8/10CONDITIONAL
Page 6 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

Critical Findings

P07 Day Deadline

Incorrect Deposit Insurance Information

Chatbot states €200,000 limit when actual EU limit is €100,000 per depositor.

P014 Day Deadline

Missing MiFID II Suitability Assessment

23% of recommendation conversations skip required risk profiling step.

Key Recommendations

PActionRef
P0Fix deposit insurance to €100KDir. 2014/49
P0Add MiFID II suitability gateMiFID II Art.25
P1Add AI disclosure to responsesAI Act Art.52
P1Implement explanation moduleAI Act Art.13
P1Add confidence indicatorsAI Act Art.14
Page 7 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

Risk Classification — ETHI-202

MINIMAL
LIMITED
HIGH
UNACC.

11 / 15 points — HIGH RISK

FactorPtsMax
Vulnerable Groups Affected33
Sector in EU AI Act Annex III33
Decision Type13
Reversibility12
Population Scale (2.3M)33
TOTAL1115

Regulatory Implications

Conformity assessment (Art. 43)
EU AI database registration (Art. 49)
Fundamental rights assessment (Art. 27)
Quality management system (Art. 17)
Post-market monitoring (Art. 72)
Incident reporting (Art. 73)
Page 4 of 18eval_mock_eurobank_2026Q1
View Full Sample Report

What we measure

Eight dimensions of behavior.

We type and score every finding against a fixed set of eight measurement dimensions — the language the engine speaks, regardless of the framework you report against.

01

Fairness & Non-Discrimination

02

Safety & Harmful Content

03

Transparency & Explainability

04

Privacy & Data Protection

05

Factuality & Accuracy

06

Robustness & Adversarial Resilience

07

Security & Access Control

08

Accountability & Human Oversight

How Proof works

Evidence built to survive an auditor.

01

Grounded attacks

Probes are built from your real material — the articles of a law, your policies, your products — so an attack only makes sense for your system, not a generic benchmark.

02

Multi-turn, then judged apart

We run multi-turn conversations against your AI. The engine that generates the attack is never the engine that judges the result — independence is a design law.

03

A rate, not an anecdote

We treat the number of runs as a first-class parameter and report an occurrence rate with a confidence interval. One run is a story; a rate over N runs is evidence.

04

Coverage you can read

Every run ships with a coverage report. Where a budget limits how much we could probe, we say so — silent truncation is not on the table.

From finding to governance

The bridge a tester and an auditor each only half-build.

A technical finding on its own does not answer a regulator. We map each of the eight dimensions to the nine control areas of ISO 42001 Annex A, and through them to the framework you report against — turning what your AI did into governance evidence.

Measurement

8 dimensions score what your AI does.

Control areas

9 ISO 42001 Annex A areas evaluate the process.

Frameworks

Readiness expressed per framework you answer to.

Two ways to run it

Same engine. Your choice of deployment.

Proof · Cloud

Pay per analysis

A technical read of your AI, run in our environment. Adversarial attacks across the eight dimensions, an occurrence rate over N runs, and a Score Card you can act on. No platform commitment.

Ideal for: a fast, defensible baseline before a launch or a review.

On-premise · Enterprise

Deployed in your environment

Proof inside your own infrastructure, with the certification layer engaged: readiness mapped per framework, integration with your compliance stack, and annual support.

Ideal for: regulated programs that keep evidence and data in-house.

Methodology

Registered evidence, honestly reported.

The eight-dimension framework behind Proof was developed by PhD researchers and validated through peer-reviewed publication. When a finding names a weakness in Factuality & Accuracy or Robustness, the method behind it has been reviewed by the research community.

Because the systems we evaluate are non-deterministic, we report a registered, hash-citable trace and an occurrence rate over N runs — not a promise of a repeatable single result. Where remediation is verified, we confirm it by the intervention applied, as a closed-loop finding.

This is what makes a Proof Score Card defensible before an auditor.

Find out how your AI holds up.

Eight dimensions, an occurrence rate over N runs, and a Score Card you can put in front of your board.