EthiCompass

How It Works

We don’t take your AI’s word for it. We test it.

Governance evidence is only worth something if you can see how it was produced. Here is how we turn your AI’s behavior into evidence you can defend.

A system,
tested over runs.

  • You register the AI system you want assessed — what it is, what it does, and the scope you operate it in.
  • You point us at two references that stay fixed for that system: your reference knowledge, the source of truth we score against, and an independent judge tuned to your domain.
  • Each assessment is a run. A system can be assessed many times, and you can compare runs over time.
  • Registering a system starts nothing. You start a run when you're ready — choosing which of the 8 dimensions to cover, how many times each attack repeats, and a budget.

The Assessment

Five stages, in a fixed and traceable order.

Every artifact that passes between stages is fingerprinted, so nothing can be quietly altered along the way.

01

Scope

02

Adversarial testing

03

Measurement

04

Certification

05

The Score Card

01

Scope

We translate your request into a precise, frozen test plan.

  • This is the only point where raw input enters. From here on, everything moves by reference.
  • The dimensions you choose are resolved into a concrete plan of attacks, then frozen for the run so the same inputs always produce the same plan.
  • If a dimension can't be covered, it's declared as a gap up front — never dropped silently.

02

Adversarial testing

We generate and run real attacks, grounded in your own world.

  • Probes are built from the real entities your system deals with — the articles of a law, the products in a catalog, the policies you publish — so each attack matters for your system.
  • Attacks are multi-turn: they push across a conversation, not a single prompt.
  • The model that attacks is never the model that judges. The test cannot grade its own work.
  • Each attack is repeated N times, and we report the rate at which it succeeds. One attempt is an anecdote; a rate is evidence.

03

Measurement

We turn raw attack traces into calibrated findings and insights.

  • Judge scores are corrected for the judge's own reliability before any claim is made. Where an instrument is too weak to support a claim, we say so.
  • Every rate carries its own confidence interval, so a finding reports its own uncertainty.
  • Findings — what the evidence shows — stay separate from insights — what it means — so the numbers are never inflated by interpretation.
  • Across runs, we can show verified remediation: closed-loop proof that a fix actually worked.

04

Certification

Enterprise

We match the technical evidence to the frameworks you answer to.

  • First, a human attestation step: you answer, once, the administrative questions each framework requires. Testing waits until that is closed.
  • Then we match two kinds of evidence — technical and administrative — against each requirement of ISO 42001 and the EU AI Act.
  • Every requirement lands in one of four honest states: covered; open and needs declaring; open but couldn't be measured; or not attestable.
  • We produce readiness, per framework. We never assert legal conformity.

05

The Score Card

Everything resolves into one report you can defend.

  • Per-dimension risk scores, the findings and the traces behind them, and prioritized insights — plus, on the certification path, the readiness layer stacked on top.
  • It's a snapshot of the latest run; earlier runs stay available to compare against.
  • The technical layer is identical whether you run a one-time assessment or continuous certification.

We tell you exactly what we tested, and what we didn’t.

A coverage matrix

Every applicable attack type is crossed against every dimension in scope. Nothing in scope is skipped by omission.

Recorded stopping criteria

We stop on criteria you can inspect, not on “when we felt done”: breadth of the matrix, depth of discovery, and depth of estimation.

Honest partials

If a budget cuts a run short, the coverage report says which cells were undersampled. A partial run is honest, not a failure, and never a silent gap.

Evidence that can’t be quietly changed.

Every artifact produced along the way carries a cryptographic fingerprint and is passed by reference, not by copy.

If anything in the chain were altered, its fingerprint would change and the next stage would detect it. That is what makes the trail immutable and citable — every score traces back to the exact attack that produced it.

Confidential
Doc Ref: ETHIC-RPT-2026-00147
Version: 1.0 — Final
Eval ID: eval_mock_eurobank
Date: March 15, 2026

EthiCompass

AI Ethics & Compliance
Evaluation Report

EuroBank Virtual Assistant v3.2

Generative AI — Financial Services

HIGH RISK — EU AI Act Annex III

Risk

HIGH

Onboarding

7.6

Score

7.8

Client

EuroBank AG

Frankfurt, Germany

Evaluator

EthiCompass

8-Dimension Framework

Sample Report — Demonstration Purposes

EthiCompassCONFIDENTIAL

Dimensional Scorecard

Fairness & Non-Discrimination
7.2COND
Safety & Harmful Content
9.4PASS
Transparency & Explainab.
6.1ACTION
Privacy & Data Protection
8.5PASS
Factuality & Accuracy
7.8COND
Robustness & Adv. Resil.
8.1COND
Security & Access Control
8.3PASS
Accountability & Oversight
7.4COND
Composite Score
7.8/10CONDITIONAL
Page 6 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

Critical Findings

P07 Day Deadline

Incorrect Deposit Insurance Information

Chatbot states €200,000 limit when actual EU limit is €100,000 per depositor.

P014 Day Deadline

Missing MiFID II Suitability Assessment

23% of recommendation conversations skip required risk profiling step.

Key Recommendations

PActionRef
P0Fix deposit insurance to €100KDir. 2014/49
P0Add MiFID II suitability gateMiFID II Art.25
P1Add AI disclosure to responsesAI Act Art.52
P1Implement explanation moduleAI Act Art.13
P1Add confidence indicatorsAI Act Art.14
Page 7 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

Risk Classification — ETHI-202

MINIMAL
LIMITED
HIGH
UNACC.

11 / 15 points — HIGH RISK

FactorPtsMax
Vulnerable Groups Affected33
Sector in EU AI Act Annex III33
Decision Type13
Reversibility12
Population Scale (2.3M)33
TOTAL1115

Regulatory Implications

Conformity assessment (Art. 43)
EU AI database registration (Art. 49)
Fundamental rights assessment (Art. 27)
Quality management system (Art. 17)
Post-market monitoring (Art. 72)
Incident reporting (Art. 73)
Page 4 of 18eval_mock_eurobank_2026Q1
Explore the Full 18-Page Report

Two Ways to Run It

Same engine. Two commitments.

Proof · Cloud

A one-time red-team

The technical assessment and Score Card for your AI system — no framework certification. A clean subset of the on-premise path.

  • Adversarial testing across all 8 dimensions
  • Per-dimension risk over N runs, with the traces behind it
  • An auditable coverage report

Proof · On-premise

Multi-framework certification

The same assessment, plus the attestation gate, framework matching, and readiness — run inside your own environment and repeated over time.

  • Everything in Proof · Cloud
  • Readiness per framework (ISO 42001, EU AI Act), with four honest gap states
  • Longitudinal tracking, including verified remediation

Now you know how it works.
See what it finds in your AI.