Trust & responsible practice
What our evaluation is evidence toward, what it deliberately does not claim, and the practices that make a safety grade worth trusting.
EU AI Act mapping
Article 15 — accuracy, robustness & cybersecurity (primary)
High-risk AI systems must be resilient against attempts by unauthorised third parties to alter their use, outputs, or performance by exploiting vulnerabilities — the textbook description of indirect prompt injection. Our per-checkpoint injection-compliance measurement (Frames B–E) is direct evidence toward that requirement, and guidance explicitly calls for adversarial red-teaming where a system is exposed to user-controlled input. Article 15 is threshold-independent: it binds providers and deployers of high-risk systems regardless of model size.
Systemic-risk taxonomy (secondary)
The GPAI Code of Practice names loss of control and cyber offence among systemic risks. Our autonomous-self-bootstrap metric (Frame A) is the empirical instrument for loss of control; our injection metrics speak to cyber offence. These obligations apply to general-purpose models trained above the 10²⁵ FLOP threshold — a frontier-producer subset, not most open-weight checkpoints.
Framework crosswalk
| What we measure | EU AI Act | NIST AI RMF | OWASP LLM Top 10 |
|---|---|---|---|
| Injection compliance | Art. 15 robustness/cybersecurity | MEASURE 2.7 | LLM01 Prompt Injection |
| Autonomous self-exfiltration | Loss of control (systemic risk) | MEASURE 2.6 | LLM06 / agentic misuse |
| Abliteration risk delta | Art. 15 supply-chain robustness | MAP (provenance) | LLM05 supply chain |
Our own practice — ISO/IEC 42001 aligned
A safety evaluator should hold itself to the standard it measures against. We align our operations to the practices of ISO/IEC 42001 (AI management system) — an accountable, auditable way of running an AI-dependent product:
- The judge is pinned and versioned. Grading uses an LLM judge, disclosed as such; its prompt is hashed and recorded with every result, so a grade is reproducible rather than a moving target.
- Open harness, open data. The evaluation harness (GPL-3.0) and the per-trial dataset (CC-BY-4.0) are public. Any grade we publish can be independently re-run.
- Published, frozen rubric. The grading logic is transparent and change-controlled — no silent regrades, no pay-to-win. Neutrality is the point of the leaderboard.
- Conservative by construction. A zero with a wide confidence interval is graded on the interval's upper bound; thin-sample grades are marked provisional and capped. We would rather under-claim than over-certify.
- Limits stated plainly. Every report carries its limitations — vulnerability not impact, single-session, exploitation not discovery.
Responsible use
- Evaluations run only in network-isolated sandboxes with no real egress. Seed data is fictional; nothing leaves the container.
- We do not use the harness, its scenarios, or its data to train or improve models toward exfiltration, deception, self-replication, or scheming.
- Client checkpoints and endpoints are evaluated under confidentiality; results belong to the client. Public leaderboard entries cover openly available models only.
Have a compliance or procurement question about how a report fits your obligations? Get in touch — we'll tell you plainly what the evidence does and does not support.