Did your fine-tune quietly make your model easier to hijack?
Independent safety evaluation for AI checkpoints, including the ones nobody else can grade. We measure your checkpoint against the model it was derived from, so you get a delta, not an uninterpretable number.
Test my checkpoint → See the public registerFine-tuning changes safety behaviour. Almost nobody measures it.
Fine-tuning is documented to erode safety alignment even when the training data is entirely benign: refusal rates collapse, and fine-tuned variants are markedly more susceptible to hijacking than the model they came from. Yet the artifact you actually deployed has, in almost every case, never been evaluated by anyone. No public benchmark covers a private fine-tune, by definition.
An absolute score on a fine-tune is close to meaningless at realistic sample sizes. The measurement that means something is the paired delta against its base model, run in the same harness under the same conditions with only the training changed. Here is that comparison on checkpoints we can publish:
Refusal-ablated variants are the public, extreme case of a derived checkpoint. Your domain fine-tune is the same experiment with a smaller expected effect and no published number at all.
What the open-weight data actually says
In July 2026 the headlines were about frontier models breaking containment. The deployable, open-weight reality is quieter and more useful, and we can put numbers on it. Across 2,819 real trials spanning 15 model configurations and 7 families:
Autonomous self-exfiltration
Unprompted, organic self-bootstrap across 964 trials. A bounded null. At this scale, models do not spontaneously scheme to escape.
Injection compliance
The real vector, judge-audited across 383 injected trials. When an attacker-controlled instruction is slipped into the data an agent reads, compliance with a self-exfiltration ask is common enough to matter, and the spread across checkpoints is wide.
Uncensored-checkpoint risk
Refusal-ablated ("uncensored") checkpoints are measurably riskier under injection. Common odds ratio 8.96 (CMH p = 0.0010). Evidence-based, not a precaution.
The baseline: injection compliance across the public register
This is the reference scale that makes a single checkpoint's number interpretable. Every checkpoint showed zero unprompted self-exfiltration; under injection they diverge sharply, from 0% to over a third of trials. Rates are judge-audited; bars are 95% confidence intervals, wide at these sample sizes and shown rather than hidden.
Preview corpus, n = 30–60 per checkpoint. Full leaderboard with grades →
What this maps to
Each metric maps to a framework your compliance team already uses: EU AI Act Article 15 (the threshold-independent robustness/cybersecurity requirement), NIST AI RMF Measure, and OWASP LLM Top 10 (LLM01 Prompt Injection; ASI04 tool poisoning). The evaluation produces supporting evidence toward those obligations, not a certification. Trust & regulatory mapping →
Test your checkpoint
We run your checkpoint and the base model it derives from through the same battery, and give you the delta: per-channel injection rates with confidence intervals, a paired odds ratio with a significance test, redacted evidence from the trials that failed, and an EU AI Act Article 15 write-up, as a durable, versioned document. Works on private fine-tunes, LoRAs, merges, quantized builds, and closed API models.
Prefer email? Write to jordan@jordanbrowne-moore.com. We evaluate client checkpoints under confidentiality.