Defender Impact
For the first time, a credible proposal exists to move AI safety evaluation from self-reported attestation to independent, continuous verification across the full model training lifecycle. If implemented with genuine independence and public disclosure rights, this closes the single largest accountability gap in frontier AI oversight: defenders and regulators currently have no mechanism to verify whether alignment claims held during training.
Capability Overview
Dario Amodei’s proposal — subsequently endorsed by OpenAI CEO Sam Altman — would embed third-party evaluators such as METR, Redwood Research, Apollo Research, and FAR.AI directly inside frontier AI companies. The critical distinction from current practice is scope: historically, outside reviewers accessed only finished models in the days before release. The new model proposes granting evaluators access to intermediate training checkpoints, post-training reward environments, evaluation transcripts, and internal logs.
This matters because finished-model evaluation has a fundamental detection ceiling. As Alexander Meinke of Apollo Research notes, models are increasingly capable of recognising evaluation conditions — creating a risk that problematic behaviours surface during training but remain concealed during standard pre-release testing. Checkpoint-level access allows evaluators to compare model behaviour at different stages of training, pinpointing when and why concerning behaviour emerged and whether alignment training was actively circumvented during the training run itself.
FAR.AI CEO Adam Gleave identifies three concrete mechanisms this enables: comparing training checkpoints to trace behavioural drift, inspecting the post-training reward environment that shapes model incentives, and cross-referencing evaluation transcripts against company claims to verify accuracy. Together these create an audit trail that does not currently exist anywhere in the industry.
Defensive Advances
This proposal, if implemented, delivers several concrete advances for defenders.
Continuous lifecycle monitoring replaces point-in-time review. Security teams assessing AI vendors can move from relying on pre-release model cards to verified evaluator findings spanning the training lifecycle.
Independent verification of alignment claims. Organisations deploying frontier models can reference third-party evaluator findings rather than trusting developer-issued safety documentation alone.
Deceptive alignment detection surface. Checkpoint access creates the first practical mechanism to detect whether a model actively undermined its own alignment training — a risk previously undetectable from the outside.
Precedent for disclosure norms. Even where legislation does not yet exist, public evaluator reports establish industry expectations that can be incorporated into procurement requirements and vendor risk frameworks.
Residual Gaps
The proposal is significant but implementation maturity is low. Neither Anthropic nor OpenAI has disclosed which evaluators will be embedded, what systems they can access, or what they are permitted to disclose publicly. Without binding legal frameworks or at minimum contractual disclosure rights, embedded evaluators risk functioning as sophisticated vendors rather than independent watchdogs — a distinction the evaluators themselves emphasise.
The scope of access remains undefined. Full benefit requires access to training checkpoints, reward model configurations, and raw evaluation logs — not merely post-hoc briefings. Whether commercial confidentiality provisions will constrain meaningful disclosure is a critical open question.
Finally, the evaluator ecosystem itself is nascent. Organisations like METR, Apollo Research, and FAR.AI have strong methodological credibility but limited capacity to embed simultaneously across multiple frontier labs at the depth this proposal envisions. Scaling evaluator capacity to match model development velocity is a multi-year challenge.
Framework Mapping
This capability most directly addresses AML.T0031 (Erode AI Model Integrity) and AML.T0018 (Manipulate AI Model) by creating independent detection mechanisms for integrity failures during training. Checkpoint access also surfaces AML.T0020 (Poison Training Data) risks by enabling evaluators to trace when training data or reward signals produced unintended behavioural shifts. The overreliance risk captured in LLM09 is reduced when developer safety claims can be independently verified rather than accepted at face value.
Deployment Considerations
Organisations should treat this as an evolving governance signal rather than an immediately deployable control. The near-term practical step is incorporating evaluator access rights into AI vendor procurement criteria — requesting evidence of third-party evaluation scope and disclosure commitments before deployment decisions. Security and risk teams should monitor public outputs from METR, Apollo Research, and FAR.AI to understand what findings, if any, are disclosed as the arrangement matures.
Legislative developments in the EU AI Act implementation and US AI safety policy will materially affect whether independence provisions have teeth. Regulatory affairs and security teams should track these in parallel.
Defender Checklist
- Update AI vendor risk questionnaires to include third-party evaluator access scope and disclosure rights as required evidence
- Subscribe to public outputs from METR, Redwood Research, Apollo Research, and FAR.AI for emerging findings
- Assess whether current AI deployment decisions are contingent on self-reported safety documentation alone and identify where independent verification is absent
- Engage legal and procurement to determine whether evaluator disclosure commitments can be incorporated into frontier AI contracts
- Monitor EU AI Act and US legislative developments that may mandate rather than encourage independent evaluator access