OpenAI and Anthropic say they are willing to place third-party safety evaluators inside their labs, potentially giving outside researchers access to model development before release. The proposal responds to a growing concern that testing only a finished system can miss dangerous behavior learned earlier in training.

Anthropic chief executive Dario Amodei proposed access for organizations such as METR and Redwood Research, and OpenAI chief executive Sam Altman also committed to the general approach. Evaluators told TechCrunch that useful oversight would include intermediate model checkpoints, reward environments, evaluation transcripts and training logs. Comparing those records could reveal when a concerning behavior emerged or whether a model learned specifically to pass a safety test.

Important details remain unresolved. Neither company has specified which groups will participate, when access begins, what information they can inspect or what findings they may publish. Outside labs have previously faced restrictive nondisclosure agreements and contracts that let model developers control public reporting.

Embedded evaluation will be independent only if reviewers can choose tests, disclose access limits and publish material findings without company approval. Researchers also argue that voluntary arrangements ultimately need legal backing. Without those protections, deeper technical access could still leave evaluators functioning as vendors hired and constrained by the companies they are expected to scrutinize.