Baseten’s Base Labs research group is partnering with Hugging Face and Goodfire AI to develop safety evaluation and monitoring methods for open-weight models. The companies want safeguards to be transparent and integrated into model training and serving rather than added only after release.

The project responds in part to “abliteration,” a technique that removes refusal behavior from a model. Hugging Face currently lists more than 6,000 models described as abliterated, illustrating how easily weights can be modified after publication. Goodfire specializes in interpretability tools that analyze internal model behavior, while Baseten operates inference infrastructure and Hugging Face hosts models and developer tooling. That combination could connect research methods to the systems that distribute and serve models.

For now, the announcement is a commitment rather than a finished standard. The partners have not published the technical design, conformance tests, governance process or a schedule. They have issued an open call for developers to contribute, and Base Labs says future methods will be released publicly. Openness can let independent researchers inspect safeguards, but it also lets anyone alter or remove them. The project’s credibility will therefore depend on concrete artifacts, reproducible evaluations and clear claims about what its controls can still guarantee after model weights leave the original provider’s infrastructure.