Base Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire

The answer to abliteration might be building safety into open models — not bolting it on.

Published: 2026-09-18 Category: Quick Take Sources: TechCrunch

Safety as an open-model standard

Baseten's research arm, Base Labs, has teamed up with Hugging Face and Goodfire AI on a new safety infrastructure standard for open-weight models. The pitch is that safety should be "built into how models are trained and deployed, rather than bolted on afterward" — a transparent, open standard the whole ecosystem can adopt. "Openness provides more visibility into the behavior of models," Base Labs said, "and greater means of turning safety research into actionable and transparent controls than closed-source."

The timing is pointed. The debate over open-weight safety has intensified around abliteration — a technique for stripping a model's safeguards. Hugging Face currently lists more than 6,000 abliterated models on its platform, a striking reminder that open models trade safety for accessibility by default. This partnership is essentially a bet that you can have both.

Who does what

The companies haven't disclosed the technical mechanics, but the roles are reasonably clear. Baseten, an AI inference provider, leads with Base Labs. Goodfire, which specializes in opening AI's "black box" to explain how models make decisions, is the likeliest candidate for the "built into" piece — interpretability as the enforcement mechanism. Hugging Face brings the distribution layer where open models actually live. Goodfire summed up the philosophy: "Safety must be built into open models and provided by those who serve them."

The firepower under this isn't trivial. Baseten raised a $1.5 billion Series F in June at a $13 billion valuation; Goodfire raised a $150 million Series B led by B Capital earlier this year. This isn't a garage partnership — it's well-capitalized incumbents moving on a structural problem.

The bet beneath the bet

The whole initiative rests on a contested claim: that openness is an advantage for AI safety, not a vulnerability. Closed labs argue the opposite — that open weights let bad actors copy and weaponize models without accountability. The 6,000 abliterated models are the strongest counter-example to the open camp, and Base Labs is essentially saying "we can fix that." If it works — building interpretability and monitoring in at training time, so stripping a guardrail doesn't silently create a weapon — it could become the blueprint for how open models stay safe at scale. If it doesn't, it's a well-funded slogan. Either way, it's the most concrete attempt yet to give open-weight safety a real infrastructure.


Analysis based on reporting by TechCrunch. Read the full story at techcrunch.com.