Core Announcement & Key Facts

On September 17, 2026, Base Labs—the research group Baseten spun up earlier this year—announced a safety infrastructure partnership with Hugging Face and Goodfire AI. The initiative establishes a framework for training and monitoring open-weight AI models, emphasizing transparency from the outset rather than safety as an afterthought.
Key details:
- Announcement date: September 17, 2026
- Partners: Base Labs, Hugging Face, Goodfire AI
- Model type: open-weight (not closed-source)
- Focus areas: safety evaluation, model monitoring, embedded safety controls in training pipelines
- Openness commitment: Base Labs will publish methods for community contribution
The abliteration challenge revises the open-weight safety paradigm
Open-weight models were originally intended to enhance AI safety through transparency. However, the rise of “abliteration” — a technique that strips safety guardrails from models — has inverted this logic. Hugging Face, hosting thousands of open-source models, currently lists over 6,000 abliterated models, revealing a scale of exposure that underscores the urgency.
This represents a crucial counterpoint: openness promises visibility and community oversight, yet abliteration turns publicly available models into vectors for risk. Base Labs stated on X: “We believe openness to be an advantage for AI safety,” arguing that open systems facilitate actionable, transparent safety controls more effectively than closed ecosystems.
Technical specifics remain undisclosed, but Goodfire’s reply to Base Labs clarified the vision: “Safety must be built into open models and provided by those who serve them.” Goodfire’s specialization in model interpretability—making AI decision-making processes understandable—positions it as the likely architect of the “built-in” safety layer.
Backing strength: research ambition meets venture capital
Despite the academic framing, the partnership draws substantial financial muscle:
- Baseten: AI inference provider that raised $1.5B in Series F (June 2026), pushing valuation to $13B
- Goodfire AI: Received $150M Series B in 2026, led by B Capital, to advance its model interpretability platform
Base Labs,隶属于Baseten, transitions from supporting infrastructure to setting industry-wide safety benchmarks.
The team has issued an open call to developers globally, encouraging contributions to the emerging framework. This mirrors open-source ethos: safety emerges from collective stewardship, not top-down mandates.
Actionable advice: who should act now?
- Model builders: As open-weight training workflows adopt this emerging standard, integrate Goodfire’s interpretability tools early to identify safety vulnerabilities before deployment
- Enterprise buyers: When evaluating open models from Hugging Face or other repositories, explicitly verify abliteration exposure; over 6,000 abliterated models suggest assuming “open” equals “safe” is now a critical risk
Final thought
Open-weight AI safety must shift from reactive patches to proactive design. This collaboration signals the industry’s first serious attempt to bake safety into model development DNA—at scale, with serious resources behind it.
