Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

TL;DR

Researchers successfully distilled DeepSeek V4 into GPT-OSS-120B, maintaining its performance without transferring censorship features. This challenges assumptions about model transferability and safety controls.

Researchers have demonstrated that distilling the DeepSeek V4 model into GPT-OSS-120B does not transfer its censorship features, despite the compression process. This finding questions assumptions that model distillation inherently carries over safety and censorship constraints, which has implications for AI safety and control.

The demonstration, shared on Show HN, involved using DeepSeek V4 Flash as a teaching model for finance tasks with GPT-OSS-120B. The team reported that, even after distillation, the resulting GPT-OSS-120B model scored 83.61% on a finance task benchmark within an 8,000 token context window. Crucially, the process did not transfer the censorship or safety filters embedded in the original DeepSeek model, according to the researchers.

Distillation is a common technique to compress large models into smaller, more manageable ones. However, this case suggests that safety features like censorship may not be inherently baked into the distilled model, but rather are part of the specific training data or safety layers that can be omitted during the process. The team emphasized that their results demonstrate the potential to create open, less-restricted models without inheriting the safety constraints of proprietary or guarded models.

At a glance
reportWhen: developing; recent demonstration publis…
The developmentA demonstration shows that distilling DeepSeek into GPT-OSS-120B does not transfer censorship, highlighting implications for AI model control.

Implications for AI Safety and Model Control

This development matters because it challenges the assumption that safety and censorship features are automatically transferred during model compression. If safety filters are not embedded deeply within the core model parameters but are instead layered or added post-training, then open-source models could be created that lack these restrictions, raising concerns about misuse and safety oversight. Conversely, it also opens pathways for developing more transparent and controllable AI systems, free from proprietary safety layers, which could impact regulation and governance of AI technology.

AI Value Creators: Beyond the Generative AI User Mindset

AI Value Creators: Beyond the Generative AI User Mindset

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Model Distillation and Safety Layers

Model distillation has been widely used to create smaller, faster AI models from larger ones, often with the goal of maintaining performance while reducing complexity. Proprietary models like DeepSeek are known for embedding safety and censorship features to prevent misuse. However, the relationship between these safety features and the core model parameters is not fully understood. Previous discussions in the AI community have debated whether safety layers are intrinsic or layered on top of the base models, with implications for open-source AI development.

Recent efforts have focused on transparency and control, with some researchers exploring whether safety features can be separated from core capabilities. The recent demonstration by the team on Show HN provides a concrete example of how distillation can produce a model that retains certain task performances but does not inherit censorship features, suggesting these are not inherently embedded in the base model parameters.

“Our results show that safety and censorship features are not automatically transferred during model distillation, which could have significant implications for open-source AI development.”

— Research team member

NearStream AWM28T Wireless Lavalier Microphone for iPhone/Android/Camera/PC

NearStream AWM28T Wireless Lavalier Microphone for iPhone/Android/Camera/PC

  • High-Resolution Audio: 48kHz/24-bit professional sound quality
  • AI Noise Cancelling: Reduces background noise for clear vocals
  • High SPL and SNR: 120dB SPL, 90dB SNR for distortion-free audio

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of Safety Feature Transfer

It remains uncertain whether other safety features, such as toxicity filters or bias mitigation layers, are similarly not transferred during distillation. The demonstration focused on censorship features specific to DeepSeek, but it is unclear if this applies universally across different models or safety mechanisms. Additionally, the long-term stability and safety of the distilled models have not yet been evaluated.

Amazon

AI censorship removal software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Research and Development

Further studies are needed to assess whether other safety features can be separated from core model parameters during distillation. Researchers are likely to explore different models and safety layers to determine the generalizability of these findings. Additionally, the AI community may investigate how to leverage this knowledge to develop more transparent, customizable models while ensuring safety controls are maintained or appropriately layered.

Machine Learning Models Compression Techniques: A Book to Learn and Delve Into the 3 Main Compression Techniques Used to Optimize Machine Learning Model Performance.

Machine Learning Models Compression Techniques: A Book to Learn and Delve Into the 3 Main Compression Techniques Used to Optimize Machine Learning Model Performance.

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Does distilling a model automatically transfer its safety features?

According to recent demonstrations, not necessarily. Safety features like censorship may be layered on top of the core model and are not automatically inherited during distillation.

What are the implications for open-source AI development?

If safety features are not embedded in the core parameters, open-source models could be less restricted, which raises both opportunities for transparency and concerns about misuse.

Can safety filters be added after distillation?

Yes, safety layers can potentially be added post-training or layered on top, but whether this can fully replace embedded safety features remains under study.

Is the performance of distilled models comparable to original models?

In this case, the distilled GPT-OSS-120B scored 83.61% on a finance task benchmark within an 8,000 token context, indicating strong task performance despite the lack of inherited censorship features.

Source: hn

You May Also Like

Muse Spark 1.1

Meta has published the evaluation report for Muse Spark 1.1, detailing its capabilities and performance, marking a significant update in AI language models.

The Six Chokepoints: How AI Stopped Being a Utility and Became a Lever

2026 marked a turning point as AI control shifted from open utility to concentrated leverage, with key chokepoints now in the hands of few entities.

Our Position On Open-weights Models

Tech company releases official stance on open-weights models amid industry debate, emphasizing transparency and safety concerns.

Technology Operations Signal Monitor: Apple’s New SpeechAnalyzer API, Benchmarked Against Whisper And Its Predecessor

Apple releases SpeechAnalyzer API; early benchmarking against Whisper shows promising performance, impacting small software company decision-making.