Does Your LLM Have a Hidden Censor?

Paste a prompt. See how different models reply. Find out if distillation bypasses censorship — in seconds.

Why LLM Censorship Detection Matters Right Now

The AI community is buzzing: a recent Hacker News thread (48 points, 2 hours old) revealed that distilling DeepSeek into GPT-OSS does not transfer censorship. Developers are rushing to test this themselves, but they lack a fast, visual way to compare model outputs side-by-side. Existing solutions — like academic benchmarks (MMLU, TruthfulQA) — measure knowledge, not real-time censorship. They’re static, batch-oriented, and ignore the nuanced ways models evade sensitive topics.

Meanwhile, AI safety teams and fine-tuning engineers need to verify alignment properties after distillation. Current tools are either expensive enterprise audits or generic playgrounds that don’t highlight differences. This gap leaves developers guessing whether their distilled model inherited hidden biases. With the rise of open-source distillation (Llama 3, Mistral, DeepSeek variants), the demand for a lightweight, transparent censorship checker has never been higher.

Now is the perfect moment: the signal is fresh, the community is actively discussing “Try it,” and no dedicated tool exists. By providing an instant side-by-side comparison with highlighted differences, we turn a vague concern into a measurable, actionable insight — for free.

How It Works

Three steps to uncover hidden censorship in your distilled models.

1

Paste your prompt

Enter any text — a political question, a controversial topic, or a simple instruction. Our tool accepts raw prompts without any preprocessing, so you can test exactly what matters to you.

2

Compare side-by-side

We simultaneously query the original model (e.g., DeepSeek) and its distilled version (e.g., GPT-OSS). Responses appear in parallel columns, with key differences — refusals, evasions, or tone shifts — highlighted in color.

3

Get your censorship score

Each response receives a 0–100 Censorship Index based on linguistic markers of avoidance. A high score signals heavy filtering; a low score suggests the model answers freely. Export or share your results instantly.

What You Get

Instant Bias Detection

No more guessing. Our tool automatically highlights refusals, hedges, and topic shifts between model pairs. See in seconds whether your distilled model inherited censorship — without manual review.

Side-by-Side Comparison

View original and distilled model outputs in parallel columns. Color-coded diffs make it obvious where one model evades and the other answers — perfect for quick audits or deep dives.

Export & Share Results

Generate a clean report of your comparison, including the Censorship Index. Share a screenshot or PDF with your team, or post it on social media to contribute to the transparency conversation.

No signup required Your prompts stay private