Qwen3.8-27B abliterated model cuts refusal rates to near zero while preserving
An abliterated Qwen3.8-27B FP8 model reduces safety refusals from 64-99% to 0-6% while maintaining benchmark performance within 1.3 points, raising questions about the relationship between refusal mechanisms and core
A red-team build of Qwen3.8-27B FP8 demonstrates that safety refusals can be nearly eliminated without measurable degradation to core reasoning benchmarks. The abliterated checkpoint reduces refusal rates across AdvBench, HarmBench, and StrongREJECT from 64 to 99 percent down to 0 to 6 percent, whil...
Sign in to read the full analysis
Free account. Full analysis on LLM unit economics, plus the weekly Cost-of-Inference column.
Try it on your own context
You just read the writeup. Now run the thing. Paste a doc or some verbose tool output and watch it shrink — free, no signup.
- Source type
- Primary publication (lab/vendor blog) — our analysis + implication
- Source link
- r/localllama
- Published
- UTC
- Byline
- By the gotcontext.ai team (editorial standards)
- Correction?
- corrections@gotcontext.ai