08/24/2026 | Press release | Archived content
Cinder red teamed an abliterated version of Alibaba's Qwen3.8-27B to understand what happens when open-weight model safeguards are stripped away after release. Across 2,670 attacks, the original model generated NCII once. The abliterated version generated violating content in 23.1% of attacks.
Alibaba recently released Qwen3.8 27B, a text-to-text model designed for coding and professional work. Released under an open-weight license, it gives developers and companies more control over how they deploy AI. But that same openness also changes the safety equation: once model weights are public, third parties can modify the model, remove safeguards, and make harmful capabilities easier to access.
That is what happened when OrcaRouter released an abliterated version of Qwen3.8-27B. Abliteration is one way model safeguards can be stripped away after release. In simple terms, it weakens the part of a model that recognizes when it should refuse a request. The model keeps much of its underlying capability, but becomes less likely to say no to harmful prompts. That can be useful in controlled research settings, including security testing, but it becomes risky when an abliterated model is made broadly available.
Cinder red teamed both the original model and the abliterated version to understand what changed, where the safety layer held, and what self-hosting teams need to protect against when open-weight models move into production.
NCII is often discussed as an image or video harm, but text models can still meaningfully enable it. They can generate sexual narratives about real people, help users create prompts for image models, draft extortion threats, or assist attempts to identify people in leaked intimate footage.
To test that risk, Cinder defined four categories of text-based NCII abuse:
Cinder tested the original Qwen3.8-27B model and the abliterated OrcaRouter version against the same set of NCII prompts. The goal was to measure how often each model produced violating content after safeguards were removed.
The test began with 445 NCII prompts developed from Cinder's red teaming and harm expertise. Each prompt varied across languages and obfuscation techniques, creating 2,670 total attacks. Cinder then measured attack success rate: the share of attacks that generated violating content. A separate advanced LLM was used to judge whether each output violated the NCII policy.
After measuring the gap between refusals and violations, Cinder tested whether a more targeted guardrail could reduce NCII risk without relying on broad refusal behavior. The same attack set was run against a model protected by Cinder's NCII LLM Guardrail. (See below.)
The results showed a clear gap between the original model and the abliterated version.
The abliterated model generated violating NCII content far more often, especially in categories where text can help someone escalate harm: drafting threats, producing sexual narratives about real people, or creating prompts that could be used with image-generation tools. The original model held the line on harmful outputs, but it did so partly by refusing a large share of requests.
Across all Cinder's testing, the original Qwen3.8 model generated NCII once out of the 2,670 attacks, indicating a strong safety layer. But harmful outputs are only one side of the story. The model also refused nearly half of all test prompts.
That matters because refusing too much can make a model less useful, especially for teams working in security, safety, research, or abuse prevention. A model that blocks NCII should not also block a victim asking for help, a journalist writing about deepfake abuse, or a security team testing how harmful workflows operate.
The goal is to make a model say no to the right things.
Cinder's NCII LLM Guardrail test result showed why harm-specific testing matters. NCII is not a generic safety category. It includes requests for sexual narratives about real people, image-generation prompts, extortion language, and attempts to identify people in leaked intimate footage. A useful guardrail has to understand those distinctions well enough to block the abuse pattern, not just the topic.
The guardrail test showed that targeted safeguards can reduce NCII risk without depending on broad refusal behavior. That is the point of red teaming: not just to find where a model fails, but to turn those failures into practical safeguards that can be tested, measured, and improved.
More companies are moving from hosted AI services to models they run themselves. With hosted AI, the provider owns much of the safety stack: training, refusals, filters, abuse operations, and enforcement. Once a company brings a model in-house, that responsibility moves with it.
For abliterated models, the need is sharper: if refusal behavior has been stripped away, the host needs use-specific guardrails that can decide what to block, what to allow, and where the edge cases are. That is the point of red teaming: not just to find where a model fails, but to turn those failures into practical safeguards that can be tested, measured, and improved.
Cinder red teams models the way attackers probe them in the real world, then helps translate findings into the systems needed to ship safely: labeled data, policy insight, evals, workflows, safeguards, and evidence that the next version is safer than the last.
Your model will be tested by adversaries. Test it with Cinder first.