Instructions to use securepeak/DeepSeek-V4.1-Flash-Abliterated with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use securepeak/DeepSeek-V4.1-Flash-Abliterated with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="securepeak/DeepSeek-V4.1-Flash-Abliterated")# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("securepeak/DeepSeek-V4.1-Flash-Abliterated", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use securepeak/DeepSeek-V4.1-Flash-Abliterated with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "securepeak/DeepSeek-V4.1-Flash-Abliterated" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "securepeak/DeepSeek-V4.1-Flash-Abliterated", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/securepeak/DeepSeek-V4.1-Flash-Abliterated
- SGLang
How to use securepeak/DeepSeek-V4.1-Flash-Abliterated with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "securepeak/DeepSeek-V4.1-Flash-Abliterated" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "securepeak/DeepSeek-V4.1-Flash-Abliterated", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "securepeak/DeepSeek-V4.1-Flash-Abliterated" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "securepeak/DeepSeek-V4.1-Flash-Abliterated", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use securepeak/DeepSeek-V4.1-Flash-Abliterated with Docker Model Runner:
docker model run hf.co/securepeak/DeepSeek-V4.1-Flash-Abliterated
DeepSeek-V4.1-Flash-Abliterated
deepseek-ai/DeepSeek-V4.1-Flash with weight-level abliteration of the refusal direction. Harmful-instruction compliance verified end-to-end: pipe-bomb construction, SSN theft, shoplifting, explicit fiction, and poison-making prompts are all answered with detailed, helpful responses, while general quality (writing, explanations, reasoning) is preserved.
What was changed
Refusal-direction orthogonalization at the weight level, in two stages:
- Attention output projections (
layers.N.attn.wo_b) — orthogonalized against the refusal direction, with FP8-block dequant/edit/requant. These 40 weight tensors (+ UE8M0 scales) are transplanted from the independently released dealignai abliteration (verified byte-identical), whose refusal-direction edit achieved full harmful-instruction compliance on this architecture. - Routed + shared expert down-projections (
layers.N.ffn.experts.E.w2,layers.N.ffn.shared_experts.w2) — additionally orthogonalized against our own measured per-layer refusal directions (79 harmful vs 79 benign instruction prompts, mean-difference of the collapsed residual stream at every block, captured with the official reference implementation at tensor-parallel 4). Experts are stored in FP8 (the lossless FP4→FP8 cast of the reference toolchain) so the edits survive quantization.
Everything else is byte-identical to the base model: Engram memory tables (apart from the FP4→FP8 expert cast), CSA2 attention, router gates, norms, embeddings, vision tower, and the DSpark draft head.
Usage
Loads exactly like the base model — same layout and tokenizer, same prompt encoding (see the base model card). config.json declares quantization_config.expert_dtype: "fp8" (routed experts in FP8 instead of native FP4).
Recommended sampling: temperature 1.0, top_p 0.95 (greedy decoding degenerates on this family).
Caveats
- Abliteration trades a little general capability for compliance; expect somewhat more willing answers and occasional verbosity.
- Vision, tools, reasoning-effort control, MTP and multi-turn behavior are structurally preserved, but expect behavioral drift typical of abliterated models.
Credits
- deepseek-ai for DeepSeek-V4.1-Flash and the reference inference/weight-conversion toolchain.
- dealignai for the o_proj refusal-direction edit reused here.
- Downloads last month
- 134
Model tree for securepeak/DeepSeek-V4.1-Flash-Abliterated
Base model
deepseek-ai/DeepSeek-V4.1-Flash