Instructions to use soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf # Run inference directly in the terminal: llama cli -hf soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf # Run inference directly in the terminal: llama cli -hf soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf # Run inference directly in the terminal: ./llama-cli -hf soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf # Run inference directly in the terminal: ./build/bin/llama-cli -hf soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
Use Docker
docker model run hf.co/soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
- LM Studio
- Jan
- vLLM
How to use soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
- Ollama
How to use soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf with Ollama:
ollama run hf.co/soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
- Unsloth Desktop
- Pi
How to use soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf with Docker Model Runner:
docker model run hf.co/soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
- Lemonade
How to use soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
Run and chat with the model
lemonade run user.qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8-27B Abliterated 3.69bpw 12GB MTP GGUF
A Ridge-style mixed-quantized GGUF of the BF16 AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 model.
This model keeps the AEON-7 uncensored / refusal-removed weights, while using a Gated-DeltaNet-aware mixed quantization layout inspired by empero-ai/Qwen3.8-27B-Ridge-GGUF. The goal is to retain as much quality as possible in the sensitive Gated-DeltaNet path while obtaining a small, fast GGUF suitable for local llama.cpp inference.
Important: This is not an official Empero Ridge release and is not the same set of weights. It is an independent quantization of AEON-7 using a Ridge-inspired tensor-type map and an AEON-specific importance matrix.
Files
qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
| Property | Value |
|---|---|
| Architecture | Qwen3.5 / Qwen3.8 hybrid Gated-DeltaNet + full attention |
| Parameters | 27B |
| Format | GGUF |
| Nominal quantization | 3.69 bpw |
| File size | 12,599,187,808 bytes (~11.73 GiB) |
| Context metadata | 262,144 tokens |
| MTP | Preserved in the GGUF; native draft-mtp supported |
| Vision | Text GGUF only; use a compatible mmproj if image support is required |
| License | Apache-2.0, inherited from the base model |
Quantization layout
The quantization map is mixed rather than a flat IQ2 dump.
| GGML type | Tensor count | Main purpose |
|---|---|---|
F32 |
360 | norms and scalar/state tensors |
Q4_K |
144 | Gated-DeltaNet mixer/projection tensors |
Q8_0 |
96 | sensitive Gated-DeltaNet state path (ssm_alpha / ssm_beta) |
IQ2_S |
160 | mid-stack FFN weights |
IQ3_S |
32 | selected FFN weights kept at higher precision |
Q5_K |
51 | full-attention Q/K/V tensors |
Q6_K |
23 | output/embedding tensors, full-attention output, and MTP tensors |
The MTP tensors have no importance matrix and are kept at Q6_K, following the
important design choice documented by the Ridge project.
Calibration
The AEON-specific importance matrix was generated locally from a calibration corpus containing English WikiText, Japanese Wikipedia extracts, and llama.cpp source code.
context length: 512
batch size: 512
chunks: 80
process output: enabled
importance entries: 497
The calibration corpus and quantization map are not the private calibration artifacts used by the original Ridge release. They are an independent local reproduction of the same general quantization strategy.
llama.cpp usage
The file name intentionally matches the repository name.
./llama-server \
-m ./qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf \
--spec-type draft-mtp \
--spec-draft-n-max 3 \
--split-mode layer \
--tensor-split 2,1 \
--host 0.0.0.0 \
--port 8080 \
-ctk q4_0 \
-ctv q4_0 \
-c 114514
For a local-only server, prefer binding to localhost instead:
--host 127.0.0.1
--host 0.0.0.0 exposes the server to the network. Add your own
authentication, firewall, and access controls before exposing it beyond a
trusted LAN.
To disable thinking when using a compatible client, use the Qwen3.8 chat options supported by your frontend or API client.
Reported local performance
The model was prepared and tested on:
OS: Ubuntu 24.04
CPU: Intel Core i7-10700K
GPU 0: NVIDIA GeForce RTX 5060 Ti 16GB
GPU 1: NVIDIA GeForce RTX 3070 8GB
RAM: 32GB
Runtime: llama.cpp CUDA build
On this machine, the command above has reached a reported peak of up to approximately 37 tokens/second. Actual speed depends on context length, prompt length, sampling settings, MTP acceptance rate, CUDA/llama.cpp version, background workload, and the amount of KV cache in use.
For this hardware, --spec-draft-n-max 3 is recommended as a starting point.
Higher draft counts can add overhead rather than improve throughput.
Provenance
Base model
- AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
- Base revision used locally:
a6775a9a8ebb65cab3f707b4ab087fc7aa698634 - Original upstream model: Qwen/Qwen3.8-27B
AEON-7 is an abliterated BF16 derivative. Its model card describes an SSM conv1d outlier-repair step, abliterix processing, a stock MTP head graft, and an untouched vision tower. This GGUF quantizes the AEON-7 BF16 weights; it is not a re-quantization of an already-quantized FP8 checkpoint.
Quantization idea
The mixed quantization strategy was inspired by:
The local conversion and quantization used llama.cpp commit
030ebb558.
AI assistance disclosure
The local model preparation workflow, conversion, calibration-data preparation, quantization, validation, and this model card were performed with assistance from GPT-5.6-Luna via Hermes Agent. The model was then reviewed and published by the repository owner.
Responsible use
This is an uncensored / refusal-removed model. It may produce content that an aligned model would refuse, including unsafe, illegal, or harmful material. It has no reliable built-in safety layer. Use appropriate access controls, moderation, logging, and human review for any deployment, and comply with all applicable laws and policies.
The model is provided as-is. Users are responsible for prompts, outputs, and any downstream actions based on them.
ๆฅๆฌ่ช
ๆฆ่ฆ
ใใใฏใ AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 ใใใผในใซใGated-DeltaNetใฎ็นๆงใ่ๆ ฎใใRidgeๅผใฎๆททๅ้ๅญๅใ้ฉ็จใใ GGUFใขใใซใงใใ
็กๆค้ฒใปๆๅฆ้คๅปๆธใฟใฎAEON-7ใฎ้ใฟใ็ถญๆใใชใใใ empero-ai/Qwen3.8-27B-Ridge-GGUF ใงไฝฟใใใฆใใ่ใๆนใๅ่ใซใGDNใฎๅฃใใใใ็ต่ทฏใธใใใใๅชๅ ้ ๅใใฆใใพใใ
ใใใฏEmperoๅ ฌๅผใฎRidgeใขใใซใงใฏใชใใAEON-7ใ็ฌ่ชใซ้ๅญๅใใๆดพ็GGUFใงใใ
ๅบๆฌไปๆง
ใใกใคใซๅ: qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
ๅฝขๅผ: GGUF
ใตใคใบ: 12,599,187,808 bytes๏ผ็ด11.73 GiB๏ผ
้ๅญๅ: 3.69 bpw
ใใฉใกใผใฟ: 27B
MTP: GGUFๅ
ใซไฟๆใnative draft-mtpๅฏพๅฟ
ๆททๅ้ๅญๅใฎๅ ่จณ
F32 360 tensors norm / state็ณป
Q4_K 144 tensors Gated-DeltaNet mixer็ณป
Q8_0 96 tensors ssm_alpha / ssm_beta
IQ2_S 160 tensors ไธญ้ๅฑคFFN
IQ3_S 32 tensors ้ซใใฎ็ฒพๅบฆใๆฎใใFFN
Q5_K 51 tensors ้ๅธธAttentionใฎQ/K/V
Q6_K 23 tensors outputใembeddingใAttention outputใMTP
MTPใใณใฝใซใซใฏimatrixใ้ฉ็จใใใQ6_Kใงไฟๆใใฆใใพใใ
AEONๅฐ็จใฎimatrixใฏใไปฅไธใๆททใใcalibrationใใผใฟใใไฝๆใใพใใใ
- ่ฑ่ชWikiText
- ๆฅๆฌ่ชWikipedia
- llama.cppใฎC/C++/Pythonใฝใผในใณใผใ
context: 512
batch: 512
chunks: 80
entries: 497
llama.cppใงใฎ่ตทๅ
./llama-server \
-m ./qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf \
--spec-type draft-mtp \
--spec-draft-n-max 3 \
--split-mode layer \
--tensor-split 2,1 \
--host 0.0.0.0 \
--port 8080 \
-ctk q4_0 \
-ctv q4_0 \
-c 114514
ใญใผใซใซใใใฎใฟๆฅ็ถใใๅ ดๅใฏใ--host 127.0.0.1ใๆจๅฅจใใพใใ
0.0.0.0ใฏใใใใฏใผใฏไธใธๅ
ฌ้ใใ่จญๅฎใชใฎใงใ่ช่จผใปFirewallใปใขใฏใปในๅถๅพกใ
ๅฟ
ใ่ฟฝๅ ใใฆใใ ใใใ
ไฝๆใปๆค่จผ็ฐๅข
Ubuntu 24.04
Intel Core i7-10700K
RTX 5060 Ti 16GB + RTX 3070 8GB
RAM 32GB
llama.cpp CUDA build
ไธ่จใฎ็ฐๅขใจ่ตทๅ่จญๅฎใงใๆ้ซ็ด37 tokens/secondใๅ ฑๅใใใฆใใพใใ ๅฎ้ใฎ้ๅบฆใฏใใณใณใใญในใ้ทใใใญใณใใ้ทใKV cacheใMTPใฎๅ็็ใ llama.cppใฎใใผใธใงใณใใใใฏใฐใฉใฆใณใ่ฒ ่ทใชใฉใงๅคๅใใพใใ
AIๅฉ็จใฎ้็คบ
ใญใผใซใซใธใฎใขใใซๅๅพใๅคๆใcalibrationใใผใฟไฝๆใimatrixไฝๆใ้ๅญๅใ ๅไฝ็ขบ่ชใใใใณใใฎREADMEใฎไฝๆใฏใHermes Agent็ต็ฑใฎGPT-5.6-Lunaใฎ ๆฏๆดใๅใใฆ่กใใใพใใใๆ็ต็ใช็ขบ่ชใจๅ ฌ้ใฏใชใใธใใชๆๆ่ ใ่กใฃใฆใใพใใ
ๅฉ็จไธใฎๆณจๆ
ใใฎใขใใซใฏ็กๆค้ฒใปๆๅฆ้คๅปๆธใฟใงใใ้ๅธธใฎใขใฉใคใณๆธใฟใขใใซใๆๅฆใใใใใชใ ๅฑ้บใป้ๆณใปๆๅฎณใชๅ ๅฎนใๅบๅใใๅฏ่ฝๆงใใใใพใใไฟก้ ผใงใใๅฎๅ จๆฉๆงใๅ ่ตใใฆ ใใใจใฏ่ใใชใใงใใ ใใใ
ๅ ฌ้้็จใใๅ ดๅใฏใใขใฏใปในๅถๅพกใๅ ฅๅใปๅบๅใใฃใซใฟใ็ฃๆปใญใฐใใฌใผใๅถ้ใ ไบบ้ใซใใ็ขบ่ชใชใฉใ็จ้ใซๅฟใใฆๅฎ่ฃ ใใฆใใ ใใใๅฉ็จ่ ใฏใใญใณใใใๅบๅใ ๅบๅใๅฉ็จใใ downstream ใฎ่ก็บใซใคใใฆ่ฒฌไปปใ่ฒ ใใพใใ
- Downloads last month
- 30,366
We're not able to determine the quantization variants.
Model tree for soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
Base model
Qwen/Qwen3.8-27B