Qwen3.8-27B Abliterated 3.69bpw 12GB MTP GGUF

A Ridge-style mixed-quantized GGUF of the BF16 AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 model.

This model keeps the AEON-7 uncensored / refusal-removed weights, while using a Gated-DeltaNet-aware mixed quantization layout inspired by empero-ai/Qwen3.8-27B-Ridge-GGUF. The goal is to retain as much quality as possible in the sensitive Gated-DeltaNet path while obtaining a small, fast GGUF suitable for local llama.cpp inference.

Important: This is not an official Empero Ridge release and is not the same set of weights. It is an independent quantization of AEON-7 using a Ridge-inspired tensor-type map and an AEON-specific importance matrix.


Files

qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
Property Value
Architecture Qwen3.5 / Qwen3.8 hybrid Gated-DeltaNet + full attention
Parameters 27B
Format GGUF
Nominal quantization 3.69 bpw
File size 12,599,187,808 bytes (~11.73 GiB)
Context metadata 262,144 tokens
MTP Preserved in the GGUF; native draft-mtp supported
Vision Text GGUF only; use a compatible mmproj if image support is required
License Apache-2.0, inherited from the base model

Quantization layout

The quantization map is mixed rather than a flat IQ2 dump.

GGML type Tensor count Main purpose
F32 360 norms and scalar/state tensors
Q4_K 144 Gated-DeltaNet mixer/projection tensors
Q8_0 96 sensitive Gated-DeltaNet state path (ssm_alpha / ssm_beta)
IQ2_S 160 mid-stack FFN weights
IQ3_S 32 selected FFN weights kept at higher precision
Q5_K 51 full-attention Q/K/V tensors
Q6_K 23 output/embedding tensors, full-attention output, and MTP tensors

The MTP tensors have no importance matrix and are kept at Q6_K, following the important design choice documented by the Ridge project.

Calibration

The AEON-specific importance matrix was generated locally from a calibration corpus containing English WikiText, Japanese Wikipedia extracts, and llama.cpp source code.

context length: 512
batch size:     512
chunks:         80
process output: enabled
importance entries: 497

The calibration corpus and quantization map are not the private calibration artifacts used by the original Ridge release. They are an independent local reproduction of the same general quantization strategy.


llama.cpp usage

The file name intentionally matches the repository name.

./llama-server \
  -m ./qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf \
  --spec-type draft-mtp \
  --spec-draft-n-max 3 \
  --split-mode layer \
  --tensor-split 2,1 \
  --host 0.0.0.0 \
  --port 8080 \
  -ctk q4_0 \
  -ctv q4_0 \
  -c 114514

For a local-only server, prefer binding to localhost instead:

--host 127.0.0.1

--host 0.0.0.0 exposes the server to the network. Add your own authentication, firewall, and access controls before exposing it beyond a trusted LAN.

To disable thinking when using a compatible client, use the Qwen3.8 chat options supported by your frontend or API client.


Reported local performance

The model was prepared and tested on:

OS:       Ubuntu 24.04
CPU:      Intel Core i7-10700K
GPU 0:    NVIDIA GeForce RTX 5060 Ti 16GB
GPU 1:    NVIDIA GeForce RTX 3070 8GB
RAM:      32GB
Runtime:  llama.cpp CUDA build

On this machine, the command above has reached a reported peak of up to approximately 37 tokens/second. Actual speed depends on context length, prompt length, sampling settings, MTP acceptance rate, CUDA/llama.cpp version, background workload, and the amount of KV cache in use.

For this hardware, --spec-draft-n-max 3 is recommended as a starting point. Higher draft counts can add overhead rather than improve throughput.


Provenance

Base model

AEON-7 is an abliterated BF16 derivative. Its model card describes an SSM conv1d outlier-repair step, abliterix processing, a stock MTP head graft, and an untouched vision tower. This GGUF quantizes the AEON-7 BF16 weights; it is not a re-quantization of an already-quantized FP8 checkpoint.

Quantization idea

The mixed quantization strategy was inspired by:

The local conversion and quantization used llama.cpp commit 030ebb558.


AI assistance disclosure

The local model preparation workflow, conversion, calibration-data preparation, quantization, validation, and this model card were performed with assistance from GPT-5.6-Luna via Hermes Agent. The model was then reviewed and published by the repository owner.


Responsible use

This is an uncensored / refusal-removed model. It may produce content that an aligned model would refuse, including unsafe, illegal, or harmful material. It has no reliable built-in safety layer. Use appropriate access controls, moderation, logging, and human review for any deployment, and comply with all applicable laws and policies.

The model is provided as-is. Users are responsible for prompts, outputs, and any downstream actions based on them.


ๆ—ฅๆœฌ่ชž

ๆฆ‚่ฆ

ใ“ใ‚Œใฏใ€ AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 ใ‚’ใƒ™ใƒผใ‚นใซใ€Gated-DeltaNetใฎ็‰นๆ€งใ‚’่€ƒๆ…ฎใ—ใŸRidgeๅผใฎๆททๅˆ้‡ๅญๅŒ–ใ‚’้ฉ็”จใ—ใŸ GGUFใƒขใƒ‡ใƒซใงใ™ใ€‚

็„กๆคœ้–ฒใƒปๆ‹’ๅฆ้™คๅŽปๆธˆใฟใฎAEON-7ใฎ้‡ใฟใ‚’็ถญๆŒใ—ใชใŒใ‚‰ใ€ empero-ai/Qwen3.8-27B-Ridge-GGUF ใงไฝฟใ‚ใ‚Œใฆใ„ใ‚‹่€ƒใˆๆ–นใ‚’ๅ‚่€ƒใซใ€GDNใฎๅฃŠใ‚Œใ‚„ใ™ใ„็ตŒ่ทฏใธใƒ“ใƒƒใƒˆใ‚’ๅ„ชๅ…ˆ้…ๅˆ†ใ—ใฆใ„ใพใ™ใ€‚

ใ“ใ‚ŒใฏEmperoๅ…ฌๅผใฎRidgeใƒขใƒ‡ใƒซใงใฏใชใใ€AEON-7ใ‚’็‹ฌ่‡ชใซ้‡ๅญๅŒ–ใ—ใŸๆดพ็”ŸGGUFใงใ™ใ€‚

ๅŸบๆœฌไป•ๆง˜

ใƒ•ใ‚กใ‚คใƒซๅ: qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf
ๅฝขๅผ:       GGUF
ใ‚ตใ‚คใ‚บ:     12,599,187,808 bytes๏ผˆ็ด„11.73 GiB๏ผ‰
้‡ๅญๅŒ–:     3.69 bpw
ใƒ‘ใƒฉใƒกใƒผใ‚ฟ: 27B
MTP:        GGUFๅ†…ใซไฟๆŒใ€native draft-mtpๅฏพๅฟœ

ๆททๅˆ้‡ๅญๅŒ–ใฎๅ†…่จณ

F32    360 tensors  norm / state็ณป
Q4_K   144 tensors  Gated-DeltaNet mixer็ณป
Q8_0    96 tensors  ssm_alpha / ssm_beta
IQ2_S  160 tensors  ไธญ้–“ๅฑคFFN
IQ3_S   32 tensors  ้ซ˜ใ‚ใฎ็ฒพๅบฆใ‚’ๆฎ‹ใ—ใŸFFN
Q5_K   51 tensors  ้€šๅธธAttentionใฎQ/K/V
Q6_K   23 tensors  outputใ€embeddingใ€Attention outputใ€MTP

MTPใƒ†ใƒณใ‚ฝใƒซใซใฏimatrixใ‚’้ฉ็”จใ›ใšใ€Q6_KใงไฟๆŒใ—ใฆใ„ใพใ™ใ€‚

AEONๅฐ‚็”จใฎimatrixใฏใ€ไปฅไธ‹ใ‚’ๆททใœใŸcalibrationใƒ‡ใƒผใ‚ฟใ‹ใ‚‰ไฝœๆˆใ—ใพใ—ใŸใ€‚

  • ่‹ฑ่ชžWikiText
  • ๆ—ฅๆœฌ่ชžWikipedia
  • llama.cppใฎC/C++/Pythonใ‚ฝใƒผใ‚นใ‚ณใƒผใƒ‰
context: 512
batch:   512
chunks:  80
entries: 497

llama.cppใงใฎ่ตทๅ‹•

./llama-server \
  -m ./qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf \
  --spec-type draft-mtp \
  --spec-draft-n-max 3 \
  --split-mode layer \
  --tensor-split 2,1 \
  --host 0.0.0.0 \
  --port 8080 \
  -ctk q4_0 \
  -ctv q4_0 \
  -c 114514

ใƒญใƒผใ‚ซใƒซใ‹ใ‚‰ใฎใฟๆŽฅ็ถšใ™ใ‚‹ๅ ดๅˆใฏใ€--host 127.0.0.1ใ‚’ๆŽจๅฅจใ—ใพใ™ใ€‚ 0.0.0.0ใฏใƒใƒƒใƒˆใƒฏใƒผใ‚ฏไธŠใธๅ…ฌ้–‹ใ™ใ‚‹่จญๅฎšใชใฎใงใ€่ช่จผใƒปFirewallใƒปใ‚ขใ‚ฏใ‚ปใ‚นๅˆถๅพกใ‚’ ๅฟ…ใš่ฟฝๅŠ ใ—ใฆใใ ใ•ใ„ใ€‚

ไฝœๆˆใƒปๆคœ่จผ็’ฐๅขƒ

Ubuntu 24.04
Intel Core i7-10700K
RTX 5060 Ti 16GB + RTX 3070 8GB
RAM 32GB
llama.cpp CUDA build

ไธŠ่จ˜ใฎ็’ฐๅขƒใจ่ตทๅ‹•่จญๅฎšใงใ€ๆœ€้ซ˜็ด„37 tokens/secondใŒๅ ฑๅ‘Šใ•ใ‚Œใฆใ„ใพใ™ใ€‚ ๅฎŸ้š›ใฎ้€Ÿๅบฆใฏใ€ใ‚ณใƒณใƒ†ใ‚ญใ‚นใƒˆ้•ทใ€ใƒ—ใƒญใƒณใƒ—ใƒˆ้•ทใ€KV cacheใ€MTPใฎๅ—็†็އใ€ llama.cppใฎใƒใƒผใ‚ธใƒงใƒณใ€ใƒใƒƒใ‚ฏใ‚ฐใƒฉใ‚ฆใƒณใƒ‰่ฒ ่ทใชใฉใงๅค‰ๅŒ–ใ—ใพใ™ใ€‚

AIๅˆฉ็”จใฎ้–‹็คบ

ใƒญใƒผใ‚ซใƒซใธใฎใƒขใƒ‡ใƒซๅ–ๅพ—ใ€ๅค‰ๆ›ใ€calibrationใƒ‡ใƒผใ‚ฟไฝœๆˆใ€imatrixไฝœๆˆใ€้‡ๅญๅŒ–ใ€ ๅ‹•ไฝœ็ขบ่ชใ€ใŠใ‚ˆใณใ“ใฎREADMEใฎไฝœๆˆใฏใ€Hermes Agent็ตŒ็”ฑใฎGPT-5.6-Lunaใฎ ๆ”ฏๆดใ‚’ๅ—ใ‘ใฆ่กŒใ‚ใ‚Œใพใ—ใŸใ€‚ๆœ€็ต‚็š„ใช็ขบ่ชใจๅ…ฌ้–‹ใฏใƒชใƒใ‚ธใƒˆใƒชๆ‰€ๆœ‰่€…ใŒ่กŒใฃใฆใ„ใพใ™ใ€‚

ๅˆฉ็”จไธŠใฎๆณจๆ„

ใ“ใฎใƒขใƒ‡ใƒซใฏ็„กๆคœ้–ฒใƒปๆ‹’ๅฆ้™คๅŽปๆธˆใฟใงใ™ใ€‚้€šๅธธใฎใ‚ขใƒฉใ‚คใƒณๆธˆใฟใƒขใƒ‡ใƒซใŒๆ‹’ๅฆใ™ใ‚‹ใ‚ˆใ†ใชใ€ ๅฑ้™บใƒป้•ๆณ•ใƒปๆœ‰ๅฎณใชๅ†…ๅฎนใ‚’ๅ‡บๅŠ›ใ™ใ‚‹ๅฏ่ƒฝๆ€งใŒใ‚ใ‚Šใพใ™ใ€‚ไฟก้ ผใงใใ‚‹ๅฎ‰ๅ…จๆฉŸๆง‹ใ‚’ๅ†…่”ตใ—ใฆ ใ„ใ‚‹ใจใฏ่€ƒใˆใชใ„ใงใใ ใ•ใ„ใ€‚

ๅ…ฌ้–‹้‹็”จใ™ใ‚‹ๅ ดๅˆใฏใ€ใ‚ขใ‚ฏใ‚ปใ‚นๅˆถๅพกใ€ๅ…ฅๅŠ›ใƒปๅ‡บๅŠ›ใƒ•ใ‚ฃใƒซใ‚ฟใ€็›ฃๆŸปใƒญใ‚ฐใ€ใƒฌใƒผใƒˆๅˆถ้™ใ€ ไบบ้–“ใซใ‚ˆใ‚‹็ขบ่ชใชใฉใ‚’็”จ้€”ใซๅฟœใ˜ใฆๅฎŸ่ฃ…ใ—ใฆใใ ใ•ใ„ใ€‚ๅˆฉ็”จ่€…ใฏใƒ—ใƒญใƒณใƒ—ใƒˆใ€ๅ‡บๅŠ›ใ€ ๅ‡บๅŠ›ใ‚’ๅˆฉ็”จใ—ใŸ downstream ใฎ่กŒ็‚บใซใคใ„ใฆ่ฒฌไปปใ‚’่ฒ ใ„ใพใ™ใ€‚

Downloads last month
30,366
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ 1 Ask for provider support

Model tree for soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf

Base model

Qwen/Qwen3.8-27B
Quantized
(33)
this model

Space using soyaakinohara/qwen3.8-27b-abliterated-3.69bpw-12GB-MTP.gguf 1