Qwen3-Next-80B-A3B-Instruct, Red Lite expert mix (19.3 GB)

A 2-bit-class GGUF of Qwen/Qwen3-Next-80B-A3B-Instruct. It has the same size and the same tensor types as Bartowski's IQ2_XXS file, but better quality: the 11 expert layers kept at the higher precision (IQ2_XS) are chosen by the importance matrix instead of a fixed rule.

Made for DwarfStar Red Lite, a native Metal runtime for this one model on Apple Silicon. It is a standard GGUF and loads in llama.cpp too.

file size SHA-256
Qwen3-Next-80B-A3B-Instruct-RedLite-E3.gguf 19.30 GB b62067d6f28c52c07f8fca55196e8a63916264e40c6641f6a8d2e710768d291c

What is different

Bartowski's IQ2_XXS.

  • Routed experts: IQ2_XS (2.31 bits per weight) on layers 0–5 and 43–47, IQ1_M (1.75 bits) on the other 37 layers.
  • Everything else: attention and DeltaNet projections IQ2_XXS / Q4_K / Q6_K / Q8_0, shared experts Q6_K, token embedding Q2_K, output head Q5_K.

This file.

  • Every non-expert tensor has exactly the type it has in Bartowski's file.
  • The 11 IQ2_XS expert layers are 37–47 instead. Those are the layers whose experts receive the most input energy in the importance matrix: down-projection input energy grows from 1.3 at layer 0 to 565 at layer 47.

Same size, same speed: the types are the same, only their placement changes.

Quality, measured

Every variant was quantized from Bartowski's Q8_0 with his imatrix.gguf and the same llama.cpp. Perplexity was measured on a fixed 111-chunk corpus at context 512, the answers on 235 prompts against Qwen's own API.

Bartowski's scheme, reproduced this file
size 19.30 GB 19.30 GB
perplexity 16.370 16.216 (βˆ’0.94 %)
paired per-chunk test vs the reproduced scheme β€” t = βˆ’3.91, better on 72 / 111 chunks
greedy answers identical to Qwen's API (235 prompts) 2 4
mean share of words matching the API from the start 12.6 % 13.6 %

Two other mixes were tried and are not better:

  • IQ2_XS on layers 0–2 and 40–47: perplexity 16.326, t = βˆ’1.28;
  • IQ2_XS on every down projection, gate/up at IQ1_M: 3 % larger, perplexity 16.423.

On the M4 Max 48 GiB, with every expert resident, this file passes Red Lite's full regression suite (52/52) and its parity checks against the pinned llama.cpp (logits, greedy tokens, long context).

Records and method: docs/REDLITE_DEV54_QUANT_MIX.md.

Use

  • Red Lite (24 GiB and larger Apple Silicon Macs): put the file in models/ and run ./bin/redlite chat.
  • llama.cpp: llama-cli -m Qwen3-Next-80B-A3B-Instruct-RedLite-E3.gguf -cnv. It needs a build with Qwen3-Next support.

Reproduce

python3 scripts/dev/quant_mix.py --like Qwen_Qwen3-Next-80B-A3B-Instruct-IQ2_XXS.gguf \
    --q8 Qwen_Qwen3-Next-80B-A3B-Instruct-Q8_0-00001-of-00003.gguf \
    --imatrix Qwen_Qwen3-Next-80B-A3B-Instruct-imatrix.gguf --iq2xs-layers 37-47 --out E3.gguf

The script is in the Red Lite repository; it calls llama-quantize with one type per tensor.

Credits and limits

  • Model: Qwen team, Apache-2.0.
  • Q8_0 source, importance matrix and the reference IQ2_XXS layout: bartowski.
  • Quantizer: llama.cpp.
  • Limits: quality was measured by perplexity on one corpus and by agreement with Qwen's API, not by task benchmarks. Only four layer mixes were compared, at one size.
Downloads last month
-
GGUF
Model size
80B params
Architecture
qwen3next
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for alfodaniello/Qwen3-Next-80B-A3B-Instruct-RedLite-GGUF

Quantized
(76)
this model