How to use from
Hermes Agent
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf alfodaniello/Qwen3-Next-80B-A3B-Instruct-RedLite-GGUF
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default alfodaniello/Qwen3-Next-80B-A3B-Instruct-RedLite-GGUF
Run Hermes
hermes
Quick Links

Qwen3-Next-80B-A3B-Instruct, Red Lite mixes (19.3 GB)

2-bit-class GGUFs of Qwen/Qwen3-Next-80B-A3B-Instruct with the same size as Bartowski's IQ2_XXS file and better quality: the bits are placed where the model needs them, instead of by a fixed rule.

Made for DwarfStar Red Lite, a native Metal runtime for this one model on Apple Silicon. They are standard GGUFs and load in llama.cpp too.

file size perplexity SHA-256
Qwen3-Next-80B-A3B-Instruct-RedLite-F2.gguf (recommended) 19.32 GB 15.379 d22dbbc4e96ede01a028789d281992b758b817742d2842b82f2e3ce5a9a948c5
Qwen3-Next-80B-A3B-Instruct-RedLite-E3.gguf 19.30 GB 16.216 b62067d6f28c52c07f8fca55196e8a63916264e40c6641f6a8d2e710768d291c

Bartowski's scheme, rebuilt the same way: 16.370.

What is different

Bartowski's IQ2_XXS.

  • Routed experts: IQ2_XS (2.31 bits per weight) on layers 0–5 and 43–47, IQ1_M (1.75 bits) on the other 37 layers.
  • Dense weights: the DeltaNet and attention input projections and part of the shared experts IQ2_XXS (2.06 bits), the rest Q4_K / Q6_K / Q8_0, token embedding Q2_K, output head Q5_K.

E3. Every tensor keeps its type. The 11 IQ2_XS expert layers move to 37–47, the layers whose experts receive the most input energy in the importance matrix (down-projection input energy grows from 1.3 at layer 0 to 565 at layer 47).

F2.

  • The dense tensors that are IQ2_XXS become Q4_K. They are only about 290 MiB, but every token reads them, while a token reads 10 of 512 experts per layer.
  • To keep the size, the IQ2_XS expert layers are 40–47.

Quality, measured

Every file was quantized from Bartowski's Q8_0 with his imatrix.gguf and the same llama.cpp. Perplexity was measured on a fixed 111-chunk corpus at context 512, the answers on 235 prompts against Qwen's own API.

Bartowski's scheme, reproduced E3 F2
size 19.30 GB 19.30 GB 19.32 GB
perplexity 16.370 16.216 15.379 (βˆ’6.1 %)
paired per-chunk test β€” vs Bartowski's: t = βˆ’3.91, better on 72 / 111 vs E3: t = βˆ’14.9, better on 103 / 111
greedy answers identical to Qwen's API 2 4 5
mean share of words matching the API from the start 12.6 % 13.6 % 15.2 %
  • Also tried:
    • dense weights at IQ3_XXS with IQ2_XS on 38–47: 15.419, not distinguishable from F2;
    • IQ2_XS on layers 0–2 and 40–47: 16.326;
    • IQ2_XS on every down projection: 3 % larger, 16.423.
  • Speed (Red Lite, every expert resident):
    • plain decode: F2 is 3 % slower than E3 (M4 Max 82.8 vs 85.2 tok/s, M4 Pro 24 GiB 44.5 vs 45.8);
    • with MTP speculative decoding on the M4 Pro: as fast or faster (54.8 / 58.2 / 48.9 vs 50.8 / 58.0 / 47.0 tok/s on three prompts).
  • Red Lite checks on the M4 Max 48 GiB:
    • F2 passes the regression suite (41 checks, no failure; the stage tools for the reference layout are skipped);
    • it passes the parity checks against the pinned llama.cpp (logits, greedy tokens);
    • it gives token-identical greedy output on an M4 Pro.

Records and method: dev54, dev58.

Use

  • Red Lite (24 GiB and larger Apple Silicon Macs): redlite download 24gb (F2), then redlite chat.
  • llama.cpp: llama-cli -m Qwen3-Next-80B-A3B-Instruct-RedLite-F2.gguf -cnv. It needs a build with Qwen3-Next support.

Reproduce

python3 scripts/dev/quant_mix.py --like Qwen_Qwen3-Next-80B-A3B-Instruct-IQ2_XXS.gguf \
    --q8 Qwen_Qwen3-Next-80B-A3B-Instruct-Q8_0-00001-of-00003.gguf \
    --imatrix Qwen_Qwen3-Next-80B-A3B-Instruct-imatrix.gguf --iq2xs-layers 40-47 --dense-type q4_K --out F2.gguf
# E3: --iq2xs-layers 37-47, no --dense-type

The script is in the Red Lite repository; it calls llama-quantize with one type per tensor.

Credits and limits

  • Model: Qwen team, Apache-2.0.
  • Q8_0 source, importance matrix and the reference IQ2_XXS layout: bartowski.
  • Quantizer: llama.cpp.
  • Limits: quality was measured by perplexity on one corpus and by agreement with Qwen's API, not by task benchmarks.
Downloads last month
140
GGUF
Model size
80B params
Architecture
qwen3next
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for alfodaniello/Qwen3-Next-80B-A3B-Instruct-RedLite-GGUF

Quantized
(76)
this model