Qwen3.8-27B QUASAR NVFP4 for NInfer

An all-NVFP4 NInfer artifact of Qwen3.8-27B, built from the quantization-aware-trained QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4 checkpoint, with the DFlash2 drafter included. Community build; not affiliated with the NInfer, Qwen, QUASAR or DFlash2 authors.

Versus the official NInfer artifact

Same RTX 5090 (600 W), NInfer build and serving flags as neroued/Qwen3.8-27B-nvfp4-NInfer:

Official This artifact
DFlash2 round time (decode) 16.1 ms 13.8 ms (≈ +16% tok/s)
Prefill, 16.7k-token prompt 11.3–11.5k tok/s 13.7k tok/s
Prose decode (same acceptance) 175.6 tok/s 204.2 tok/s
VRAM with the command below 30.8 GiB 27.0 GiB
GPQA-Diamond 90.40% (179/198) 89.90% (178/198)
Perplexity, ninfer-ppl-1m-v1 quick 4.3122 4.3955 (+1.9%)

Qwen's BF16 card reports 89.2 on GPQA-Diamond. With 198 samples (±2.1 points) the three GPQA figures are indistinguishable. The perplexity cost is largest on English reference text (+3.8%).

Run

Requires NInfer built from source (tested at d44ab584), Linux, an RTX 5090 and CUDA 13.1+.

hf download jesdga/Qwen3.8-27B-QUASAR-DFlash2-nvfp4-NInfer qwen3_8_27b_quasar.ninfer --local-dir models

./build/apps/ninfer-serve models/qwen3_8_27b_quasar.ninfer \
  --host 127.0.0.1 --port 8080 \
  --max-context 262144 --kv-capacity 262144 --default-max-tokens 262144 \
  --max-concurrency 1 --kv-dtype fp8 \
  --device-state-slots 2 --host-state-slots 8 --host-kv-mib 8192 \
  --spec dflash2 --draft-tokens 7 --lm-head-draft --preserve-thinking

Add --vision for images and video; it still fits the full 262,144-token KV pool.

qwen3_8_27b_quasar.ninfer: 19,624,516,612 bytes, SHA-256 5d683c403dbe121e230a6730cead81e0f9756387d528702d08f55aa14d1de159.

Contents

  • All 256 Text projections (MLP, GDN, attention): NVFP4 imported bit-exactly from QUASAR, including its activation scales, so prefill uses FP4 activations.
  • Embedding and output head: row-scaled FP8 from QUASAR's BF16 tensors.
  • GDN gating: BF16, decoded from QUASAR.
  • DFlash2, MTP, Vision, proposal head: the official artifact's formats.

How it was measured

Decode and prefill: one request at a time, thinking on, FP8 KV, prefill chunk 1,024, on fixed conversational, JSON, essay and 16.7k-token-document prompts. Round time is the comparable figure; tokens/s also depends on how many draft tokens each sampled response accepts. Official values are two server sessions, this artifact one; expect about ±2%.

GPQA-Diamond: NInfer's eval/ with EvalScope 1.10.0 and the official card's GPQA settings (thinking, temperature 1.0, top-p 0.95, top-k 20, seed 42, INT8 KV), using DFlash2 instead of MTP. Reports and configs are in eval/.

Not tested: MTP, image/video quality, and benchmarks other than GPQA-Diamond.

Rebuild

From an NInfer checkout at d44ab584, with QUASAR at revision 15d2e47bffe5d8ad23928879f8f7d2f74909e259 and DFlash2 at 50307d4c4cde6860d4eee73e2547cd786fe8e8a4:

python3 -m tools.convert \
  --model sources/Qwen3.8-27B-QUASAR-NVFP4 \
  --recipe qwen3_8_27b_quasar.py \
  --source dflash2=sources/Qwen3.8-27B-DFlash2 \
  --components text,vision,mtp,dflash2 \
  --resource chat_template.jinja=tools/chat_templates/qwen3_8.jinja \
  --proposal --name qwen3.8-27b --device cpu \
  --out models/qwen3_8_27b_quasar.ninfer

The recipe is qwen3_8_27b_quasar.py; the per-object record is conversion-report.json.

License

Apache-2.0, as are the Qwen3.8-27B base model, the QUASAR checkpoint and the DFlash2 drafter.

Downloads last month
194
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jesdga/Qwen3.8-27B-QUASAR-DFlash2-nvfp4-NInfer

Base model

Qwen/Qwen3.8-27B
Quantized
(10)
this model

Evaluation results