Instructions to use jesdga/Qwen3.8-27B-QUASAR-DFlash2-nvfp4-NInfer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NInfer
How to use jesdga/Qwen3.8-27B-QUASAR-DFlash2-nvfp4-NInfer with NInfer:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Qwen3.8-27B QUASAR NVFP4 for NInfer
An all-NVFP4 NInfer artifact of Qwen3.8-27B, built from the quantization-aware-trained QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4 checkpoint, with the DFlash2 drafter included. Community build; not affiliated with the NInfer, Qwen, QUASAR or DFlash2 authors.
Versus the official NInfer artifact
Same RTX 5090 (600 W), NInfer build and serving flags as neroued/Qwen3.8-27B-nvfp4-NInfer:
| Official | This artifact | |
|---|---|---|
| DFlash2 round time (decode) | 16.1 ms | 13.8 ms (≈ +16% tok/s) |
| Prefill, 16.7k-token prompt | 11.3–11.5k tok/s | 13.7k tok/s |
| Prose decode (same acceptance) | 175.6 tok/s | 204.2 tok/s |
| VRAM with the command below | 30.8 GiB | 27.0 GiB |
| GPQA-Diamond | 90.40% (179/198) | 89.90% (178/198) |
Perplexity, ninfer-ppl-1m-v1 quick |
4.3122 | 4.3955 (+1.9%) |
Qwen's BF16 card reports 89.2 on GPQA-Diamond. With 198 samples (±2.1 points) the three GPQA figures are indistinguishable. The perplexity cost is largest on English reference text (+3.8%).
Run
Requires NInfer built from source (tested at
d44ab584),
Linux, an RTX 5090 and CUDA 13.1+.
hf download jesdga/Qwen3.8-27B-QUASAR-DFlash2-nvfp4-NInfer qwen3_8_27b_quasar.ninfer --local-dir models
./build/apps/ninfer-serve models/qwen3_8_27b_quasar.ninfer \
--host 127.0.0.1 --port 8080 \
--max-context 262144 --kv-capacity 262144 --default-max-tokens 262144 \
--max-concurrency 1 --kv-dtype fp8 \
--device-state-slots 2 --host-state-slots 8 --host-kv-mib 8192 \
--spec dflash2 --draft-tokens 7 --lm-head-draft --preserve-thinking
Add --vision for images and video; it still fits the full 262,144-token KV pool.
qwen3_8_27b_quasar.ninfer: 19,624,516,612 bytes, SHA-256
5d683c403dbe121e230a6730cead81e0f9756387d528702d08f55aa14d1de159.
Contents
- All 256 Text projections (MLP, GDN, attention): NVFP4 imported bit-exactly from QUASAR, including its activation scales, so prefill uses FP4 activations.
- Embedding and output head: row-scaled FP8 from QUASAR's BF16 tensors.
- GDN gating: BF16, decoded from QUASAR.
- DFlash2, MTP, Vision, proposal head: the official artifact's formats.
How it was measured
Decode and prefill: one request at a time, thinking on, FP8 KV, prefill chunk 1,024, on fixed conversational, JSON, essay and 16.7k-token-document prompts. Round time is the comparable figure; tokens/s also depends on how many draft tokens each sampled response accepts. Official values are two server sessions, this artifact one; expect about ±2%.
GPQA-Diamond: NInfer's eval/ with EvalScope 1.10.0 and the official card's GPQA settings
(thinking, temperature 1.0, top-p 0.95, top-k 20, seed 42, INT8 KV), using DFlash2 instead of MTP.
Reports and configs are in eval/.
Not tested: MTP, image/video quality, and benchmarks other than GPQA-Diamond.
Rebuild
From an NInfer checkout at d44ab584, with QUASAR at revision
15d2e47bffe5d8ad23928879f8f7d2f74909e259 and DFlash2 at 50307d4c4cde6860d4eee73e2547cd786fe8e8a4:
python3 -m tools.convert \
--model sources/Qwen3.8-27B-QUASAR-NVFP4 \
--recipe qwen3_8_27b_quasar.py \
--source dflash2=sources/Qwen3.8-27B-DFlash2 \
--components text,vision,mtp,dflash2 \
--resource chat_template.jinja=tools/chat_templates/qwen3_8.jinja \
--proposal --name qwen3.8-27b --device cpu \
--out models/qwen3_8_27b_quasar.ninfer
The recipe is qwen3_8_27b_quasar.py; the per-object record is
conversion-report.json.
License
Apache-2.0, as are the Qwen3.8-27B base model, the QUASAR checkpoint and the DFlash2 drafter.
- Downloads last month
- 194
Model tree for jesdga/Qwen3.8-27B-QUASAR-DFlash2-nvfp4-NInfer
Evaluation results
- Accuracy (0-shot on GPQA-DiamondNInfer EvalScope 1.10.089.900