Instructions to use unsloth/Qwen3.8-27B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use unsloth/Qwen3.8-27B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M # Run inference directly in the terminal: llama cli -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M # Run inference directly in the terminal: llama cli -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M
Use Docker
docker model run hf.co/unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M
- LM Studio
- Jan
- Ollama
How to use unsloth/Qwen3.8-27B-GGUF with Ollama:
ollama run hf.co/unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M
- Unsloth Desktop
- Pi
How to use unsloth/Qwen3.8-27B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use unsloth/Qwen3.8-27B-GGUF with Docker Model Runner:
docker model run hf.co/unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M
- Lemonade
How to use unsloth/Qwen3.8-27B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M
Run and chat with the model
lemonade run user.Qwen3.8-27B-GGUF-UD-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use unsloth/Qwen3.8-27B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use unsloth/Qwen3.8-27B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
We did it! Qwen3.8-27B is now the #1 most liked GGUF of all time! π
pinnedππ₯ 44
#127 opened 27 days ago
by
danielhanchen
Introducing Unsloth Dynamic v3 Qwen3.8
pinnedππ₯ 55
37
#74 opened about 2 months ago
by
danielhanchen
Uniform GGUF quants silently break Qwen3.8-27B's deep thinking β reproduced on llama.cpp AND vLLM (short tasks unaffected)
#135 opened about 1 hour ago
by
Jerryccccc
stops generating randomly
#134 opened 3 days ago
by
AndreyPanfilov
q8_K_L
#132 opened 16 days ago
by
lobstertot
Which GGUF quantization would you recommend for an RTX 5080 (16GB VRAM)?
2
#131 opened 18 days ago
by
pengwenzhi
Qwen3.8 27B VRAM requirements per quant
#130 opened 18 days ago
by
yash-711
improve unsloth
3
#129 opened 24 days ago
by
l0rdraiden
spec-draft-n-max = 4 error but works with 2 no error
#126 opened 28 days ago
by
Utku92
Vision + DFLASH2 fix?
π 3
#125 opened 29 days ago
by
Josephur
MTP with --spec-draft-hf in llama.cpp
3
#124 opened 30 days ago
by
ZhaolunYin
GGUF versions available via Torrent
1
#123 opened 30 days ago
by
baiomys
Qwen 3.8 IQ2 XXS - new version Lower perfomances
4
#122 opened about 1 month ago
by
GrahamShandor
Community benchmark: Qwen3.8-27B GGUF on an RTX 5060 Ti 16GB
1
#121 opened about 1 month ago
by
alexzhang0118
incorrect quantization
#120 opened about 1 month ago
by
Utku92
Intel Arc B70 results: 32 T/s with UD-Q6_K, SYCL, MTP
π 2
3
#114 opened about 1 month ago
by
Testererer
The Q8_K_XL seems to have the wrong file type
π 1
#113 opened about 1 month ago
by
superslowsoftware
Chat template
4
#112 opened about 1 month ago
by
Akzk
Tested with BrainBench-llama, the qwen3.8-27b-UD-*: they are comparable to the best Claude Opus 4.6.
π₯ 2
1
#111 opened about 1 month ago
by
WhiteDan64
Native C laptop-CPU runtime for the Qwen3.8-27B Dynamic V3 GGUFs
#109 opened about 1 month ago
by
shyringo
Will any other model get UD3 update?
βπ 3
1
#108 opened about 1 month ago
by
Gavin-chen
removed Qwen3.8-27B-IQ4_XS.gguf vs replaced Qwen3.8-27B-UD-IQ4_XS.gguf.
π₯ 4
6
#104 opened about 1 month ago
by
userJ345
i goon to unsloth
π 1
11
#100 opened about 2 months ago
by
Hellomaniamcoollol
What's the use of mtp-Qwen3.8-27B-Q4_0.gguf?
7
#97 opened about 2 months ago
by
cnayan01
Could not download MTP drafter: RemoteEntryNotFoundError: 404 Client Error
π₯β 2
3
#95 opened about 2 months ago
by
CyberTod
Qwen in the UD-IQ3_XXS quant is honestly pretty impressive on my RTX 5070 Ti with 32 GB of 6400 MT/s RAM.
π 1
4
#94 opened about 2 months ago
by
Nevermorye
KV-cache KLD scales with model fidelity, not quant family β three null results and one metric trap
1
#89 opened about 2 months ago
by
Knappy
Imatrix now available!
π₯π 15
1
#88 opened about 2 months ago
by
danielhanchen
# VRAM is linear in context: a formula that predicts any Qwen3.8-27B GGUF config (and settles the 150K debate)
π₯ 1
1
#87 opened about 2 months ago
by
Knappy
Why delete Qwen3.8-27B-IQ4_NL.gguf with huggingface_hub?
π 6
5
#86 opened about 2 months ago
by
puchuu
Even at low Q6 quant it's very good!
ππ€ 3
3
#84 opened about 2 months ago
by
auf1r2
The model loads successfully via llama.cpp, but it crashes with an abnormal exit as soon as inference begins. Why is this happening?
π€ 1
2
#83 opened about 2 months ago
by
Matheartcis
Can i ask the sweet point for 12gb cards?
4
#81 opened about 2 months ago
by
AsThirtyThree
DFlash2?
π 5
1
#79 opened about 2 months ago
by
guarism0
What is mtp-Qwen3.8-27B-Q4_0.gguf used for?
2
#78 opened about 2 months ago
by
artden111
Qwen3.8-27B on Dual RTX 3060 β llama.cpp Benchmark
5
#77 opened about 2 months ago
by
marjimgu
`Qwen3.8-27B-UD-IQ1_S.gguf`Has anyone tried this?
8
#76 opened about 2 months ago
by
artden111
NVFP4/MXFP4 GGUF Version?
π 3
2
#75 opened about 2 months ago
by
CYISNOTHERE
Qwen3.8-27B: Q5_K_M, RTX 5060ti 16gb x 2, total 32gb, 30t/s, 215k context
π€ 1
15
#72 opened about 2 months ago
by
Artem7799
Squeezing Qwen 3.8 27B into a Single 16 GB GPU β Almost 42 tok/s, 64K Context
π 2
7
#70 opened about 2 months ago
by
Ataa
Crash when used with llama.cpp on image loading.
3
#69 opened about 2 months ago
by
MalcolmMielle
Running Qwen3.8-27B on a Single RTX 3090 (24 GB) 45-70t/s with 150k context
π₯ 3
5
#68 opened about 2 months ago
by
SergeySS8
Qwen3.8-27B only 18 tok/s vs Qwen3.6-27B 62 tok/s on A100 80GB + vLLM β expected?
π 4
5
#66 opened about 2 months ago
by
Alecone
Qwen3.8-27B on an RTX 4070 Ti 12GB: IQ2 vs Q2, 16K context, and MTP
π 2
1
#65 opened about 2 months ago
by
Hugosmr
Enable Reasoning Level with LM Studio and VS Code GitHub Copilot Chat
β€οΈ 6
1
#63 opened about 2 months ago
by
asage-me
Running Qwen3.8-27B Q4_K_S on an RTX 3060 12GB with 96K Context + MTP (~10 t/s)
π 3
13
#61 opened about 2 months ago
by
Hjx2
UD_IQ1_XXXS possible like you did with the big one?
π₯ 1
4
#60 opened about 2 months ago
by
TheWegemann
Qwen3.8-27B UD-Q2_K_XL: unexpectedly slow ROCm prefill; ssm_alpha/ssm_beta are IQ1_M
π₯ 1
3
#59 opened about 2 months ago
by
djtrondheim