Instructions to use orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M
Use Docker
docker model run hf.co/orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M
- Ollama
How to use orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF with Ollama:
ollama run hf.co/orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF with Docker Model Runner:
docker model run hf.co/orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M
- Lemonade
How to use orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-Uncensored-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Q6_K looks reachable: the PLE tensor is 51.2B, not ~66B
Re: "Why no Q6_K / Q8_0 / F16 / BF16?" — the premise is off, and Q6_K looks reachable.
The tensor is 51.2B, not ~66B. per_layer_token_embd.weight is [160, 320001536]. Check that needs no assumptions about block overhead: the base model's BF16 GGUF has a single shard of 102,400,491,712 bytes = 51,200,245,760 x 2, plus a 192-byte shard header. At 66B it would be 132 GB.
So the real sizes are Q6_K ~ 42.0 GB, Q8_0 ~ 54.4 GB, BF16 ~ 102.4 GB. Q6_K fits under 50 GB, and under your ~48 GB split.
There is no 50 GB per-file limit on the Hub. The docs say <200 GB recommended, 500 GB hard. The 50 GB figure is CloudFront's single-request download limit — real, but a delivery constraint: larger files need hf_xet or ranged requests (huggingface_hub#3868).
And the upstream Unsloth GGUFs do not cap at Q4 — that repo has UD-Q5_K_XL, UD-Q6_K_XL, Q8_0 and BF16. Their Q5/Q6/Q8 all carry the PLE table at Q8_0 in a 54.4 GB shard (identical byte size in all three), and BF16 ships it as a single 102.4 GB file. Worth noting their UD-Q5_K_XL keeps the table at Q8_0, where a stock Q5_K_M puts it at Q5_1 (~38.4 GB).
If the actual blocker is the quantizer bumping the embedding table to Q8_0 at Q6, current master lets you name the tensor directly:
--tensor-type per_layer_token_embd=q6_k
(src/llama-quant.cpp: "per_layer_token_embd follows --token-embedding-type by default, but it is a large separate table, so let an explicit --tensor-type name it")
You're right on the core facts: per_layer_token_embd is [160, 320001536] = 51.2B (BF16 shard 102,400,491,712 B), not ~66B; the 50 GB figure is a CloudFront single-request download limit, not a storage cap (xet / ranged handles larger files); and Unsloth ships up to Q6_K/Q8_0/BF16, not Q4. Our card's tensor size, the "50 GB per-file limit", and the "matches Unsloth's Q4 cap" wording were wrong, and we're correcting them.
One correction on the Q6_K sizing, though. That tensor can't actually be stored as Q6_K: its quantized dimension is 160, and the K-quants (Q4_K/Q5_K/Q6_K) require the row length to be divisible by 256. 160 isn't, so llama-quantize falls back. For Q4_K/Q5_K it falls back to the legacy block-32 types (Q4_1 32 GB / Q5_1 ~38 GB), which is exactly why the lineup could reach Q5 under 50 GB. But Q6_K has no block-32 legacy equivalent, so 54.4 GB)** and has to sit in its own shard. That's precisely what Unsloth's UD-Q6_K_XL does — its shard 00003 is 54.40 GB. So there's no Q6_K with every shard under 50 GB; the accurate framing is "a >50 GB shard is fine as long as it's fetched with xet/ranged," which is your actual point.per_layer_token_embd falls back to **Q8_0 (
We're adding Q6_K on exactly that basis (same structure as upstream: ~6 shards, one ~54 GB PLE shard served via xet), and fixing the card's incorrect claims.