Q6_K looks reachable: the PLE tensor is 51.2B, not ~66B

#2
by yhn112 - opened

Re: "Why no Q6_K / Q8_0 / F16 / BF16?" — the premise is off, and Q6_K looks reachable.

The tensor is 51.2B, not ~66B. per_layer_token_embd.weight is [160, 320001536]. Check that needs no assumptions about block overhead: the base model's BF16 GGUF has a single shard of 102,400,491,712 bytes = 51,200,245,760 x 2, plus a 192-byte shard header. At 66B it would be 132 GB.

So the real sizes are Q6_K ~ 42.0 GB, Q8_0 ~ 54.4 GB, BF16 ~ 102.4 GB. Q6_K fits under 50 GB, and under your ~48 GB split.

There is no 50 GB per-file limit on the Hub. The docs say <200 GB recommended, 500 GB hard. The 50 GB figure is CloudFront's single-request download limit — real, but a delivery constraint: larger files need hf_xet or ranged requests (huggingface_hub#3868).

And the upstream Unsloth GGUFs do not cap at Q4 — that repo has UD-Q5_K_XL, UD-Q6_K_XL, Q8_0 and BF16. Their Q5/Q6/Q8 all carry the PLE table at Q8_0 in a 54.4 GB shard (identical byte size in all three), and BF16 ships it as a single 102.4 GB file. Worth noting their UD-Q5_K_XL keeps the table at Q8_0, where a stock Q5_K_M puts it at Q5_1 (~38.4 GB).

If the actual blocker is the quantizer bumping the embedding table to Q8_0 at Q6, current master lets you name the tensor directly:

--tensor-type per_layer_token_embd=q6_k

(src/llama-quant.cpp: "per_layer_token_embd follows --token-embedding-type by default, but it is a large separate table, so let an explicit --tensor-type name it")

OrcaRouter org

You're right on the core facts: per_layer_token_embd is [160, 320001536] = 51.2B (BF16 shard 102,400,491,712 B), not ~66B; the 50 GB figure is a CloudFront single-request download limit, not a storage cap (xet / ranged handles larger files); and Unsloth ships up to Q6_K/Q8_0/BF16, not Q4. Our card's tensor size, the "50 GB per-file limit", and the "matches Unsloth's Q4 cap" wording were wrong, and we're correcting them.

One correction on the Q6_K sizing, though. That tensor can't actually be stored as Q6_K: its quantized dimension is 160, and the K-quants (Q4_K/Q5_K/Q6_K) require the row length to be divisible by 256. 160 isn't, so llama-quantize falls back. For Q4_K/Q5_K it falls back to the legacy block-32 types (Q4_1 32 GB / Q5_1 ~38 GB), which is exactly why the lineup could reach Q5 under 50 GB. But Q6_K has no block-32 legacy equivalent, so per_layer_token_embd falls back to **Q8_0 (54.4 GB)** and has to sit in its own shard. That's precisely what Unsloth's UD-Q6_K_XL does — its shard 00003 is 54.40 GB. So there's no Q6_K with every shard under 50 GB; the accurate framing is "a >50 GB shard is fine as long as it's fetched with xet/ranged," which is your actual point.

We're adding Q6_K on exactly that basis (same structure as upstream: ~6 shards, one ~54 GB PLE shard served via xet), and fixing the card's incorrect claims.

Sign up or log in to comment