Qwen2.5-14B-Instruct โ€” Pollard

Pollard shrank this model: 29.60 GB (f16) โ†’ 3.92 GB โ€” 87% smaller, 7.6ร— down.

The smallest rung here; larger, higher-fidelity rungs are listed below.

format this model's size
f16 29.60 GB
Q8_0 ~15.69 GB
Q6_K ~12.14 GB
Q4_K_M ~8.58 GB
PollardMix (this repo's IQ1_KT) 3.92 GB

Pollard builds of Qwen/Qwen2.5-14B-Instruct made with Pollard Weights โ€” a ladder of measured-allocation quants (bits placed by per-layer sensitivity, not a uniform crush).

Standard GGUF โ€” runs in stock llama.cpp / ik_llama.cpp, Ollama, LM Studio, except where noted. IQ1_KT needs ik_llama.cpp: its allocation puts ik_llama-only atoms on the tensors it protects. The rest run anywhere.

Model details

Parameter count ~14.8B
Architecture qwen2
Input support text
imatrix yes โ€” see calibration
Perplexity measured yes โ€” table below

Which file should I choose?

Every rung is the same weights, sized to a different RAM budget by the measured allocation. Pick the largest one that fits your machine with room for context:

  • ~14 GB RAM / VRAM โ†’ Q6_K (12.12 GB). stock atoms โ€” runs anywhere. See Speed below: 12.12 GB does not stay resident on a 16 GB card
  • ~11 GB RAM / VRAM โ†’ IQ4_XS (8.71 GB). stock atoms โ€” best size/quality point on this ladder
  • ~9 GB RAM / VRAM โ†’ IQ3_S (7.38 GB). stock atoms
  • ~6 GB RAM / VRAM โ†’ IQ1_KT (3.92 GB). (ik_llama.cpp) trellis atoms โ€” ik_llama.cpp only

Available files (WikiText-2 raw test, ctx 2048, 145 chunks)

f16 reference PPL 4.9718 ยฑ0.0288.

file PPL size tok/s Mean KLD runs in notes
Qwen2.5-14B-Instruct-Pollard-IQ1_KT.gguf 8.2662 3.92 GB 53.1 โ€” ik_llama trellis atoms โ€” ik_llama.cpp only
Qwen2.5-14B-Instruct-Pollard-IQ3_S.gguf 5.5004 7.38 GB 92.6 โ€” any llama.cpp stock atoms
Qwen2.5-14B-Instruct-Pollard-IQ4_XS.gguf 5.1595 8.71 GB 83.3 โ€” any llama.cpp stock atoms โ€” best size/quality point on this ladder
Qwen2.5-14B-Instruct-Pollard-Q6_K.gguf 4.9914 12.12 GB โ€” โ€” any llama.cpp stock atoms โ€” runs anywhere. See Speed below: 12.12 GB does not stay resident on a 16 GB card

tok/s measured on an RTX 5070 Ti (16 GB), full GPU offload, Windows/CUDA.

The numbers (WikiText-2 raw, ctx 2048, 145 chunks)

Perplexity with error bars, and the trellis rung against the uniform quants it competes with. Sizes are the published byte counts, in GB:

build PPL size bpw Mean KLD Median KLD top-1
f16 reference 4.9718 ยฑ0.0288 29.60 GB 16.00 โ€” โ€” โ€”
Q6_K 4.9914 ยฑ0.0290 12.12 GB 6.57 โ€” โ€” โ€”
IQ4_XS 5.1595 ยฑ0.0302 8.71 GB 4.72 โ€” โ€” โ€”
IQ3_S 5.5004 ยฑ0.0327 7.38 GB 4.00 โ€” โ€” โ€”
uniform IQ2_KT (2-bit ceiling) 6.92 4.62 GB 2.50 0.353 0.128 76.85%
IQ1_KT (PollardMix) 8.2662 ยฑ0.0534 3.92 GB 2.12 0.552 0.253 70.79%
uniform IQ1_KT (1-bit baseline) 9.65 3.61 GB 1.96 0.714 0.363 66.41%

Q6_K costs 0.4% perplexity against the f16 reference at 41% of its size, and IQ4_XS costs 3.8% at 29%. IQ3_S is the first rung where the cost is visible rather than academic, at +10.6%.

PollardMix beats the uniform 1-bit trellis quant on every metric โ€” PPL โˆ’14%, Mean KLD โˆ’23%, Median KLD โˆ’30%, top-1 +4.4 pts โ€” at +8.5% size, and stays under the 2-bit ceiling. It also passes a chat-coherence gate (explanation, code, reasoning, creative continuation all coherent).

PollardMix vs uniform 1-bit โ€” 7B and 14B

Speed

Decode speed on the same machine, full GPU offload:

file size tok/s
IQ4_XS 8.71 GB 83.3
IQ3_S 7.38 GB 92.6
IQ1_KT (PollardMix) 3.92 GB 53.1
Q6_K 12.12 GB not yet measured cleanly

Two rows there are worth reading rather than skimming.

IQ1_KT is the smallest file and the slowest measured rung โ€” 53.1 tok/s against IQ3_S's 92.6. Trellis atoms trade decode compute for bytes, so on a GPU that is not memory-starved the small rung loses. Reach for IQ1_KT when the model would not otherwise fit; reach for IQ3_S or IQ4_XS when it does.

Q6_K is listed as unmeasured on purpose. At 12.12 GB it does not stay resident on a 16 GB card once the desktop, the KV cache and the compute buffers are accounted for, and every attempt so far was taken while the GPU was shared โ€” yielding figures between 1.8 and 8.4 tok/s that measure contention, not the model. A number that low reads as a property of the quant, which it is not, so the cell stays empty until the measurement is clean.

Allocation (the surgery)

The trellis rung, tensor by tensor:

tensor role atom
expert / FFN body (gate, up) IQ1_KT crushed
attention k, v IQ1_KT crushed
attention q, output IQ2_KT protected
ffn_down (residual writer) IQ2_KT protected
first-2 / last-2 blocks IQ2_KT protected
token embeddings Q4_K kept
output head Q6_K kept
norms F32 kept

Prompt format

ChatML, the Qwen2.5-Instruct template:

<|im_start|>system You are a helpful assistant.<|im_end|> <|im_start|>user {prompt}<|im_end|> <|im_start|>assistant


The template is embedded in the GGUF metadata, so llama.cpp, Ollama and LM Studio apply it for you.

Download a specific file

pip install -U "huggingface_hub[cli]"
hf download PollardWeights/Qwen2.5-14B-Instruct-Pollard \
  --include "Qwen2.5-14B-Instruct-Pollard-IQ1_KT.gguf" --local-dir ./

How to run

IQ1_KT is built on ik_llama-only atoms, so it runs with ik_llama.cpp:

llama-cli    -m Qwen2.5-14B-Instruct-Pollard-IQ1_KT.gguf -ngl 99 -p "Explain why the sky is blue."
llama-server -m Qwen2.5-14B-Instruct-Pollard-IQ1_KT.gguf -ngl 99

For stock llama.cpp, Ollama or LM Studio, use Q6_K instead:

llama-server -hf PollardWeights/Qwen2.5-14B-Instruct-Pollard:Q6_K
llama-cli    -m Qwen2.5-14B-Instruct-Pollard-Q6_K.gguf -ngl 99 -p "Explain why the sky is blue."

imatrix (calibration)

The importance matrix (qwen14b_calib3.imatrix, included) was computed on Calib-3.0: 13.9 MB of mixed prose, code, reasoning and dialogue, 60 chunks at ctx 512, computed CPU-only. Checked disjoint from the WikiText-2 test split used for the perplexity table โ€” none of the 1,723 test lines over 200 characters appears anywhere in the calibration corpus, so the numbers above are not measured on text the allocation was tuned on.

ARM / AVX

llama.cpp repacks weights into an interleaved layout at load time for faster inference on ARM and AVX machines โ€” no special file needed, online repacking covers these quants. The old Q4_0_4_4/4_8/8_8 variants are not required.

Errata

  • IQ1_KT carries ik_llama-only atoms and needs ik_llama.cpp to run; stock llama.cpp rejects any ggml type above 42 outright. Checked with pollard-ggufcheck, from the files' tensor types rather than their names.
  • Measured allocation places bits by per-layer sensitivity under a size budget.
  • Single machine; replication invited.

Credits & license

Built with Pollard Weights โ€” frontier models, small hardware, no compromise.

Downloads last month
2,746
GGUF
Model size
15B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for PollardWeights/Qwen2.5-14B-Instruct-Pollard

Base model

Qwen/Qwen2.5-14B
Quantized
(197)
this model