GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD
GLM-5.3-Flash with every precision choice stored in the checkpoint, so serving needs no load-time conversion:
Main model:
Routed expert projections - NVFP4, quantization-aware distillation (QAD)
Attention projections - MXFP8 (402: KDA + MLA/indexer, Spark selection)
Shared experts - MXFP8 (129, Spark selection)
MTP (layer 45):
Routed expert projections - NVFP4 (W4A16 group)
Attention / shared - MXFP8 (Spark selection)
Routed-expert E4M3 block scales - lossless CSF compression
Provenance
- Routed experts:
local-inference-lab/GLM-5.3-Flash-NVFP4revisionmtp-bf16(cfd47bd7680e68408924df09b179d5bed25b2ae9), the QAD checkpoint, unchanged. - MXFP8: the 531 projections of
local-inference-lab/GLM-5.3-Flash-NVFP4-Spark's MXFP8 groups, quantized from the QAD checkpoint's BF16 weights with vLLM'smxfp8_e4m3_quantize(E4M3 values, UE8M0 scale per 32) - the same values vLLM's online MXFP8 produces. - MTP routed experts: quantized from the QAD BF16 MTP experts with vLLM's online
NVFP4 MoE quantizer (
_quantize_moe_weight_to_nvfp4). - CSF: codec
byte-window4-fixed-stream-u24-exceptions/1, 36,288 scale matrices, 19.03 GB of scales stored in 9.83 GB.verification.jsonrecords exact reconstruction of every source shard (SHA-256 of reconstructed and compressed files).
Layout
The lil-nvfp4-csf-checkpoint/1 container:
tensors/- 44 safetensors shards; routed-expert scales are CSF-compressedmetadata/- config (ModelOptMIXED_PRECISION, 574 quantized layers), tokenizer, chat template, generation configconfig.json- a copy ofmetadata/config.jsonat the top of the repository, where the Hub counts downloads; runtimes readmetadata/manifest.json,build-contract.json,receipts/,verification.jsonLICENSE,NOTICE,LICENSES/(upstream license texts),CITATION.cff,REUSE.toml, and the SHA-256 of every file inSHA256SUMSandlil-manifest.json
Serving
Use a runtime with the NVFP4-CSF decoder (Karmic Kraken beta images: vLLM
--quantization nvfp4_csf --load-format nvfp4_csf). The served model directory
holds the files of metadata/, with config.json's quantization_config
replaced by:
{
"quant_method": "nvfp4_csf",
"format_version": 1,
"checkpoint_root": "/path/to/this/snapshot",
"source_quantization_config": { "...": "metadata/config.json quantization_config" }
}
For exact arithmetic on GLM-5.3-Flash's checksum-style prompts we run the routed
experts as W4A16 with FP32 router-weight combine
(VLLM_B12X_MOE_FP4_FORCE_A16=1, B12X_W4A16_FP32_TOPK_WEIGHTS=1).
Measured on 2x RTX PRO 6000 Blackwell Max-Q (TP2, DCP2, MTP with 3 speculative tokens): model memory 83.2 GiB per GPU; decode about 170 tok/s at concurrency 1 and 490-540 tok/s at concurrency 8 (CSF decoder version dependent); prefill about 7.6K tok/s at 32K context. Needle-checksum (500 identical 8K prompts at temperature 0, concurrency 8, no prefix-cache reuse): 0.5% near misses, against 0.25% for the same QAD experts with BF16 attention and shared experts (the difference is within run-to-run noise).
Upstream license
GLM-5.3-Flash (zai-org/GLM-5.3-Flash-BF16) is licensed by Z.AI under the MIT License. Its license text is in LICENSES/MIT.txt, and NOTICE names the upstream materials. This checkpoint is licensed under the Local Inference Lab License, Version 1.0 (LICENSE); see License and attribution below.
License and attribution
GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD is licensed under the Local Inference Lab License, Version 1.0, which reproduces the terms and conditions of the Apache License, Version 2.0, and adds conditions that restrict them. It is not the Apache License, and these files are not open source. Copyright (c) 2026 Local Inference Lab, Inc., a non-profit organization, 311 West Main Street, Grayson, Kentucky 41143. Revisions published up to and including revision 9ca06e24effde995982c439091722e9e76db82bc were released under the MIT License, and copies of those revisions remain under it. The Local Inference Lab License, Version 1.0 applies to later revisions.
- No reuploads. Do not upload, mirror, or redistribute these files, or any Substantially Similar Copy of them (including renamed, re-sharded, re-packaged, metadata-stripped, converted, or dequantized copies). Link to https://huggingface.co/local-inference-lab/GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD instead.
- Attribution at the top. Every README, model card, or other landing page for a project, model, dataset, application, service, or distribution that contains, is derived from, or runs GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD must begin with this Attribution Notice as its first paragraph (a single title line may come before it):
This model is based on GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD by Local Inference Lab, Inc., a non-profit organization, available at https://huggingface.co/local-inference-lab/GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD. GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD is licensed under the Local Inference Lab License, Version 1.0.
Plain-text form:
This model is based on GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD by Local Inference Lab, Inc., a non-profit organization, available at https://huggingface.co/local-inference-lab/GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD. GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD is licensed under the Local Inference Lab License, Version 1.0.
- Keep the marks. Do not remove the license metadata, the Identifying Marks ("Ні пуху, ні пера", LIL-CANARY-F795-6543-0661-0749), or the copyright notices embedded in these files.
- Breach. A breach of the no-reupload or attribution terms ends the license immediately (Section 5.1), and Local Inference Lab, Inc. may ask hosting services to remove the material (Section 5.2).
File integrity: SHA256SUMS and lil-manifest.json list the SHA-256 of every file.
- Downloads last month
- 35
Model tree for local-inference-lab/GLM-5.3-Flash-NVFP4-MXFP8-CSF-QAD
Base model
zai-org/GLM-5.3-Flash-BF16