Instructions to use local-inference-lab/GLM-5.3-Flash-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use local-inference-lab/GLM-5.3-Flash-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="local-inference-lab/GLM-5.3-Flash-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("local-inference-lab/GLM-5.3-Flash-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("local-inference-lab/GLM-5.3-Flash-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use local-inference-lab/GLM-5.3-Flash-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "local-inference-lab/GLM-5.3-Flash-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "local-inference-lab/GLM-5.3-Flash-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/local-inference-lab/GLM-5.3-Flash-NVFP4
- SGLang
How to use local-inference-lab/GLM-5.3-Flash-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "local-inference-lab/GLM-5.3-Flash-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "local-inference-lab/GLM-5.3-Flash-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "local-inference-lab/GLM-5.3-Flash-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "local-inference-lab/GLM-5.3-Flash-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use local-inference-lab/GLM-5.3-Flash-NVFP4 with Docker Model Runner:
docker model run hf.co/local-inference-lab/GLM-5.3-Flash-NVFP4
Add community tool-calling evaluation: tool-eval-bench v2.7.0 hard mode 94/100 (seeds 42+43)
Same-instrument variance note (2026-10-05, same day). We re-ran this exact recipe through a per-request injection shim (a local proxy adding chat_template_kwargs {"reasoning_effort":"max","clear_thinking":true} to each request; server config untouched at canonical effort-high) to validate the alternative delivery path. Injection was verified effective (temp-0 effort-low probe through the shim matches direct-low exactly: reasoning 57 = 57 chars, vs direct-high 574).
Scores through the shim: seed 42 = 93, seed 43 = 91, and an identical-conditions seed-43 repeat = 96 — against the server-side-default receipt of 94/94 above.
Takeaway: with temp 0 + fixed seed, single-run tool-eval-bench scores on a shared serving stack are not rankable within ~±3 points — server-side batch composition (request interleaving, incl. other tenants' traffic) flips borderline scenarios, and 30-turn agentic trajectories amplify single-token flips. The stable signal across all 5 runs is the deterministic fail set: TC-43 (5/5 runs), TC-68 (5/5), TC-51 (4/5); everything else wobbles run-to-run. The 94/94 receipt stands as the leaderboard-recipe datapoint; the ±1 comparison against the EXL3 reference (95/95) should be read as "within the same band", not as a ranking.
Concurrency scaling + repeat-averaged protocol (2026-10-05, update 2).
Follow-up to the variance note above: we tested whether concurrency itself affects the noise. Same seed/flags, only --parallel varied, all through the validated per-request injection shim (effort-max behavior, server config untouched):
- parallel 4: 91 / 94 / 96 across 3 identical-condition runs (5-pt spread) — the noisy regime
- parallel 24: 94 / 94 (n=2)
- parallel 88 (all 88 scenarios at once; vLLM queues at
max_num_seqs32): see below
New protocol we're adopting: run at high concurrency (let vLLM handle batching — no artificial caps), 3 repeats per seed, report per-seed averages. At parallel 88 one 88-scenario suite takes 66–125 s, so 3 repeats × 2 seeds ≈ 8 minutes total.
Results, parallel 88, 3 repeats per seed:
| Seed | Runs | Mean |
|---|---|---|
| 42 | 94 / 94 / 97 | 95.0 |
| 43 | 95 / 94 / 93 | 94.0 |
| all 6 | 93–97 | 94.5 |
Observations:
- High concurrency collapses most of the run-to-run churn (earlier 4/4 runs within 1 pt at C24/C88; 6-run spread now 93–97 vs 91–96 at parallel 4) and is ~2x faster wall-clock — but does not eliminate variance (one 97 outlier in 6 runs).
- The deterministic fail set is unchanged at every concurrency: TC-68 (6/6), TC-43 (6/6), TC-51 (5/6), TC-80 (4/6); every other fail wobbled ≤2/6.
- Per-conversation median turn rises with concurrency (~1.5 s → ~10–12 s); scores unaffected.
- No timeout risk: max scenario duration 85 s against the 600 s budget at any concurrency tested.
Practical recommendation for anyone running this instrument on a shared stack: don't rank single-run scores (±3–5 run-to-run at low parallelism); run at parallel ≥ 24, 3 repeats, report per-seed averages. The 94/94 server-side-default receipt above stands as the original leaderboard-recipe datapoint; the repeat-averaged high-concurrency read of the same checkpoint is 95.0 (seed 42) / 94.0 (seed 43), overall 94.5.
Official-endpoint reference (update 3): original weights via z.ai API — 90.5 avg, with a serving-stack signature in Structured Output.
We ran the identical protocol (parallel 88, 3 repeats per seed, temp 0, reasoning effort max — pinned via z.ai's documented top-level reasoning_effort param, which defaults to max for GLM-5.3-Flash) against the official z.ai endpoint (glm-5.3-flash @ api.z.ai — the original checkpoint as served by z.ai):
| Seed | Runs | Mean |
|---|---|---|
| 42 | 93 / 91 / 88 | 90.7 |
| 43 | 91 / 90 / 90 | 90.3 |
| all 6 | 88–93 | 90.5 |
Cost: $0.64 for all six runs (3.02M input / 0.37M output tokens).
The ~4-point gap vs the community quants (94.5 local NVFP4, 95 EXL3 reference) is concentrated in one category: Structured Output scores 6/12 (50%) on the official endpoint vs 10/12 (83%) on the local stack — TC-65–69 fail 6/6 on official ("called the tool correctly but final output is not valid JSON") while passing locally. TC-45 is an endpoint-capability artifact ("Endpoint does not enforce tool_choice='required'"). The universal fails TC-43/TC-68 fail on both stacks.
Reading:
- The official hosted serving is not a clean BF16-checkpoint reference for this instrument — its template/system-prompt/structured-output handling differs enough to flip an entire category. Outside that category, official serving is roughly at parity with the community quants.
- Therefore the 94–95 quant scores stand: on this instrument there is no measurable quant-degradation signal outside the serving-stack-confounded category, and the deterministic weak spots (TC-43, TC-68) are shared by the original weights.
- If anyone wants a true BF16 reference, it needs the original weights on a local vLLM stack (not the hosted API) — nobody in the community has the hardware for BF16 Flash (~790 GB); the NVIDIA BF16 rows in their card table remain the best published anchor.
Practical notes for replicating: z.ai PAYG works fine for this (top up at the console → glm-5.3-flash @ api.z.ai/api/paas/v4; their reasoning_effort defaults to max; interleaved thinking with tools is default-on) at ~$0.10/run. Gotcha: their API has no /v1 path segment — clients that append /v1/chat/completions will 404.
Original-weights control on the same stack: zai FP8 blockwise originals — matched-88 mean 94.8 (n=6), deterministic trio reproduced
Complement to the receipt above (NVFP4 checkpoint 94/94, seeds 42+43): we ran the original zai-org/GLM-5.3-Flash FP8 weights (DeepSeek-style blockwise, 306 GiB) through the identical leaderboard recipe on the same serving stack. Purpose: isolate the checkpoint's quant effect with the engine held constant.
Serving stack (matched to the receipt): Karmic Kraken beta digest-pinned sha256:2230db60afb4fd06dc2ef2b7f7f27dc7f78d7a50be70a08461b4df9fa5e90732 (same digest as the receipt's run), TP4 on 4× RTX PRO 6000 (on-demand rental), MTP3, FP8 KV, gpu-memory-utilization 0.86, max-model-len 65536. Quant-adapter differences forced by the FP8-original checkpoint: MoE backend triton (mode-config patched off marlin), linear backend cutlass, load-format: safetensors, ViT attention TORCH_SDPA, FLASHINFER_CUDA_ARCH_LIST=12.0f, VLLM_USE_DEEP_GEMM=0, LMCache L2 off. Server-side --default-chat-template-kwargs {"reasoning_effort":"max","clear_thinking":true} — same mechanism as the receipt. Wrapper llm_decode_bench.py v0.7.7: "differs_from_leaderboard": [] on all runs.
Results (hardmode, seeds 42/43 × 3 reps, temp 0, parallel 4, max-turns 30, timeout 600):
| Run | final | fails |
|---|---|---|
| seed 42 r1 | 92 | TC-21, 43, 51, 68, 89, 92 |
| seed 42 r2 | 93 | TC-43, 51, 68, 89, 92 |
| seed 42 r3 | 91 | TC-21, 43, 49, 51, 68, 89, 92 |
| seed 43 r1 | 92 | TC-21, 43, 51, 68, 92 |
| seed 43 r2 | 92 | TC-21, 43, 51, 68, 89, 92 |
| seed 43 r3 | 94 | TC-21, 43, 51, 68, 89, 92 |
Version note: these runs executed on the wrapper's current pinned tool-eval-bench 2.7.1.dev14+g570951a77 (92 scenarios = the receipt's 88 + TC-89/90/91/92), while the committed receipt above used v2.7.0 (88). Normalizing to the matched 88-scenario subset (per-scenario points over the 88 shared IDs, max 176): 94.9 / 96.0 / 92.6 / 94.3 / 94.9 / 96.0 → mean 94.8 (n=6), vs the receipt's 94.3 (166/176, both seeds).
Reading:
- The deterministic fail trio TC-43 / TC-51 / TC-68 reproduces 6/6 on the original weights — identical to the NVFP4 receipt's signature. It is a property of the model+stack, not of the quantization.
- Matched-set comparison (94.8 vs 94.3) is within the single-run spread documented for this benchmark (±3 on same-condition repeats); headline: the NVFP4 checkpoint costs ~nothing on agentic tool-calling vs the original FP8 weights.
- FP8 originals additionally fail TC-21 (4/6) which the NVFP4 receipt passed; the receipt's TC-49 fail appears once (1/6). The new-set scenarios TC-92 (6/6) and TC-89 (5/6) fail on the originals — not comparable against the v2.7.0 receipt.
Disclosure: ~2.4 h on 4× RTX PRO 6000 rented on-demand (vast.ai) for the 6 runs + serving; no API costs (local serving). Raw per-run JSONs available on request (kept privately to avoid publishing scenario content).