Text Generation
Transformers
Safetensors
English
Chinese
glm5_next
image-text-to-text
quantized
nvfp4
mxfp8
modelopt
distillation
quantization-aware-distillation
conversational
multimodal
8-bit precision
Instructions to use local-inference-lab/GLM-5.3-Flash-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use local-inference-lab/GLM-5.3-Flash-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="local-inference-lab/GLM-5.3-Flash-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("local-inference-lab/GLM-5.3-Flash-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("local-inference-lab/GLM-5.3-Flash-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use local-inference-lab/GLM-5.3-Flash-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "local-inference-lab/GLM-5.3-Flash-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "local-inference-lab/GLM-5.3-Flash-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/local-inference-lab/GLM-5.3-Flash-NVFP4
- SGLang
How to use local-inference-lab/GLM-5.3-Flash-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "local-inference-lab/GLM-5.3-Flash-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "local-inference-lab/GLM-5.3-Flash-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "local-inference-lab/GLM-5.3-Flash-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "local-inference-lab/GLM-5.3-Flash-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use local-inference-lab/GLM-5.3-Flash-NVFP4 with Docker Model Runner:
docker model run hf.co/local-inference-lab/GLM-5.3-Flash-NVFP4
Add community multimodal evaluation: MMMU Pro vision-only 76.71% (temp 1.0/top_p 0.95 NV-protocol, 1730 items, 4x RTX PRO 6000 Karmic Kraken stack)
#7
by malaiwah - opened
README.md
CHANGED
|
@@ -91,3 +91,34 @@ NVFP4/MXFP8 ModelOpt layout, including per-expert activation scales.
|
|
| 91 |
The tokenizer, upstream chat template, generation settings, vision assets
|
| 92 |
and MTP tensors are included. KV cache and runtime buffers require memory
|
| 93 |
in addition to the model weights.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 91 |
The tokenizer, upstream chat template, generation settings, vision assets
|
| 92 |
and MTP tensors are included. KV cache and runtime buffers require memory
|
| 93 |
in addition to the model weights.
|
| 94 |
+
|
| 95 |
+
## Community multimodal evaluation (2026-10-04) — MMMU Pro (vision-only), NV-protocol sampling
|
| 96 |
+
|
| 97 |
+
Independent multimodal accuracy measurement of checkpoint `175ae8ce` contributed by [@malaiwah](https://huggingface.co/malaiwah). **Text benchmarks (GSM8K / GPQA Diamond / MMLU-Pro) are covered separately in [PR #6](https://huggingface.co/local-inference-lab/GLM-5.3-Flash-NVFP4/discussions/6)** — this section adds the first multimodal datapoint for the exact checkpoint this card describes. Raw per-item JSONL available on request.
|
| 98 |
+
|
| 99 |
+
### Methodology
|
| 100 |
+
|
| 101 |
+
- Dataset: [MMMU/MMMU_Pro](https://huggingface.co/datasets/MMMU/MMMU_Pro), config `vision` (question + all options rendered into the image; 10-choice), official `test` split, N = 1,730 (full coverage, 0 request errors after a resume pass).
|
| 102 |
+
- Prompt: official `mmmu-pro/prompts.yaml` **CoT-vision** template ("Write out the multiple-choice question in the image and then solve it… 'Answer: $LETTER'").
|
| 103 |
+
- Parsing: mirrors the official `parse_multi_choice_response` (last `Answer:` → unique option-letter match, layered fallbacks); 3 parse failures, each being an item that hit the `max_tokens` cap in a reasoning loop and never emitted the marker — counted as incorrect, matching how the sibling-table protocol treats truncations.
|
| 104 |
+
- Sampling: **temperature 1.0 / top_p 0.95** (sibling-table protocol, identical to the card's NVIDIA eval settings), `max_tokens` 327,680, no other overrides — server defaults enforced. 16-way concurrent requests against the OpenAI-compatible endpoint; the vision tower passes through unquantized (quantization applies to routed-expert/dense-MLP linear ops only), so this measures the end-to-end multimodal pipeline + LLM-side quantization.
|
| 105 |
+
- Token profile: mean 3,470 completion tokens (thinking on, effort `high`), p95 11,403, 6.0M completion tokens total.
|
| 106 |
+
- Runner: custom minimal OpenAI-compatible client (per-item JSONL records: pred, latency, completion tokens), smoke-validated before launch.
|
| 107 |
+
|
| 108 |
+
### Serving stack
|
| 109 |
+
|
| 110 |
+
Identical to the text-evaluation section in PR #6: 4× RTX PRO 6000 Blackwell, vLLM `0.1.dev22021+g93dabce32`, digest-pinned `karmic-kraken-beta` image, MTP ×3 + LMCache engine-driven, ModelOpt 0.45.0, 1M context.
|
| 111 |
+
|
| 112 |
+
### Results
|
| 113 |
+
|
| 114 |
+
| Benchmark | Items | Sampling | Accuracy | Wilson 95% CI | Notes |
|
| 115 |
+
| --- | --- | --- | --- | --- | --- |
|
| 116 |
+
| MMMU Pro (vision) | 1730 | temp 1.0 / top_p 0.95, max_tokens 327,680 | **76.71%** (1327/1730) | 74.66–78.64 | 0 request errors; 3 max-token cap hits (reasoning loops, counted incorrect) |
|
| 117 |
+
|
| 118 |
+
Reference anchors (NVIDIA's card table for this model family, GB200): BF16 baseline **0.7688**, NVIDIA ModelOpt-NVFP4 **0.763** (different quantization recipe — ModelOpt PTQ, not this QAD distillation).
|
| 119 |
+
|
| 120 |
+
### Reading
|
| 121 |
+
|
| 122 |
+
- **76.71% sits between the two NVIDIA anchors and within the Wilson CI of both** (74.66–78.64 covers 0.763 and 0.7688): no measurable multimodal degradation for this NVFP4 QAD distillation vs BF16 at MMMU-Pro vision resolution on consumer hardware.
|
| 123 |
+
- Per-subject spread is wide (Electronics 96.6% / Finance 95.0% best; Music 45.0% / Diagnostics & Lab Medicine 50.0% worst) — consistent with the subject-level variance reported for other frontier models on MMMU-Pro.
|
| 124 |
+
- First published multimodal datapoint for this checkpoint; text-benchmark context in PR #6 (GSM8K 97.35%, GPQA 88.69% NV-protocol mean, MMLU-Pro 83.70%).
|