Add community multimodal evaluation: MMMU Pro vision-only 76.71% (temp 1.0/top_p 0.95 NV-protocol, 1730 items, 4x RTX PRO 6000 Karmic Kraken stack)

#7
by malaiwah - opened
Files changed (1) hide show
  1. README.md +31 -0
README.md CHANGED
@@ -91,3 +91,34 @@ NVFP4/MXFP8 ModelOpt layout, including per-expert activation scales.
91
  The tokenizer, upstream chat template, generation settings, vision assets
92
  and MTP tensors are included. KV cache and runtime buffers require memory
93
  in addition to the model weights.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
91
  The tokenizer, upstream chat template, generation settings, vision assets
92
  and MTP tensors are included. KV cache and runtime buffers require memory
93
  in addition to the model weights.
94
+
95
+ ## Community multimodal evaluation (2026-10-04) — MMMU Pro (vision-only), NV-protocol sampling
96
+
97
+ Independent multimodal accuracy measurement of checkpoint `175ae8ce` contributed by [@malaiwah](https://huggingface.co/malaiwah). **Text benchmarks (GSM8K / GPQA Diamond / MMLU-Pro) are covered separately in [PR #6](https://huggingface.co/local-inference-lab/GLM-5.3-Flash-NVFP4/discussions/6)** — this section adds the first multimodal datapoint for the exact checkpoint this card describes. Raw per-item JSONL available on request.
98
+
99
+ ### Methodology
100
+
101
+ - Dataset: [MMMU/MMMU_Pro](https://huggingface.co/datasets/MMMU/MMMU_Pro), config `vision` (question + all options rendered into the image; 10-choice), official `test` split, N = 1,730 (full coverage, 0 request errors after a resume pass).
102
+ - Prompt: official `mmmu-pro/prompts.yaml` **CoT-vision** template ("Write out the multiple-choice question in the image and then solve it… 'Answer: $LETTER'").
103
+ - Parsing: mirrors the official `parse_multi_choice_response` (last `Answer:` → unique option-letter match, layered fallbacks); 3 parse failures, each being an item that hit the `max_tokens` cap in a reasoning loop and never emitted the marker — counted as incorrect, matching how the sibling-table protocol treats truncations.
104
+ - Sampling: **temperature 1.0 / top_p 0.95** (sibling-table protocol, identical to the card's NVIDIA eval settings), `max_tokens` 327,680, no other overrides — server defaults enforced. 16-way concurrent requests against the OpenAI-compatible endpoint; the vision tower passes through unquantized (quantization applies to routed-expert/dense-MLP linear ops only), so this measures the end-to-end multimodal pipeline + LLM-side quantization.
105
+ - Token profile: mean 3,470 completion tokens (thinking on, effort `high`), p95 11,403, 6.0M completion tokens total.
106
+ - Runner: custom minimal OpenAI-compatible client (per-item JSONL records: pred, latency, completion tokens), smoke-validated before launch.
107
+
108
+ ### Serving stack
109
+
110
+ Identical to the text-evaluation section in PR #6: 4× RTX PRO 6000 Blackwell, vLLM `0.1.dev22021+g93dabce32`, digest-pinned `karmic-kraken-beta` image, MTP ×3 + LMCache engine-driven, ModelOpt 0.45.0, 1M context.
111
+
112
+ ### Results
113
+
114
+ | Benchmark | Items | Sampling | Accuracy | Wilson 95% CI | Notes |
115
+ | --- | --- | --- | --- | --- | --- |
116
+ | MMMU Pro (vision) | 1730 | temp 1.0 / top_p 0.95, max_tokens 327,680 | **76.71%** (1327/1730) | 74.66–78.64 | 0 request errors; 3 max-token cap hits (reasoning loops, counted incorrect) |
117
+
118
+ Reference anchors (NVIDIA's card table for this model family, GB200): BF16 baseline **0.7688**, NVIDIA ModelOpt-NVFP4 **0.763** (different quantization recipe — ModelOpt PTQ, not this QAD distillation).
119
+
120
+ ### Reading
121
+
122
+ - **76.71% sits between the two NVIDIA anchors and within the Wilson CI of both** (74.66–78.64 covers 0.763 and 0.7688): no measurable multimodal degradation for this NVFP4 QAD distillation vs BF16 at MMMU-Pro vision resolution on consumer hardware.
123
+ - Per-subject spread is wide (Electronics 96.6% / Finance 95.0% best; Music 45.0% / Diagnostics & Lab Medicine 50.0% worst) — consistent with the subject-level variance reported for other frontier models on MMMU-Pro.
124
+ - First published multimodal datapoint for this checkpoint; text-benchmark context in PR #6 (GSM8K 97.35%, GPQA 88.69% NV-protocol mean, MMLU-Pro 83.70%).