Expand evaluation table with NVIDIA references and independent results

#4
by siertum - opened

Adds NVIDIA FP8/NVFP4 reference scores and an independently measured QAD row. The notes identify protocol differences and link the complete Mia runtime report. No model files or metadata are changed.

I simplified the proposed model-card change so that the Evaluation section now contains only one compact comparison table. The methodology and comparison caveats are kept here for maintainers instead of making the card verbose.

Methodology notes

  • The NVIDIA NVFP4 values are references copied from the NVIDIA model card, not reruns on the independent host. NVIDIA did not publish an identical harness for every score.
  • The local-inference-lab QAD — model-card results row preserves the existing author-reported scores: AA-LCR 79.4, GPQA Diamond 89.9, and Tool Eval Bench 91.
  • The independent run used checkpoint revision 7c4f1bc1a2d6847e0cbc01ac6b823f00251de8dd with MiaAI-Lab runtime commit e74e7af934c19799eca5a2c0dc9f97bc6decd784 and vLLM 0.1.dev20073+g8e685d198 on one ASUS Ascent GX10.
  • Serving used the native 262,144-token context, FP8 KV cache, BF16 recurrent state, MTP3, and MAX_NUM_SEQS=8.
  • GPQA Diamond: 181/198 = 91.41%, 3,188,456 total tokens, concurrency 3.
  • IFBench: 242/294 = 82.31%, 2,483,251 total tokens, one response per prompt, concurrency 4. NVIDIA's protocol may use a different number of repeats.
  • Tool Eval Bench: 165/174 points = 95/100 across all 88 hard-mode scenarios. One infrastructure scenario (TC-45) was excluded because the endpoint did not enforce tool_choice=required.
  • AA-LCR 79.4 was not rerun independently.

The complete runtime/evaluation report is linked from the independent-results row: https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark/pull/61

lukealonso changed pull request status to merged

Sign up or log in to comment