Instructions to use local-inference-lab/Qwen3.8-Flash-Next-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use local-inference-lab/Qwen3.8-Flash-Next-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="local-inference-lab/Qwen3.8-Flash-Next-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("local-inference-lab/Qwen3.8-Flash-Next-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("local-inference-lab/Qwen3.8-Flash-Next-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use local-inference-lab/Qwen3.8-Flash-Next-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "local-inference-lab/Qwen3.8-Flash-Next-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "local-inference-lab/Qwen3.8-Flash-Next-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/local-inference-lab/Qwen3.8-Flash-Next-NVFP4
- SGLang
How to use local-inference-lab/Qwen3.8-Flash-Next-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "local-inference-lab/Qwen3.8-Flash-Next-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "local-inference-lab/Qwen3.8-Flash-Next-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "local-inference-lab/Qwen3.8-Flash-Next-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "local-inference-lab/Qwen3.8-Flash-Next-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use local-inference-lab/Qwen3.8-Flash-Next-NVFP4 with Docker Model Runner:
docker model run hf.co/local-inference-lab/Qwen3.8-Flash-Next-NVFP4
Expand evaluation table with NVIDIA references and independent results
Adds NVIDIA FP8/NVFP4 reference scores and an independently measured QAD row. The notes identify protocol differences and link the complete Mia runtime report. No model files or metadata are changed.
I simplified the proposed model-card change so that the Evaluation section now contains only one compact comparison table. The methodology and comparison caveats are kept here for maintainers instead of making the card verbose.
Methodology notes
- The NVIDIA NVFP4 values are references copied from the NVIDIA model card, not reruns on the independent host. NVIDIA did not publish an identical harness for every score.
- The
local-inference-lab QAD — model-card resultsrow preserves the existing author-reported scores: AA-LCR 79.4, GPQA Diamond 89.9, and Tool Eval Bench 91. - The independent run used checkpoint revision
7c4f1bc1a2d6847e0cbc01ac6b823f00251de8ddwith MiaAI-Lab runtime commite74e7af934c19799eca5a2c0dc9f97bc6decd784and vLLM0.1.dev20073+g8e685d198on one ASUS Ascent GX10. - Serving used the native 262,144-token context, FP8 KV cache, BF16 recurrent state, MTP3, and
MAX_NUM_SEQS=8. - GPQA Diamond: 181/198 = 91.41%, 3,188,456 total tokens, concurrency 3.
- IFBench: 242/294 = 82.31%, 2,483,251 total tokens, one response per prompt, concurrency 4. NVIDIA's protocol may use a different number of repeats.
- Tool Eval Bench: 165/174 points = 95/100 across all 88 hard-mode scenarios. One infrastructure scenario (TC-45) was excluded because the endpoint did not enforce
tool_choice=required. - AA-LCR 79.4 was not rerun independently.
The complete runtime/evaluation report is linked from the independent-results row: https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark/pull/61