Cosmos3-Nano: one unified NVFP4/FP8 model

This is a single joint Cosmos3 checkpoint, not a repository containing two independent models. One root tensor index covers the understanding and generation pathways, with the original Cosmos3 pipeline components.

  • Understanding/language pathway: 252 calibrated NVFP4 projections, native W4A4.
  • Generation pathway: 252 FP8 projections, native W8A8; first/last three diffusion iterations retain the upstream A16 activation policy.
  • Shared vision encoder, embeddings, language head, normalization and sensitive modules: BF16. VAE, scheduler and tokenizers are retained.
  • 3,181 indexed tensors; 14,569,813,872 tensor bytes (13.57 GiB), plus VAE, sound tokenizer and ancillary files.

The generation transformer consumes the NVFP4 understanding pathway inside the same model. This is not external text-prompt chaining.

Source and licensing

Original developer: NVIDIA. Original model: nvidia/Cosmos3-Nano, revision e59a53c25979a090fa8706c9acc0c254a6e89b92. Framework/plugin revision: cf5d68c00d97ccd2480a2320ed652b92dec63102.

The OpenMDW-1.1 license, original model card and original safety/bias/privacy/explainability notices are retained. Modifications: calibrated quantization, merging, canonical alias mapping, joint indices and mixed-format metadata. All calibrated NVFP4 and retained FP8 payload bytes match their source checkpoints; 498 shared BF16 tensors were verified byte-for-byte.

Verified execution and limitations

  • Standard vLLM freshly loaded the reasoner from this unified root and scored 101/115 on the same held-out requests. All 144 fused linears were native W4A4.
  • vLLM-Omni loaded the joint model and audited 252 NVFP4 understanding and 252 FP8 generation linears using native kernels.
  • Safety-enabled image and video generation passed: one 256 x 256 image and one 13-frame 256 x 256 video, with 12 diffusion iterations each.
  • Guardrail authorization was obtained; safety checks were not disabled.

This is bounded functional verification, not broad production-quality acceptance. The reasoner comparison remains experimental:

Model Transformed RoboVQA Synthetic Total
BF16 reasoner 96/110 5/5 101/115
Previous NVFP4 reasoner 94/110 5/5 99/115
Calibrated standalone reasoner, original run 95/110 5/5 100/115
Unified root, understanding-scoped metadata 96/110 5/5 101/115

RoboVQA evaluation has 101 distinct clips across 110 rows. Calibration used 272 rows, including 224 assistant continuations (24,973 assistant tokens). Robot dataset attribution: NVIDIA Cosmos Reason1 SFT/benchmark releases (CC-BY-4.0), RoboVQA (Apache-2.0); dataset revisions 4691712532849e9d5c550618c2941856c0ef112e and dc73eeac52e2fba0ae237b82af5317911c5b0366. No dataset media are redistributed here. A current standalone retest scored 102/115; a generic mixed-loader text run scored 97/115. The release uses the understanding-specific NVFP4 storage contract at the root and the true joint mixed contract in transformer/config.json. Tensor bytes are unchanged. Outputs are not bitwise identical across these serving configurations/runs, and this small transformed benchmark does not establish statistical improvement.

FP8 generation calibration used eight benign self-authored prompts and 400 denoising iterations at 256 pixels / 13 video frames. Changing the understanding precision changes generated outputs: a same-seed image comparison was performed, but pixel difference is not perceptual quality acceptance. Larger resolutions, long videos, audio, action outputs and broad prompt quality remain untested. No speedup or simultaneous-server memory guarantee is made.

Download

hf download thaddeusk/Cosmos3-Nano-NVFP4-Reasoner-FP8-Generator \
  --local-dir Cosmos3-Nano-Mixed

The repository is private, so authentication is required. Prefer a Linux filesystem working copy under WSL.

Text understanding from the same model

Use the pinned reasoner environment:

uv venv reasoner-env --python 3.12 --managed-python
uv pip sync --python reasoner-env/bin/python \
  Cosmos3-Nano-Mixed/requirements-reasoner.lock
reasoner-env/bin/vllm serve ./Cosmos3-Nano-Mixed \
  --quantization modelopt_fp4 --dtype bfloat16 \
  --max-model-len 8192 --gpu-memory-utilization 0.8

The native Cosmos3 loader filters generation weights automatically. Do not select a separate reasoner/ subfolder: the repository root is the joint model.

Image/video generation from the same model

vLLM-Omni 0.30 supports the mixed linear formats but its Cosmos3 diffusion-step policy validator does not recognize modelopt_mixed. The included launcher adds strict serialized mixed-format support in a fresh isolated SDK copy. It verifies FP8/NVFP4 serialized flags and group size 16. It does not alter installed packages, weight validation, or safety guardrails.

uv venv generation-env --python 3.12 --managed-python
uv pip sync --python generation-env/bin/python \
  Cosmos3-Nano-Mixed/requirements-generation.lock
generation-env/bin/python Cosmos3-Nano-Mixed/serve_unified.py \
  --runtime-dir ./fresh-omni-runtime-copy ./Cosmos3-Nano-Mixed \
  --host 127.0.0.1 --port 8012 --init-timeout 1800 \
  --enable-layerwise-offload

Choose a fresh runtime directory. Guardrail terms/access must be accepted for nvidia/Cosmos-1.0-Guardrail.

Example image request:

curl http://127.0.0.1:8012/v1/images/generations \
  -H "Content-Type: application/json" \
  -d '{"prompt":"A red wooden cube and a blue sphere on a white tabletop in daylight.","size":"256x256","num_inference_steps":12,"guidance_scale":6,"seed":42}'
Downloads last month
23
Safetensors
Model size
13B params
Tensor type
BF16
路
U8
路
F8_E4M3
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for thaddeusk/Cosmos3-Nano-NVFP4-Reasoner-FP8-Generator

Finetuned
(35)
this model