Restore text_config.ple_embedding_dtype = nvfp4

#7
by tingtu0721 - opened

The step-5500 config.json dropped ple_embedding_dtype. The PLE n-gram table in this checkpoint is still NVFP4 (ngram_embedding.shard_*.weight + weight_scale), but without the key the qwen4_exp PLE layer in ghcr.io/nvidia-ai-iot/vllm:qwen3.8-next-jetson-thor builds it as BF16 (~97 GiB) and OOMs during model construction on Jetson AGX Thor. With this one-line change the model loads (97.84 GiB) and serves normally. Details: discussion #6.

lukealonso changed pull request status to merged

Sign up or log in to comment