DeepSeek-V4-Flash-Vision-Exp-lossless-CSF

A lossless MXFP4-CSF container of deepseek-ai/DeepSeek-V4-Flash-Vision-Exp at revision 6821d6ad3681a4b137b066b76094fa82ebd0a380.

CSF stores the routed-expert block-scale planes in a compressed form. Nothing else changes: FP4 weight nibbles, every other tensor, the tensor names, dtypes, shapes and the source metadata are all byte-identical. Decoding is integer arithmetic on bytes, with no requantization or fitting. The original safetensors shards can be restored bit-exactly.

Routed experts (43 layers x 256) - MXFP4 (E2M1, UE8M0 scale per 32)
Attention, shared experts - FP8 E4M3 (128 x 128 blocks, UE8M0 scales)
Vision encoder (32 blocks), aligner, image tokens - BF16, unchanged
MTP layers (including their routed experts), embeddings, norms - source format

Routed-expert block scales - lossless MXFP4-CSF (row-base-offset1-u24-exceptions/1)

Sizes

  • Weight files: 167.82 GB in the source, 160.37 GB here (7.45 GB saved).
  • Compressed scales: 33,024 matrices (UE8M0 block scales of the 43 x 256 x 3 main-layer routed-expert projections): 8.66 GB -> 1.21 GB (13.9%).

Provenance

  • Source: deepseek-ai/DeepSeek-V4-Flash-Vision-Exp revision 6821d6ad3681a4b137b066b76094fa82ebd0a380, 48 index-referenced safetensors shards. During export, each source shard's SHA-256 was checked against its Hugging Face LFS hash.
  • Built with trellis-quant trellis_quant.lossless_scale_checkpoint (commit 60feca330087, family deepseek_v4_flash).
  • verification.json: every original shard was rebuilt from this container, and both the rebuilt and the stored files matched their SHA-256 (33,024 scale matrices, 48 shards, passed).

Layout

lil-mxfp4-csf-checkpoint/1, codec row-base-offset1-u24-exceptions/1:

  • tensors/ - the source shard names; each routed-expert scale <name> is stored as <name>.mxfp4_csf_fixed (uint8) plus <name>.mxfp4_csf_exceptions (uint32)
  • metadata/ - byte copies of the source's config, tokenizer, index, README and LICENSE
  • config.json - a copy of metadata/config.json at the top of the repository, where the Hub counts downloads; runtimes read metadata/
  • manifest.json, build-contract.json, receipts/ (per-shard source headers and hashes), verification.json
  • LICENSE, NOTICE, LICENSES/ (upstream license texts), CITATION.cff, REUSE.toml, and the SHA-256 of every file in SHA256SUMS and lil-manifest.json

There is no top-level model.safetensors.index.json, so a plain safetensors loader will not open this directory by mistake.

Serving

Use vLLM with the MXFP4-CSF reader: --quantization mxfp4_csf --load-format mxfp4_csf. Point vLLM at a serving directory that holds the files of metadata/, with config.json's quantization_config replaced by:

{
  "...": "every key of metadata/config.json quantization_config",
  "quant_method": "mxfp4_csf",
  "format_version": 1,
  "checkpoint_root": "/path/to/this/checkpoint"
}

The weights are read from checkpoint_root; the serving directory holds only metadata.

Serving needs a runtime whose MXFP4-CSF reader knows the deepseek_v4_flash family.

Verify or restore

PYTHONPATH=/path/to/trellis-quant python3 -m trellis_quant.lossless_scale_checkpoint verify \
  --checkpoint /path/to/this/checkpoint --workers 16
PYTHONPATH=/path/to/trellis-quant python3 -m trellis_quant.lossless_scale_checkpoint restore \
  --checkpoint /path/to/this/checkpoint --destination /path/to/fresh/dir --workers 16

verify rebuilds every original shard and compares SHA-256 values. restore writes the original shards and metadata files into a fresh directory, checking every hash before a file is published. The result is an ordinary copy of the source checkpoint.

Upstream license

The source model, deepseek-ai/DeepSeek-V4-Flash-Vision-Exp, is licensed by DeepSeek under the MIT License. Its license text is in LICENSES/MIT.txt, metadata/LICENSE keeps the source copy unchanged, and NOTICE names the upstream materials. This container is licensed under the Local Inference Lab License, Version 1.0 (LICENSE); see License and attribution below.

License and attribution

DeepSeek-V4-Flash-Vision-Exp-lossless-CSF is licensed under the Local Inference Lab License, Version 1.0, which reproduces the terms and conditions of the Apache License, Version 2.0, and adds conditions that restrict them. It is not the Apache License, and these files are not open source. Copyright (c) 2026 Local Inference Lab, Inc., a non-profit organization, 311 West Main Street, Grayson, Kentucky 41143. Revisions published up to and including revision 73a48c1453077266388407e7f7a70ec190a723cf were released under the MIT License, and copies of those revisions remain under it. The Local Inference Lab License, Version 1.0 applies to later revisions.

  • No reuploads. Do not upload, mirror, or redistribute these files, or any Substantially Similar Copy of them (including renamed, re-sharded, re-packaged, metadata-stripped, converted, or dequantized copies). Link to https://huggingface.co/local-inference-lab/DeepSeek-V4-Flash-Vision-Exp-lossless-CSF instead.
  • Attribution at the top. Every README, model card, or other landing page for a project, model, dataset, application, service, or distribution that contains, is derived from, or runs DeepSeek-V4-Flash-Vision-Exp-lossless-CSF must begin with this Attribution Notice as its first paragraph (a single title line may come before it):

This model is based on DeepSeek-V4-Flash-Vision-Exp-lossless-CSF by Local Inference Lab, Inc., a non-profit organization, available at https://huggingface.co/local-inference-lab/DeepSeek-V4-Flash-Vision-Exp-lossless-CSF. DeepSeek-V4-Flash-Vision-Exp-lossless-CSF is licensed under the Local Inference Lab License, Version 1.0.

Plain-text form:

This model is based on DeepSeek-V4-Flash-Vision-Exp-lossless-CSF by Local Inference Lab, Inc., a non-profit organization, available at https://huggingface.co/local-inference-lab/DeepSeek-V4-Flash-Vision-Exp-lossless-CSF. DeepSeek-V4-Flash-Vision-Exp-lossless-CSF is licensed under the Local Inference Lab License, Version 1.0.
  • Keep the marks. Do not remove the license metadata, the Identifying Marks ("Ні пуху, ні пера", LIL-CANARY-34C7-8086-0538-78FD), or the copyright notices embedded in these files.
  • Breach. A breach of the no-reupload or attribution terms ends the license immediately (Section 5.1), and Local Inference Lab, Inc. may ask hosting services to remove the material (Section 5.2).

File integrity: SHA256SUMS and lil-manifest.json list the SHA-256 of every file.

Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for local-inference-lab/DeepSeek-V4-Flash-Vision-Exp-lossless-CSF

Quantized
(29)
this model