Runtime format audit (2026-10-06)

No quantization-container correction was needed. This family has no n-gram tensors; no n-gram file or declaration was added. This remote format audit does not grant an oMLX/MTPLX runtime profile or a successful load/generation claim.

See runtime_audit.json for pinned config/index/header bindings, architecture, physical-format findings, and applied corrections. This is development evidence; no quality, MTP exactness, speed, or certification claim is added. Historical evidence stays bound to its original revision.

AX-DeepSeek-OCR-2-MLX-AXQ-MXFP8

Development AXQuant (AXQ) MXFP8 language MLX pack of deepseek-ai/DeepSeek-OCR-2 @ aaa02f3811945a91062062994c5c4a3f4c0af2b0.

Language experts/attention/MLP and embeddings at native MXFP8 (group 32, physical repack of unrefined 8-bit allocations); vision towers, norms, MoE routers, and LM head BF16-preserved. Measured total 9.67 BPW.

Converted via MLX-VLM deepseekocr_2 from mlx-community/DeepSeek-OCR-2-bf16 @ 9946f9ac306378a3e6a86cad7d7f8be8e536f092.

Built with AXQuant 1.9.0 (plan-manual + convert --q-mode mxfp8, --allow-unmeasured). Load smoke: mlx-vlm loads 685 modules. MXFP8 is a physical repacking mode, not a measured planner method.

Claims

Claim Status
AXQuant architecture-prior / development quant Yes
OCR accuracy / document-bench scores Not claimed
Vision optimization No — SAM + Qwen2 encoder + projector BF16-preserved
Certified release No

Attribution

Base weights © DeepSeek AI (Apache-2.0). Quantization by AXQuant (development).

Downloads last month
39
Safetensors
Model size
3B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AutomatosX/AX-DeepSeek-OCR-2-MLX-AXQ-MXFP8

Quantized
(10)
this model

Collections including AutomatosX/AX-DeepSeek-OCR-2-MLX-AXQ-MXFP8