Instructions to use AutomatosX/AX-DeepSeek-OCR-2-MLX-AXQ-MXFP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use AutomatosX/AX-DeepSeek-OCR-2-MLX-AXQ-MXFP8 with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("AutomatosX/AX-DeepSeek-OCR-2-MLX-AXQ-MXFP8") config = load_config("AutomatosX/AX-DeepSeek-OCR-2-MLX-AXQ-MXFP8") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Runtime format audit (2026-10-06)
No quantization-container correction was needed. This family has no n-gram tensors; no n-gram file or declaration was added. This remote format audit does not grant an oMLX/MTPLX runtime profile or a successful load/generation claim.
See runtime_audit.json for pinned config/index/header bindings, architecture, physical-format findings, and applied corrections. This is development evidence; no quality, MTP exactness, speed, or certification claim is added. Historical evidence stays bound to its original revision.
AX-DeepSeek-OCR-2-MLX-AXQ-MXFP8
Development AXQuant (AXQ) MXFP8 language MLX pack of
deepseek-ai/DeepSeek-OCR-2
@ aaa02f3811945a91062062994c5c4a3f4c0af2b0.
Language experts/attention/MLP and embeddings at native MXFP8 (group 32, physical repack of unrefined 8-bit allocations); vision towers, norms, MoE routers, and LM head BF16-preserved. Measured total 9.67 BPW.
Converted via MLX-VLM deepseekocr_2 from
mlx-community/DeepSeek-OCR-2-bf16
@ 9946f9ac306378a3e6a86cad7d7f8be8e536f092.
Built with AXQuant 1.9.0 (plan-manual + convert --q-mode mxfp8,
--allow-unmeasured). Load smoke: mlx-vlm loads 685 modules. MXFP8 is a
physical repacking mode, not a measured planner method.
Claims
| Claim | Status |
|---|---|
| AXQuant architecture-prior / development quant | Yes |
| OCR accuracy / document-bench scores | Not claimed |
| Vision optimization | No — SAM + Qwen2 encoder + projector BF16-preserved |
| Certified release | No |
Attribution
Base weights © DeepSeek AI (Apache-2.0). Quantization by AXQuant (development).
- Downloads last month
- 39
8-bit
Model tree for AutomatosX/AX-DeepSeek-OCR-2-MLX-AXQ-MXFP8
Base model
deepseek-ai/DeepSeek-OCR-2