Quick Navigation Index

  1. Optimization History & Transparency Notice
  2. Quality Spectrum: APEX-I-NanoPlus vs. Standard Flat Quantizations
  3. Model Files & Technical Specifications
  4. Surgical Tensor Quantization Map (Audited from GGUF)
  5. Inference Quickstart
  6. CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
  7. Hardened Agentic Chat Template & Reasoning Effort
  8. Optional Support

EXPERIMENTAL PRE-RELEASE NOTICE: ENGLISH-ONLY CODING SPECIALIST

This model suite is quantized from Jab1718/qwen3.8-flash-coder-85gb-bf16, which is an intermediate experimental slice created using moe-slice (352 out of 512 routed experts were permanently pruned exclusively against English Python and SWE-bench calibration datasets).

  • English Coding Only: This model is strictly designed for programming, code completion, refactoring, and agentic tool-calling in English.
  • Severe Multilingual & General Degradation: Because conversational and multilingual experts were pruned and the upstream author has not yet released the recovery fine-tuning pass, this model severely degrades and outputs broken text in languages other than English (e.g., Spanish, French, German, etc.) or in general chit-chat.
  • Incompatible with Strata Engine: This model uses a 160-expert layout with decoupled n-gram tables; it is not compatible with Strata Engine (which requires the 512-expert monolith and 51B PLE tables). Run using stock llama.cpp (llama-server) or LM Studio.

Qwen3.8-Flash-Coder APEX-I-NanoPlus GGUF

The Next-Generation Frontier MoE · Extreme 18GB Footprint · Fast System RAM Streaming & Massive Context on 16GB–24GB VRAM

EXPLORE THE COMPLETE QWEN3.8 FLASH CODER LINEUP

These are complementary APEX-I releases, not alternate downloads of the same model:

EMPIRICAL BENCHMARK & QUALITY COMPARISON

Quantization Specification File Size (Disk) Memory Footprint (RAM/VRAM) Average BPW WikiText-2 Perplexity Delta PPL vs BF16 (%) Quality Tier Equivalent
Uncompressed BF16 Reference 85.30 GB (79.44 GiB) 79.44 GiB 16.00 BPW 30.0975 +/- 0.1200 Baseline (0.00%) Lossless Reference Baseline
APEX-I-MiniPlus V2.1 21.77 GB (20.27 GiB) 20.27 GiB 3.45 BPW 30.1495 +/- 1.0089 +0.0520 (+0.17%) Q5_K_L / Q6_K Tier Boundary
APEX-I-NanoPlus (CURRENT) 18.34 GB (17.08 GiB) 17.08 GiB 2.90 BPW 34.4199 +/- 1.1591 +4.3224 (+14.36%) Solid Q4_K_M Tier
Standard Flat Q3_K_S 20.41 GB 19.01 GiB 3.10 BPW aprox. 30.75 - 31.20 +0.65 a +1.10 (+2.9%) High syntax degradation
Generic APEX Mini (IQ2_S) 17.73 GB 16.51 GiB 2.50 BPW aprox. 31.60 - 33.10+ +1.50 a +3.00+ (+7.5%) Severe reasoning breakdown

Model Files & Technical Specifications

File Name File Size Memory Footprint BPW Description
Qwen3.8-Flash-Coder-85GB.APEX-I-NanoPlus.gguf 18.34 GB (17.08 GiB) 17.08 GiB 2.90 BPW Ultra-compact agentic coding MoE achieving solid Q4 quality with massive context capability
  • Base Model: Jab1718/qwen3.8-flash-coder-85gb-bf16
  • Parameters: 45.8B total (approx. 3.7B active per token: 10 routed experts + 1 shared expert + dense backbone; 4.9B with vocabulary embeddings)
  • Architecture: Qwen4ExpForCausalLM (48 hybrid layers, 160 MoE experts with 10 active)
  • Context Length: 262,144 tokens (native 256K)

Surgical Tensor Quantization Map (Audited from GGUF)

Layer Group Sub-Component / Tensor Qty Precision Engineering Rationale
Global Output Head output.weight 1 Q6_K Armored in high-precision Q6_K to preserve token classification.
Global Embeddings token_embd.weight 1 Q3_K Compact embedding representation across 248k vocabulary.
All Normalizations output_norm, attn_norm, ffn_norm, hc_norm, ssm_norm 146 F32 100% uncompressed numerical stability across all 48 layers.
Expert Routers blk.*.ffn_gate_inp, ffn_gate_inp_shexp 96 F32 100% uncompressed routing fidelity across 160 experts.
Attention Gates blk.*.attn_gate.weight (36 SSM Layers) 36 Q8_0 High-precision attention gating across hybrid DeltaNet recurrence layers.
Linear Attention Projections blk.*.attn_qkv.weight (36 SSM Layers) 36 Q3_K Efficient 3-bit quantization for SSM attention state inputs.
SSM Linear Output blk.*.ssm_out.weight (36 SSM Layers) 36 Q5_K 5-bit precision for linear state-space recurrence output.
Recurrent SSM Parameters blk.*.ssm_{a,alpha,beta,conv1d,dt,norm} 216 F32 Guarded in uncompressed FP32 to prevent DeltaNet recurrent state drift.
Periodic Sparse Attention blk.{3,7,...}.attn_{q,k,v}.weight (12 Layers) 36 Q4_K Quadratic attention checkpoints for deep retrieval.
Periodic Sparse Attention Output blk.{3,7,...}.attn_output.weight (12 Layers) 12 Q6_K Armored attention output projection over deep context.
QSA Sparse Attention Indexer blk.{3,7,...}.indexer.{q,k}_proj.weight 24 BF16 High-fidelity sparse indexer projections for query-stream attention routing.
QSA Indexer Norms blk.{3,7,...}.indexer.{q,k}_norm.weight 24 F32 Uncompressed indexer layer normalizations.
Shared Foundation Experts blk.*.ffn_down_shexp.weight (All 48 Layers) 48 Q5_0 ne0=640 in standard block-32; zero divisibility crashes.
Shared Foundation Experts blk.*.ffn_{gate,up}_shexp.weight (All 48 Layers) 96 Q5_K Preserves core coding knowledge active on 100% of tokens.
Highway Connections (Down/Inject) blk.*.hc_{attn,ffn}_{down,inject}.weight 192 Q6_K High-fidelity residual highway bypass.
Highway Connections (Up) blk.*.hc_{attn,ffn}_up.weight 96 Q5_0 ne0=320 in standard block-32 format.
Routed MoE Down-Projections blk.*.ffn_down_exps.weight (All 48 Layers) 48 IQ4_NL (36) / Q4_0 (12) Non-linear codebook quantization for SSM layers, Q4_0 for anchor layers.
Routed MoE Gate/Up Projections blk.*.ffn_{gate,up}_exps.weight (All 48 Layers) 96 IQ2_S (36) / IQ2_XXS (36) / IQ3_XXS (24) Scaled expert density: sub-2.5 BPW on deep layers, IQ3_XXS on anchor layers.
Residual Output Highway output_hc_{down,up}.weight 2 Q4_0 Low-rank residual highway projections at model termination.
Residual Output Highway Norm output_hc_norm.weight 1 F32 Final residual normalization anchor.

Inference Quickstart

llama-server \
  -m Qwen3.8-Flash-Coder-85GB.APEX-I-NanoPlus.gguf \
  --jinja \
  -ngl 99 \
  --ctx-size 65536 \
  --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 \
  --port 8080

CRITICAL ADVISORY FOR CODING WORKFLOWS: PREVENTING SYNTAX & TOKEN SWAPPING

Disable repeat penalties (repeat_penalty: 1.0, presence_penalty: 0.0, frequency_penalty: 0.0) and use --jinja to avoid syntax bracket substitutions.

Optional Support

Gold Ship dancing

If these releases are helpful, voluntary support is welcome at https://ko-fi.com/isvalorum.

Downloads last month
617
GGUF
Model size
43B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF

Quantized
(6)
this model

Collections including IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF