Instructions to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF # Run inference directly in the terminal: llama cli -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF # Run inference directly in the terminal: llama cli -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF # Run inference directly in the terminal: ./llama-cli -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Use Docker
docker model run hf.co/IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
- LM Studio
- Jan
- Ollama
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with Ollama:
ollama run hf.co/IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
- Unsloth Desktop
- Pi
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with Docker Model Runner:
docker model run hf.co/IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
- Lemonade
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Navigation Index
- Optimization History & Transparency Notice
- Quality Spectrum: APEX-I-NanoPlus vs. Standard Flat Quantizations
- Model Files & Technical Specifications
- Surgical Tensor Quantization Map (Audited from GGUF)
- Inference Quickstart
- CRITICAL: Coding Syntax & Repeat Penalty Advisory (Preventing Character Swapping)
- Hardened Agentic Chat Template & Reasoning Effort
- Optional Support
EXPERIMENTAL PRE-RELEASE NOTICE: ENGLISH-ONLY CODING SPECIALIST
This model suite is quantized from Jab1718/qwen3.8-flash-coder-85gb-bf16, which is an intermediate experimental slice created using
moe-slice(352 out of 512 routed experts were permanently pruned exclusively against English Python and SWE-bench calibration datasets).
- English Coding Only: This model is strictly designed for programming, code completion, refactoring, and agentic tool-calling in English.
- Severe Multilingual & General Degradation: Because conversational and multilingual experts were pruned and the upstream author has not yet released the recovery fine-tuning pass, this model severely degrades and outputs broken text in languages other than English (e.g., Spanish, French, German, etc.) or in general chit-chat.
- Incompatible with Strata Engine: This model uses a 160-expert layout with decoupled n-gram tables; it is not compatible with Strata Engine (which requires the 512-expert monolith and 51B PLE tables). Run using stock
llama.cpp(llama-server) or LM Studio.
Qwen3.8-Flash-Coder APEX-I-NanoPlus GGUF
The Next-Generation Frontier MoE · Extreme 18GB Footprint · Fast System RAM Streaming & Massive Context on 16GB–24GB VRAM
EXPLORE THE COMPLETE QWEN3.8 FLASH CODER LINEUP
These are complementary APEX-I releases, not alternate downloads of the same model:
- Qwen3.8-Flash-Coder APEX-I-MiniPlus-V2.1 — specialist agentic coding MoE tuned for Q5–Q6 quality (21.77 GB / 3.45 BPW).
- Qwen3.8-Flash-Coder APEX-I-NanoPlus — ultra-compact footprint achieving solid Q4 quality (18.34 GB / 2.90 BPW).
- Qwen3.8-Flash-Coder-85GB Lossless BF16 GGUF — uncompressed reference baseline (85.30 GB / 16.00 BPW).
EMPIRICAL BENCHMARK & QUALITY COMPARISON
Quantization Specification File Size (Disk) Memory Footprint (RAM/VRAM) Average BPW WikiText-2 Perplexity Delta PPL vs BF16 (%) Quality Tier Equivalent Uncompressed BF16 Reference 85.30 GB(79.44 GiB)79.44 GiB16.00 BPW 30.0975 +/- 0.1200Baseline (0.00%) Lossless Reference Baseline APEX-I-MiniPlus V2.1 21.77 GB(20.27 GiB)20.27 GiB3.45 BPW 30.1495 +/- 1.0089+0.0520 (+0.17%)Q5_K_L / Q6_K Tier Boundary APEX-I-NanoPlus (CURRENT) 18.34 GB(17.08 GiB)17.08 GiB2.90 BPW 34.4199 +/- 1.1591+4.3224 (+14.36%)Solid Q4_K_M Tier Standard Flat Q3_K_S 20.41 GB 19.01 GiB 3.10 BPW aprox. 30.75 - 31.20 +0.65 a +1.10 (+2.9%) High syntax degradation Generic APEX Mini (IQ2_S) 17.73 GB 16.51 GiB 2.50 BPW aprox. 31.60 - 33.10+ +1.50 a +3.00+ (+7.5%) Severe reasoning breakdown
Model Files & Technical Specifications
| File Name | File Size | Memory Footprint | BPW | Description |
|---|---|---|---|---|
Qwen3.8-Flash-Coder-85GB.APEX-I-NanoPlus.gguf |
18.34 GB (17.08 GiB) |
17.08 GiB |
2.90 BPW | Ultra-compact agentic coding MoE achieving solid Q4 quality with massive context capability |
- Base Model: Jab1718/qwen3.8-flash-coder-85gb-bf16
- Parameters: 45.8B total (approx. 3.7B active per token: 10 routed experts + 1 shared expert + dense backbone; 4.9B with vocabulary embeddings)
- Architecture:
Qwen4ExpForCausalLM(48 hybrid layers, 160 MoE experts with 10 active) - Context Length: 262,144 tokens (native 256K)
Surgical Tensor Quantization Map (Audited from GGUF)
| Layer Group | Sub-Component / Tensor | Qty | Precision | Engineering Rationale |
|---|---|---|---|---|
| Global Output Head | output.weight |
1 | Q6_K |
Armored in high-precision Q6_K to preserve token classification. |
| Global Embeddings | token_embd.weight |
1 | Q3_K |
Compact embedding representation across 248k vocabulary. |
| All Normalizations | output_norm, attn_norm, ffn_norm, hc_norm, ssm_norm |
146 | F32 |
100% uncompressed numerical stability across all 48 layers. |
| Expert Routers | blk.*.ffn_gate_inp, ffn_gate_inp_shexp |
96 | F32 |
100% uncompressed routing fidelity across 160 experts. |
| Attention Gates | blk.*.attn_gate.weight (36 SSM Layers) |
36 | Q8_0 |
High-precision attention gating across hybrid DeltaNet recurrence layers. |
| Linear Attention Projections | blk.*.attn_qkv.weight (36 SSM Layers) |
36 | Q3_K |
Efficient 3-bit quantization for SSM attention state inputs. |
| SSM Linear Output | blk.*.ssm_out.weight (36 SSM Layers) |
36 | Q5_K |
5-bit precision for linear state-space recurrence output. |
| Recurrent SSM Parameters | blk.*.ssm_{a,alpha,beta,conv1d,dt,norm} |
216 | F32 |
Guarded in uncompressed FP32 to prevent DeltaNet recurrent state drift. |
| Periodic Sparse Attention | blk.{3,7,...}.attn_{q,k,v}.weight (12 Layers) |
36 | Q4_K |
Quadratic attention checkpoints for deep retrieval. |
| Periodic Sparse Attention Output | blk.{3,7,...}.attn_output.weight (12 Layers) |
12 | Q6_K |
Armored attention output projection over deep context. |
| QSA Sparse Attention Indexer | blk.{3,7,...}.indexer.{q,k}_proj.weight |
24 | BF16 |
High-fidelity sparse indexer projections for query-stream attention routing. |
| QSA Indexer Norms | blk.{3,7,...}.indexer.{q,k}_norm.weight |
24 | F32 |
Uncompressed indexer layer normalizations. |
| Shared Foundation Experts | blk.*.ffn_down_shexp.weight (All 48 Layers) |
48 | Q5_0 |
ne0=640 in standard block-32; zero divisibility crashes. |
| Shared Foundation Experts | blk.*.ffn_{gate,up}_shexp.weight (All 48 Layers) |
96 | Q5_K |
Preserves core coding knowledge active on 100% of tokens. |
| Highway Connections (Down/Inject) | blk.*.hc_{attn,ffn}_{down,inject}.weight |
192 | Q6_K |
High-fidelity residual highway bypass. |
| Highway Connections (Up) | blk.*.hc_{attn,ffn}_up.weight |
96 | Q5_0 |
ne0=320 in standard block-32 format. |
| Routed MoE Down-Projections | blk.*.ffn_down_exps.weight (All 48 Layers) |
48 | IQ4_NL (36) / Q4_0 (12) |
Non-linear codebook quantization for SSM layers, Q4_0 for anchor layers. |
| Routed MoE Gate/Up Projections | blk.*.ffn_{gate,up}_exps.weight (All 48 Layers) |
96 | IQ2_S (36) / IQ2_XXS (36) / IQ3_XXS (24) |
Scaled expert density: sub-2.5 BPW on deep layers, IQ3_XXS on anchor layers. |
| Residual Output Highway | output_hc_{down,up}.weight |
2 | Q4_0 |
Low-rank residual highway projections at model termination. |
| Residual Output Highway Norm | output_hc_norm.weight |
1 | F32 |
Final residual normalization anchor. |
Inference Quickstart
llama-server \
-m Qwen3.8-Flash-Coder-85GB.APEX-I-NanoPlus.gguf \
--jinja \
-ngl 99 \
--ctx-size 65536 \
--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.00 \
--port 8080
CRITICAL ADVISORY FOR CODING WORKFLOWS: PREVENTING SYNTAX & TOKEN SWAPPING
Disable repeat penalties (
repeat_penalty: 1.0,presence_penalty: 0.0,frequency_penalty: 0.0) and use--jinjato avoid syntax bracket substitutions.
Optional Support
If these releases are helpful, voluntary support is welcome at https://ko-fi.com/isvalorum.
- Downloads last month
- 617
We're not able to determine the quantization variants.
Model tree for IsValorum/Qwen3.8-Flash-Coder-APEX-I-NanoPlus-GGUF
Base model
Qwen/Qwen3.8-Flash-Next