Qwen3.8-Flash-Next-Uncensored CED Projector

Linear projectors that make long-prompt prefill of Qwen3.8-Flash-Next 1.45–1.85× faster. They approximate the late layers for all but the last ~2K prompt tokens. Decode stays exact.

These files are not a model. An inference engine has to implement CED to use them. The code/ folder has the vLLM integration that produced every number on this page.

What it is

Qwen3.8-Flash-Next has 48 layers: 36 gated-delta-net (GDN) layers and 12 full-attention layers with a lightning indexer. Between layers it carries a multi-stream state: 4 hyper-connection streams × 2,560 = 10,240 values per token.

CED (approximate late-layer prefill) prefills a long prompt like this:

  1. Layers 0..S-1 run exactly on every prompt token. S is the split: 12, 16 or 24 here.
  2. A linear ridge map predicts each later layer's block input (layers S..47) from the layer-S multi-stream state (10,240 values + bias). A second map predicts the state the MTP drafter reads.
  3. The later layers only fill their caches from those predictions. GDN layers update their conv and recurrent state. Attention layers write K/V and the indexer keys. MoE, attention output and residual updates are skipped.
  4. The last prompt tokens run the full model exactly, and so does every decode step. The exact tail is at least 2,048 tokens, rounded to a cache block.

The approximated tokens never produce output. They only fill the cache that the exact tail and the answer read.

Why predict inputs instead of K/V

The projector predicts each layer's input. The model's own weights then compute the K/V, indexer keys and GDN state from it. That keeps the files smaller, and the design doesn't depend on one quant's cache-projection weights.

It was fit on one quant only: UD-IQ4_XS of the Uncensored (abliterated) variant. Other quants, the original Qwen weights and other fine-tunes are untested.

Files

File Split Layers run exactly Size (also its VRAM per GPU) Use
ced-projector-split16.safetensors 16 0–15 1.76 GiB Recommended
ced-projector-split24.safetensors 24 0–23 1.37 GiB Careful: least quality loss, smallest speedup
ced-projector-split12.safetensors 12 0–11 1.95 GiB Fastest, most quality loss
ced-projector-split16-msfa-int8.safetensors 16, multi-source 0–15 1.79 GiB (int8) Quality mode: less quality loss than split 16, less speedup

metadata.json has the checksums, the fit settings and the exact GGUF files the fit used.

Multi-source projector (quality mode, int8)

ced-projector-split16-msfa-int8.safetensors is a split-16 projector whose input is not only the layer-16 state but also the block inputs of the four full-attention layers before it (3, 7, 11 and 15): 10,240 + 4 × 2,560 = 20,480 values per token. The approximate chunk keeps those block inputs as layers 0–15 compute them and joins them with the layer-16 state in the order the file's sources metadata lists them.

It ships as int8 (symmetric, one bf16 scale per 128 input columns of a row, bf16 bias; 1.79 GiB). The bf16 fit is 3.5 GiB per GPU and did not fit next to R9V's expert cache on the second GPU. The int8 file holds exactly the tensors the grade below quantized at load, bit for bit.

In one GPU grade of both projectors, both loaded as int8, on the 16 prompts where long context clearly helps (7.5K–35K tokens, 2,048-token exact tail, eager grading server):

Projector Perplexity Long-context gain lost Prefill speedup (median) Projector time per 1K approximated tokens
split 16 ×1.049 17% 1.68× 24 ms
split 16 multi-source ×1.029 10% 1.55× 56 ms

A refit single-source control on the same training text graded like split 16, so the gain comes from the extra layers, not the extra data. R9V runs it with --ced quality; its default stays split 16. It has not yet run in R9V's compiled production server.

File contract: keys layer.16 … layer.47 (int8, [2560, 20480]) and final (int8, [10240, 20480]), each with scale.<key> (bf16, [rows, 160]) and bias.<key> (bf16, [rows]). Dequantize a map as int8 × scale per 128-column group, then y = x @ W.T + bias, where x is the sources concatenated in metadata order. code/model_overlay_multisource.py is the serving code; code/quantize_projector.py made the file from the bf16 fit (code/fit_projector.py --sources boundary,block_input_3,block_input_7,block_input_11,block_input_15).

Measured results

Hardware and stack for everything below:

  • 2× AMD Radeon AI PRO R9700, tensor parallel 2.
  • MTP with 4 draft tokens.
  • vLLM 0.26 with the R9V overlays (ROCm).
  • One request at a time, 4,096-token prefill chunks.

The ~94 GB model doesn't fit in 64 GB of VRAM, so part of the experts is served from a host-RAM cache. Absolute speeds reflect that.

Speed and quality per split

The grade used 22 prompts of 7.5K–19.5K tokens: Rust, Python, docs and prose. None of the projectors was fit on their text. The requested exact tail was 2,048 tokens, which came to 2,056–3,556 after block rounding.

Split Prefill speedup (median, prompts ≥ 12K tokens) Perplexity on long-context-dependent prompts ΔNLL, nats/token (95% CI) Long-context gain lost
24 1.45× (1.36–1.48, n=9) ×1.026 +0.026 ± 0.024 9%
16 1.70× (1.59–1.77, n=11) ×1.051 +0.050 ± 0.037 18%
12 1.85× (1.74–1.96, n=9) ×1.084 +0.081 ± 0.051 29%

How to read it:

  • Quality metric. Quality is the change in negative log-likelihood (NLL) of the last 512 prompt tokens, CED vs exact. It counts only "novel" tokens, the ones that can't be copied from earlier in the prompt. The table averages the 15 prompts where long context clearly helps (full context beats a context under 2,048 tokens). Perplexity = exp(ΔNLL).
  • Long-context gain lost. The share of the full-context NLL improvement that CED gives up.
  • Noise floor. Exact runs vary by 0.005 nats/token (median) between prompt orderings.
  • All 22 prompts. The mean change is smaller and noisier (+0.014 / +0.010 / +0.032 nats/token for splits 24 / 16 / 12), because several prompts barely use their long context.
  • How the speed was measured. These speedups come from the grading server: eager mode, one timed run per prompt and mode.
  • Where the time goes (split 16). An approximated chunk cost 261 ms per 1K tokens, vs 554 ms exact: 204 ms for input prep and layers 0–15, 23 ms for the projector, 29 ms for the cache fills.

Production config (split 16)

Same stack with the compiled model and decode CUDA graphs. CED was on for prompts of 8,192 tokens or more, with a 2,048-token exact tail.

  • Prefill: 1.70× median (1.62–1.76, n=6) on 12.5K–19.5K-token prompts. Approximated tokens cost 261–285 ms per 1K tokens, vs 512–555 ms exact.
  • A/B against a CED-off server: ~12K-token prompts went from 1,868 to 3,018 tok/s (+61.6%). ~4K-token prompts, below the 8,192 threshold, were unchanged (1,783 vs 1,789 tok/s).
  • MTP: on the first answer after a CED prefill, MTP accepted about 10% fewer tokens per step (0.895× of exact, range 0.778–0.951, n=6). A follow-up turn was back to normal (1.018×, n=2).
  • Decode: loading the projector costs no decode speed. Compared with a CED-off server, decode took 36.17 vs 36.77 ms/step on short prompts (n=6) and 38.98 vs 38.46 ms/step at ~8K context (n=4). Both are within run-to-run spread.
  • Checks:
    • An exact request after a CED request matches a fresh exact run bitwise. CED requests get their own prefix-cache namespace.
    • Multi-turn and the per-request tail setting work.
    • There were no recompiles after the first CED request.

Prefill sweep (BetterBench, split 16, CED forced on)

8 runs per size. tok/s = prompt tokens / time to first token, median.

Prompt tokens (median) Prefill tok/s Time to first token
1,557 1,900 0.82 s (shorter than the exact tail, so fully exact)
5,935 2,810 2.10 s
11,842 3,467 3.42 s
23,590 3,650 6.46 s
47,057 3,600 13.08 s

For reference, exact prefill on this stack runs at 1,783–1,868 tok/s for 4K–12K-token prompts (A/B above).

Limitations

  • One quant, one variant. The fit used UD-IQ4_XS of the Uncensored variant. Nothing else was tested. The exact GGUF files are listed in metadata.json. They were quantized locally from orcarouter's Q8_0 with Unsloth's UD-IQ4_XS recipe and are not published.
  • Text only. Prompts with images or video must be prefilled exactly. The included scheduler does this.
  • Narrow fit corpus. The fit text is English-heavy: Python source, Rust source, technical docs, and prose (license texts). That's about 55K sampled token positions for splits 12 and 16, and 74K for split 24. No chat transcripts, other languages or math. No document longer than ~20K tokens was in the fit or the quality grade.
  • Exact tail of 2,048 tokens or more. Only a 2,048-token tail was graded. The code accepts tails down to 512, but shorter tails are not validated.
  • No prompt logprobs. Approximated positions have no usable logits. Requests that need prompt logprobs should stay exact.
  • Measured one request at a time. The included scheduler approximates a prefill chunk only when it is the only request in the step. With concurrent requests, those chunks run exact.
  • Needs engine integration. Stock vLLM, llama.cpp and other engines can't use these files as they are.
  • No task-level tests. The grade measures NLL on held-out text. Task accuracy, such as tool calls or long-document QA, was not measured.

How to use

code/ has the vLLM 0.26 integration used for all the numbers above. See code/README.md. To build CED into another engine, this is the contract:

  • File. safetensors, bf16. It holds keys layer.S … layer.47, each [2560, 10241], plus final, [10240, 10241]. The split S is the lowest layer.* key, and every layer from S to 47 must be present. The metadata also records split, lambda, train_rows and heldout_mean_cosine.
  • Input x. The multi-stream hyper-connection state entering layer S, [tokens, 10240], bf16. Apply any pending combine from layer S-1 first.
  • Apply. y = x @ W[:, :10240].T + W[:, 10240]. The last column is the bias.
  • layer.L output, [tokens, 2560]. The predicted input to layer L's GDN or attention block, which is the output of that layer's attention hyper-connection mix. Run only the layer's cache writes from it:
    • GDN: q/k/v and b/a projections, then the conv and recurrent state update.
    • Attention: K/V (with RoPE as usual) into the KV cache, and the indexer key projection into the indexer cache.
  • final output, [tokens, 10240]. The predicted multi-stream state that enters the final hyper-connection mixer. Hand it to the MTP drafter for those positions.
  • Scheduling. An approximated chunk must end where the exact tail begins, and its logits are never sampled. Give CED requests their own prefix-cache namespace so exact requests never reuse their blocks.
  • Memory. Each tensor-parallel GPU holds the whole file. It is not sharded.

code/check_projector.py FILE checks a file against this contract on the CPU.

How it was made

  1. Capture. A vLLM server with the capture hooks in code/model_overlay.py prefilled English-heavy text:

    • Python standard-library source.
    • Rust source and Markdown design docs from the author's own inference-engine project.
    • Prose (open-source license texts).

    For each kept token it recorded:

    • the multi-stream state at the candidate split layers;
    • every later layer's block input;
    • the MTP drafter's input.

    Most training text kept every 4th token.

  2. Fit. code/fit_projector.py accumulates XᵀX and XᵀY in one streaming pass. It then solves a ridge regression per split, with a bias column and the ridge strength scaled to the mean Gram diagonal.

  3. Choosing λ. For splits 12 and 16, λ came from {0.03, 0.1, 0.3, 1.0}, picked by mean block-input cosine on three whole held-out documents. λ = 0.1 won for both. Split 24 is an earlier fit on different captures. There, held-out cosine slightly preferred λ = 0.03 (0.8176 vs 0.8164), but offline diagnostics favored 0.1, so it was refit with λ = 0.1.

Held-out mean cosine is the cosine between predicted and captured block input, median over tokens, averaged over late layers:

Split Held-out mean cosine
16 0.7396
12 0.7059
24 0.8164

The split-24 number comes from different held-out documents, so it isn't directly comparable with the other two.

Credit

The idea of predicting late-layer state from an early-layer boundary comes from PixelML's KVA projector for Qwen3.8-Flash-Next, which itself ports kishida's LLKVApprox. This project is an independent fit and design and is not affiliated with PixelML.

PixelML's published projector did not transfer to this model. Here its learned corrections lowered accuracy, so it isn't used.

Three things are different here:

  • a closed-form linear fit on this model's own activations;
  • it predicts each layer's input instead of cache values;
  • it adds a map for the MTP drafter's input.

Thanks to Qwen for the base model, orcarouter for the Uncensored variant, and Unsloth for the UD-IQ4_XS quantization recipe.

License

The projector weights come from Qwen3.8-Flash-Next's activations. They are released under the Qwen Community License 1.0 (LICENSE, the same text as the base model's). The Uncensored variant's publisher labels it Apache-2.0, but that variant is itself derived from Qwen's weights. This release follows the upstream Qwen license.

The code in code/ is Apache-2.0 (code/LICENSE). model_overlay.py and scheduler-ced.patch modify vLLM files, which are Apache-2.0, copyright contributors to the vLLM project.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Dyluhn/Qwen3.8-Flash-Next-Uncensored-CED-Projector

Adapter
(1)
this model