CYBER-FROST-3.8 GGUF

The queued trunk quants are uploaded. Each one passed a short Vulkan generation test. CYBER-FROST-3.8-Q2_K_S.gguf now carries the MTP head inside the file. The separate mtp- files still do not load.

Source checkpoint: Blackfrost-AI/CYBER-FROST-3.8-BF16. Architecture qwen4exp (Qwen4ExpForConditionalGeneration), from Qwen/Qwen3.8-Flash-Next. Text stack only. The vision tower was not converted.

About 180B parameters, about 6B active. 48 layers, 512 experts, 10 experts per token, one shared expert. Configured context is 262144. Smoke tests used 512 tokens of context. Long context is not tested.

The checkpoint includes one native MTP draft block, trained before the final trunk behavior pass. That block is grafted into CYBER-FROST-3.8-Q2_K_S.gguf as layer 48, Q8_0, from mtp-CYBER-FROST-3.8-Q8_0.gguf. The other trunks do not have it yet. The standalone mtp- files still do not load.

Files

Trunk files are the model. mtp- files are only a draft head. You do not chat with an mtp- file by itself.

file role size test
CYBER-FROST-3.8-Q4_K_M.gguf flat 4-bit 111.43 GB uploaded, Vulkan smoke passed
CYBER-FROST-3.8-UD-Q4_K_XL.gguf dynamic 4-bit 123.10 GB uploaded, Vulkan smoke passed
CYBER-FROST-3.8-UD-IQ4_XS.gguf dynamic IQ4_XS 121.84 GB uploaded, Vulkan smoke passed
CYBER-FROST-3.8-MXFP4_MOE.gguf expert MXFP4, other tensors Q8_0 96.65 GB uploaded, Vulkan smoke passed
CYBER-FROST-3.8-Q3_K_M.gguf 3-bit 88.56 GB uploaded, Vulkan smoke passed
CYBER-FROST-3.8-Q3_K_S.gguf 3-bit small 88.56 GB uploaded, Vulkan smoke passed
CYBER-FROST-3.8-Q2_K.gguf 2-bit 80.08 GB uploaded, Vulkan smoke passed
CYBER-FROST-3.8-Q2_K_S.gguf 2-bit small, in-file Q8 MTP head 80.08 GB uploaded, Vulkan MTP smoke passed
mtp-CYBER-FROST-3.8-Q8_0.gguf draft head, Q8_0 4.13 GB uploaded, failed to load
mtp-CYBER-FROST-3.8-Q4_K_M.gguf draft head, Q4_K_M 2.62 GB uploaded, not loaded
mtp-CYBER-FROST-3.8-Q4_0.gguf draft head, Q4_0 not built

Quant choices

No importance matrix. These dynamic files are a per-layer type mix. They are not Unsloth Dynamic 3.0, which also depends on their calibration set.

Q4_K_M is one type for most 2D weights. output.weight and token_embd.weight are Q6_K. Attention v is Q5_K.

UD-Q4_K_XL keeps the first and last layers, and attention and SSM tensors, at Q5_K or Q6_K. Output and token embeddings are Q8_0. The GGUF file-type field still says Q4_K_M.

UD-IQ4_XS uses that same layer split with IQ4_XS where the row width allows it, and IQ4_NL where it does not. No imatrix, so this will not match a calibrated IQ4_XS.

MXFP4_MOE puts MXFP4 on expert tensors and the n-gram table only. Attention, SSM, output, and embeddings stay Q8_0.

On the 3-bit and 2-bit files, two tensors cannot use a 256-wide block, so they are Q4_0:

  • Expert down projections, row length 640.
  • The n-gram table per_layer_token_embd, row length 160.

On the 4-bit files those two tensors are Q5_0 (flat Q4_K_M) or Q5_1 (dynamic files).

ssm_conv1d.weight is stored as F16. The Vulkan SSM conv shader reads float. Coherent text on the two passing files used a build that casts that kernel to F32 before the conv. A build that feeds the F16 kernel straight into that shader produces garbage.

Norms and the router stay float32.

Run a trunk

Tested on llama.cpp b1-4da6337, Radeon 680M, Vulkan, with async uploads disabled. The file is mmap'd. Experts and the n-gram table stay on CPU and fault in from disk. GTT stayed near 45 MB. Fit is off so the loader does not try to fill the GPU with the expert pool.

KV cache is f16. Quantized KV crashes this architecture.

Smoke context was 512. Prompt speed was about 0.5 to 0.7 tok/s. Generation was about 1.0 to 1.1 tok/s.

GGML_VK_DISABLE_ASYNC=1 llama-cli \
  -m CYBER-FROST-3.8-Q4_K_M.gguf \
  -ot "per_layer_token_embd=CPU,exps=CPU" \
  -cmoe \
  -ngl 99 \
  -c 512 \
  -lm mmap -fit off \
  -ctk f16 -ctv f16 \
  --jinja

UD-Q4_K_XL uses the same flags.

Draft head

CYBER-FROST-3.8-Q2_K_S.gguf is the only trunk with the head in the file. Run it with --spec-type draft-mtp. Do not pass -md.

GGML_VK_DISABLE_ASYNC=1 llama-cli \
  -m CYBER-FROST-3.8-Q2_K_S.gguf \
  -ot "per_layer_token_embd=CPU,exps=CPU" \
  -cmoe \
  -ngl 99 \
  -c 512 \
  -lm mmap -fit off \
  -ctk f16 -ctv f16 \
  --spec-type draft-mtp --spec-draft-n-max 2

This needs the qwen4exp MTP graph in the local llama.cpp build on this machine. The head reads the wide residual from before the final mixer, not a separate draft file.

mtp-CYBER-FROST-3.8-Q8_0.gguf still does not load next to a trunk. llama.cpp stops because output_hc_norm.weight is missing. The mixer weights are stored as blk.48.nextn.hc_head_norm.weight. The Q4_K_M draft was exported the same way. Do not pass either sidecar with -md.

Test log

Prompt: write a Python 3 function named add that takes two ints and returns their sum. Pass means the reply contains that function and a return.

file date build backend result
CYBER-FROST-3.8-Q4_K_M.gguf 2026-09-28 b1-4da6337 Vulkan, mmap, experts and n-gram table on CPU pass, def add with a return, 1.1 tok/s
CYBER-FROST-3.8-UD-Q4_K_XL.gguf 2026-09-28 b1-4da6337 Vulkan, mmap, experts and n-gram table on CPU pass, same function, 1.0 tok/s
CYBER-FROST-3.8-UD-IQ4_XS.gguf 2026-09-28 b1-4da6337 Vulkan, mmap, experts and n-gram table on CPU pass, same function, 0.9 tok/s
CYBER-FROST-3.8-MXFP4_MOE.gguf 2026-09-28 b1-4da6337 Vulkan, mmap, experts and n-gram table on CPU pass, same function, 1.5 tok/s
CYBER-FROST-3.8-Q3_K_M.gguf 2026-09-28 b1-4da6337 Vulkan, mmap, experts and n-gram table on CPU pass, same function, 1.4 tok/s
CYBER-FROST-3.8-Q3_K_S.gguf 2026-09-28 b1-4da6337 Vulkan, mmap, experts and n-gram table on CPU pass, same function, 1.4 tok/s
CYBER-FROST-3.8-Q2_K.gguf 2026-09-28 b1-4da6337 Vulkan, mmap, experts and n-gram table on CPU pass, same function, 1.9 tok/s
CYBER-FROST-3.8-Q2_K_S.gguf 2026-09-28 b1-4da6337 Vulkan, in-file MTP, --spec-type draft-mtp pass, def add with a return, 0.9 tok/s
mtp-CYBER-FROST-3.8-Q8_0.gguf with Q4_K_M 2026-09-28 b1-4da6337 Vulkan, then CPU fail, draft missing output_hc_norm.weight

License

Qwen Community License 1.0, same terms as the source checkpoint. See LICENSE.

Downloads last month
2,138
GGUF
Model size
177B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 0xSojalSec/CYBER-FROST-3.8-GGUF

Quantized
(10)
this model