Instructions to use 0xSojalSec/CYBER-FROST-3.8-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use 0xSojalSec/CYBER-FROST-3.8-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf 0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf 0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf 0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf 0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf 0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./llama-cli -hf 0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf 0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf 0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL
Use Docker
docker model run hf.co/0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL
- LM Studio
- Jan
- vLLM
How to use 0xSojalSec/CYBER-FROST-3.8-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "0xSojalSec/CYBER-FROST-3.8-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "0xSojalSec/CYBER-FROST-3.8-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL
- Ollama
How to use 0xSojalSec/CYBER-FROST-3.8-GGUF with Ollama:
ollama run hf.co/0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL
- Unsloth Desktop
- Pi
How to use 0xSojalSec/CYBER-FROST-3.8-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use 0xSojalSec/CYBER-FROST-3.8-GGUF with Docker Model Runner:
docker model run hf.co/0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL
- Lemonade
How to use 0xSojalSec/CYBER-FROST-3.8-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull 0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL
Run and chat with the model
lemonade run user.CYBER-FROST-3.8-GGUF-UD-Q4_K_XL
List all available models
lemonade list
- Hermes Agent
How to use 0xSojalSec/CYBER-FROST-3.8-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default 0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use 0xSojalSec/CYBER-FROST-3.8-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "0xSojalSec/CYBER-FROST-3.8-GGUF:UD-Q4_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
CYBER-FROST-3.8 GGUF
The queued trunk quants are uploaded. Each one passed a short Vulkan generation test. CYBER-FROST-3.8-Q2_K_S.gguf now carries the MTP head inside the file. The separate mtp- files still do not load.
Source checkpoint: Blackfrost-AI/CYBER-FROST-3.8-BF16. Architecture qwen4exp (Qwen4ExpForConditionalGeneration), from Qwen/Qwen3.8-Flash-Next. Text stack only. The vision tower was not converted.
About 180B parameters, about 6B active. 48 layers, 512 experts, 10 experts per token, one shared expert. Configured context is 262144. Smoke tests used 512 tokens of context. Long context is not tested.
The checkpoint includes one native MTP draft block, trained before the final trunk behavior pass. That block is grafted into CYBER-FROST-3.8-Q2_K_S.gguf as layer 48, Q8_0, from mtp-CYBER-FROST-3.8-Q8_0.gguf. The other trunks do not have it yet. The standalone mtp- files still do not load.
Files
Trunk files are the model. mtp- files are only a draft head. You do not chat with an mtp- file by itself.
| file | role | size | test |
|---|---|---|---|
CYBER-FROST-3.8-Q4_K_M.gguf |
flat 4-bit | 111.43 GB | uploaded, Vulkan smoke passed |
CYBER-FROST-3.8-UD-Q4_K_XL.gguf |
dynamic 4-bit | 123.10 GB | uploaded, Vulkan smoke passed |
CYBER-FROST-3.8-UD-IQ4_XS.gguf |
dynamic IQ4_XS | 121.84 GB | uploaded, Vulkan smoke passed |
CYBER-FROST-3.8-MXFP4_MOE.gguf |
expert MXFP4, other tensors Q8_0 | 96.65 GB | uploaded, Vulkan smoke passed |
CYBER-FROST-3.8-Q3_K_M.gguf |
3-bit | 88.56 GB | uploaded, Vulkan smoke passed |
CYBER-FROST-3.8-Q3_K_S.gguf |
3-bit small | 88.56 GB | uploaded, Vulkan smoke passed |
CYBER-FROST-3.8-Q2_K.gguf |
2-bit | 80.08 GB | uploaded, Vulkan smoke passed |
CYBER-FROST-3.8-Q2_K_S.gguf |
2-bit small, in-file Q8 MTP head | 80.08 GB | uploaded, Vulkan MTP smoke passed |
mtp-CYBER-FROST-3.8-Q8_0.gguf |
draft head, Q8_0 | 4.13 GB | uploaded, failed to load |
mtp-CYBER-FROST-3.8-Q4_K_M.gguf |
draft head, Q4_K_M | 2.62 GB | uploaded, not loaded |
mtp-CYBER-FROST-3.8-Q4_0.gguf |
draft head, Q4_0 | not built |
Quant choices
No importance matrix. These dynamic files are a per-layer type mix. They are not Unsloth Dynamic 3.0, which also depends on their calibration set.
Q4_K_M is one type for most 2D weights. output.weight and token_embd.weight are Q6_K. Attention v is Q5_K.
UD-Q4_K_XL keeps the first and last layers, and attention and SSM tensors, at Q5_K or Q6_K. Output and token embeddings are Q8_0. The GGUF file-type field still says Q4_K_M.
UD-IQ4_XS uses that same layer split with IQ4_XS where the row width allows it, and IQ4_NL where it does not. No imatrix, so this will not match a calibrated IQ4_XS.
MXFP4_MOE puts MXFP4 on expert tensors and the n-gram table only. Attention, SSM, output, and embeddings stay Q8_0.
On the 3-bit and 2-bit files, two tensors cannot use a 256-wide block, so they are Q4_0:
- Expert down projections, row length 640.
- The n-gram table
per_layer_token_embd, row length 160.
On the 4-bit files those two tensors are Q5_0 (flat Q4_K_M) or Q5_1 (dynamic files).
ssm_conv1d.weight is stored as F16. The Vulkan SSM conv shader reads float. Coherent text on the two passing files used a build that casts that kernel to F32 before the conv. A build that feeds the F16 kernel straight into that shader produces garbage.
Norms and the router stay float32.
Run a trunk
Tested on llama.cpp b1-4da6337, Radeon 680M, Vulkan, with async uploads disabled. The file is mmap'd. Experts and the n-gram table stay on CPU and fault in from disk. GTT stayed near 45 MB. Fit is off so the loader does not try to fill the GPU with the expert pool.
KV cache is f16. Quantized KV crashes this architecture.
Smoke context was 512. Prompt speed was about 0.5 to 0.7 tok/s. Generation was about 1.0 to 1.1 tok/s.
GGML_VK_DISABLE_ASYNC=1 llama-cli \
-m CYBER-FROST-3.8-Q4_K_M.gguf \
-ot "per_layer_token_embd=CPU,exps=CPU" \
-cmoe \
-ngl 99 \
-c 512 \
-lm mmap -fit off \
-ctk f16 -ctv f16 \
--jinja
UD-Q4_K_XL uses the same flags.
Draft head
CYBER-FROST-3.8-Q2_K_S.gguf is the only trunk with the head in the file. Run it with --spec-type draft-mtp. Do not pass -md.
GGML_VK_DISABLE_ASYNC=1 llama-cli \
-m CYBER-FROST-3.8-Q2_K_S.gguf \
-ot "per_layer_token_embd=CPU,exps=CPU" \
-cmoe \
-ngl 99 \
-c 512 \
-lm mmap -fit off \
-ctk f16 -ctv f16 \
--spec-type draft-mtp --spec-draft-n-max 2
This needs the qwen4exp MTP graph in the local llama.cpp build on this machine. The head reads the wide residual from before the final mixer, not a separate draft file.
mtp-CYBER-FROST-3.8-Q8_0.gguf still does not load next to a trunk. llama.cpp stops because output_hc_norm.weight is missing. The mixer weights are stored as blk.48.nextn.hc_head_norm.weight. The Q4_K_M draft was exported the same way. Do not pass either sidecar with -md.
Test log
Prompt: write a Python 3 function named add that takes two ints and returns their sum. Pass means the reply contains that function and a return.
| file | date | build | backend | result |
|---|---|---|---|---|
CYBER-FROST-3.8-Q4_K_M.gguf |
2026-09-28 | b1-4da6337 | Vulkan, mmap, experts and n-gram table on CPU | pass, def add with a return, 1.1 tok/s |
CYBER-FROST-3.8-UD-Q4_K_XL.gguf |
2026-09-28 | b1-4da6337 | Vulkan, mmap, experts and n-gram table on CPU | pass, same function, 1.0 tok/s |
CYBER-FROST-3.8-UD-IQ4_XS.gguf |
2026-09-28 | b1-4da6337 | Vulkan, mmap, experts and n-gram table on CPU | pass, same function, 0.9 tok/s |
CYBER-FROST-3.8-MXFP4_MOE.gguf |
2026-09-28 | b1-4da6337 | Vulkan, mmap, experts and n-gram table on CPU | pass, same function, 1.5 tok/s |
CYBER-FROST-3.8-Q3_K_M.gguf |
2026-09-28 | b1-4da6337 | Vulkan, mmap, experts and n-gram table on CPU | pass, same function, 1.4 tok/s |
CYBER-FROST-3.8-Q3_K_S.gguf |
2026-09-28 | b1-4da6337 | Vulkan, mmap, experts and n-gram table on CPU | pass, same function, 1.4 tok/s |
CYBER-FROST-3.8-Q2_K.gguf |
2026-09-28 | b1-4da6337 | Vulkan, mmap, experts and n-gram table on CPU | pass, same function, 1.9 tok/s |
CYBER-FROST-3.8-Q2_K_S.gguf |
2026-09-28 | b1-4da6337 | Vulkan, in-file MTP, --spec-type draft-mtp |
pass, def add with a return, 0.9 tok/s |
mtp-CYBER-FROST-3.8-Q8_0.gguf with Q4_K_M |
2026-09-28 | b1-4da6337 | Vulkan, then CPU | fail, draft missing output_hc_norm.weight |
License
Qwen Community License 1.0, same terms as the source checkpoint. See LICENSE.
- Downloads last month
- 2,138
2-bit
3-bit
4-bit
8-bit
Model tree for 0xSojalSec/CYBER-FROST-3.8-GGUF
Base model
Qwen/Qwen3.8-Flash-Next