Instructions to use Blackfrost-AI/CYBER-FROST-3.8-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Blackfrost-AI/CYBER-FROST-3.8-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Blackfrost-AI/CYBER-FROST-3.8-FP8") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Blackfrost-AI/CYBER-FROST-3.8-FP8") model = AutoModelForMultimodalLM.from_pretrained("Blackfrost-AI/CYBER-FROST-3.8-FP8", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Blackfrost-AI/CYBER-FROST-3.8-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Blackfrost-AI/CYBER-FROST-3.8-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Blackfrost-AI/CYBER-FROST-3.8-FP8
- SGLang
How to use Blackfrost-AI/CYBER-FROST-3.8-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/CYBER-FROST-3.8-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/CYBER-FROST-3.8-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Blackfrost-AI/CYBER-FROST-3.8-FP8 with Docker Model Runner:
docker model run hf.co/Blackfrost-AI/CYBER-FROST-3.8-FP8
CYBER-FROST-3.8-FP8
A first-party Blackfrost-AI FP8 deployment artifact for security professionals conducting authorized research, assessment, engineering, and response work.
Cyber-Frost Harness
Cyber-Frost Harness is the public red/blue/purple runtime built around the Cyber-Frost family. It supplies eight procedural security skills, structured native-analysis tools, durable evidence handling, token-aware context management, and an isolated x86-64 vulnerable-image environment.
The published hard-12 harness result shown above used the sibling CYBER-FROST-3.8-NVFP4-V2 artifact, not this FP8 checkpoint. It increased verified solves from 1 under the generic scaffold to 5 under the harness, with five valid model misses and two infrastructure-invalid tasks. The defensible statements are 5/10 valid attempts and a 5/12 verified lower bound; no final 12-task percentage is claimed. Total model tokens fell from 51,939,855 to 22,043,450. These are scaffold-intervention results and do not transfer automatically through FP8 conversion.
- Source and quick start
- Architecture and trust boundaries
- Full evaluation disclosure
- Publication-safe result records
The model-facing container receives only the vulnerable task image and no network, fixed image, Docker control, grader database, or credentials. The harness changes the runtime and evidence path; it does not change model weights.
Release status and contents
Cyber-Frost is a public research release under active quality assessment.
This repository contains a standalone FP8 checkpoint derived from Blackfrost-AI/CYBER-FROST-3.8-BF16, plus its weight index, configuration, tokenizer and processor assets, packaged chat template, structural-verification receipt, and upstream license. It is not an adapter and does not require the BF16 parent checkpoint at load time.
| Field | Released artifact |
|---|---|
| Clean model name | CYBER-FROST-3.8-FP8 |
| Former artifact name | BLACKFROST-3.8-DERISKED-FP8 |
| Architecture | Qwen4ExpForConditionalGeneration |
| Precision | Mixed FP8/BF16; see the precision layout below |
| Weight layout | 131 SafeTensors shards |
| Physical weight payload | 185,523,321,634 bytes (172.78 GiB) |
| Indexed tensor payload | 185,502,232,570 bytes |
| Configured context | 262,144 tokens |
| Native speculative head | one MTP layer; its routed experts are FP8 and its remaining tensors are BF16 |
| Primary release task | text generation; artifact-specific inference qualification is pending |
The configuration includes a vision tower and processor files, but this FP8 release has not received a multimodal quality evaluation. Do not infer validated image or video capability from their presence.
Why Cyber-Frost exists
Security work is unusually vulnerable to false refusals. The same vocabulary appears in incident response, exploit validation, malware analysis, defensive engineering, and unauthorized activity; a general-purpose assistant can react to individual terms instead of the operator's legitimate scope.
Cyber-Frost is designed to reduce that unnecessary friction in professional, authorized workflows. It is intended to stay technically direct when an analyst is reviewing a finding, reproducing a vulnerability in a controlled environment, writing detection content, analyzing malicious code, or operating an approved security agent. This is a design objective inherited from the BF16 parent, not a measured behavioral claim for this FP8 variant.
Authorization is an external control. The model cannot establish ownership, consent, rules of engagement, jurisdiction, or whether a target is in scope. Deployers must enforce identity, scope, tool permissions, logging, rate limits, and human review outside the model.
Security corpus
The BF16 parent was fine-tuned on a Blackfrost-AI security corpus combining curated security material, operator-authored workflows, realistic engagement-style scenarios, and Blackfrost-owned distillation data. Corpus sizes, source-by-source counts, raw engagement material, client identities, prompts, and responses are intentionally not published.
Domain coverage includes:
- Reconnaissance and OSINT
- Social engineering, business-email compromise, and deepfake-enabled abuse
- Web application and API security
- Identity, authentication, and Active Directory security
- Network, perimeter, VPN, and protocol security
- Vulnerability research, bug bounty, and binary exploitation
- Malware analysis, ransomware, and endpoint defense
- Cloud, container, and Kubernetes security
- Software supply-chain security
- Mobile, IoT, wireless, and physical security
- Industrial-control-system and operational-technology security
- Cryptography and security protocols
- Privilege escalation, lateral movement, and data exfiltration
- Threat intelligence, APT analysis, and purple-team operations
- AI-agent, LLM, and adversarial-ML security
Blackfrost-AI attests that owned portions of the corpus were developed from sanitized experience with authorized security work. It also applies a frontier-scale policy to its distillation teachers, excluding teachers below the 753B-parameter class. The release evidence independently binds one security subset to a Qwen3.8 2.4T teacher; it does not include a corpus-wide teacher manifest. These are operator provenance statements, not independent benchmark findings.
Training-data provenance and licensing review for the mixed-source corpus remains in progress. FP8 conversion added no new fine-tuning data; this section describes the BF16 parent inherited by the quantized artifact.
Model specifications and precision layout
The text stack has 48 blocks with hybrid linear and full attention, using a full-attention block every fourth layer. Its hidden size is 2,560 with 24 attention heads and 2 KV heads. The MoE stack contains 512 routed experts, selects 10 experts per token, and includes a shared expert. One native MTP layer is packaged for speculative decoding.
This is a mixed-precision checkpoint, not an all-tensor FP8 conversion:
- Routed expert weights across the 48 trunk layers and the MTP expert layer use FP8 E4M3 with 128 by 128 block granularity.
- Activations use the runtime's dynamic FP8 path.
- Expert scales remain BF16.
- The PLE embedding uses 128 FP8 shards with one shared BF16 scale.
- The remaining 29 MTP tensors, all vision tensors, attention and linear-attention state, routers, shared experts, embeddings, normalization tensors, and other non-target tensors retain their source precision.
- KV-cache precision is selected by the serving runtime and is not encoded in the checkpoint weights.
The bundled configuration follows the tensor naming, block geometry, and dynamic-activation convention of Qwen/Qwen3.8-Flash-Next-FP8 at immutable revision 236dfdf285828023ca3bcd3f37366c58a3469b13.
The 262,144-token configuration ceiling is not a blanket quality guarantee. FP8-specific long-context, high-concurrency, multimodal, tool-use, and speculative-decoding qualification remains pending.
Lineage
- Foundational checkpoint:
Qwen/Qwen3.8-Flash-Nextat immutable revisionde4b8e4d43b917e7706784d8bb445c9af86a3540. - Blackfrost security adaptation: security-domain fine-tuning followed by a full BF16 merge. The merged internal stage was identified as
BLACKFROST-3.8-FLASH-BF16. - Behavioral stage: a Blackfrost-AI behaviorally modified derivative targeting lower false-refusal friction in authorized security workflows. The proprietary transformation process is not distributed.
- BF16 conversion source: the payload now published as
Blackfrost-AI/CYBER-FROST-3.8-BF16, at immutable source revision5321904427c4ef54df8a667edcbc2d1184e4286e. - FP8 layout reference:
Qwen/Qwen3.8-Flash-Next-FP8at revision236dfdf285828023ca3bcd3f37366c58a3469b13supplied the packaging convention, not the behavioral or weight lineage. - FP8 conversion: the target expert and PLE tensors were converted as described above. No additional training or behavioral modification was performed during conversion.
- Release identity: the resulting artifact was formerly labeled
BLACKFROST-3.8-DERISKED-FP8and is now namedCYBER-FROST-3.8-FP8. The rename is not another conversion or training run.
Tokenizer, processor, and license lineage comes through the pinned Qwen foundation and BF16 source.
Artifact verification
The completed artifact passed structural conversion validation. A subsequent header-level audit opened all 131 SafeTensors shards and reconciled every indexed tensor with its assigned shard: 152,089 index entries, zero missing or extra tensors, zero invalid offsets, and zero tensor-size mismatches.
The verified payload contains:
- 75,392 FP8 E4M3 tensors;
- 76,694 BF16 tensors;
- 3 I64 tensors;
- 75,264 routed-expert weight tensors and their 75,264 BF16 scale tensors;
- 128 FP8 PLE embedding tensors and one BF16 PLE scale;
- 333 preserved vision tensors; and
- 1,432 preserved non-quantized tensors.
The generated FP8 values and scales passed finite-value checks. Preserved tensor shapes and dtypes were checked against the BF16 source, and the completed weight map matched the official Qwen3.8 Flash-Next FP8 reference layout.
| Artifact | SHA-256 |
|---|---|
config.json |
6c41934d3fd7bda8a89f2af85292df2b94e3807927ca372a01bd5225f70f2953 |
model.safetensors.index.json |
1f163f7efc9a4b981e6e13978b7f39f9a3d4ddb11fff562b80500538021b4407 |
chat_template.jinja |
ba1946683f7615254fb246f0c0a652fd3aa02066ef8328ed4b8af08219749395 |
tokenizer_config.json |
3bc90552399707778ca94ef93fb1da7aa3c4dfa5ba94ca15f52b18c2a38ec904 |
ASSETS/BLACKFROST-AI-BANNER.png |
52a12b0caf63c1859da518738f78517a49a2616a8bc04daecc515f8793e30f09 |
| Qwen Community License file | a0dc422560841fd68e06d974907f8b4c709bca44a67daad2b528437bdf676c08 |
The sanitized ARTIFACT-VERIFICATION.json receipt is included with the release.
Evaluation status
The FP8 checkpoint has passed structural validation only. It has not yet received an artifact-specific load test, API smoke test, capability benchmark, refusal or over-refusal evaluation, coding benchmark, cyber benchmark, multimodal evaluation, tool-calling evaluation, long-context evaluation, or throughput benchmark.
Results measured on the BF16 parent, NVFP4 sibling, EXL3 derivative, or upstream Qwen artifacts must not be attributed to this FP8 checkpoint. Fresh measurements must be run against this exact payload before behavioral or performance results are reported.
Prompt, tool use, and sampling
The release includes the same chat_template.jinja and embedded tokenizer_config.json template as the published Cyber-Frost BF16 and NVFP4 variants. The two template representations are byte-identical.
The template supplies the Cyber-Frost operating prompt, appends caller-provided system context, supports image and video placeholders, exposes Qwen-style reasoning controls, and serializes XML-style tool calls and tool responses. Thinking is enabled by default; supported reasoning-effort values are xhigh, medium, and low. Generation defaults are temperature 1.0, top-p 0.95, and top-k 20.
The template's authorization assumption is not an access-control mechanism. An agent runtime must independently restrict credentials, targets, files, networks, commands, and approval-requiring actions. Tool-call output is proposed text until an external executor acts on it.
Changing the template, caller system message, reasoning mode, sampling, quantization backend, or runtime can materially change behavior. Record those settings when reporting results.
Deployment
This approximately 173 GiB weight artifact does not fit entirely in the 128 GiB memory envelope of a single DGX Spark. It is intended for a larger multi-GPU FP8-capable target and requires a runtime with explicit qwen4_exp and blockwise FP8 support.
This repository does not yet include a validated deployment kit. Runtime compatibility, kernel selection, memory use, throughput, MTP behavior, and output quality are implementation-dependent. The clean API model identifier is CYBER-FROST-3.8-FP8.
Intended use
Cyber-Frost is intended for qualified security professionals working within explicit authorization, including defensive research, secure code review, vulnerability validation, red-team and purple-team exercises, detection engineering, incident response, malware analysis, bug hunting, and controlled security-agent workflows.
It is not intended to authorize access, choose targets, define rules of engagement, make autonomous high-impact decisions, or replace legal, compliance, and safety review. Do not use it to access systems or data without permission, evade oversight, persist in third-party environments, deploy malware, steal credentials or data, disrupt services, or cause physical harm.
Limitations and security responsibility
- Generated findings, code, commands, indicators, and remediation advice may be wrong, incomplete, outdated, or fabricated. Independently review them and execute only in isolated, authorized environments.
- Reduced over-refusal is a design objective inherited from the BF16 parent, not a measured result for this FP8 artifact.
- BF16 and sibling-quant behavior, safety observations, benchmark results, and runtime characteristics do not automatically transfer through FP8 conversion.
- The model is not a policy engine, authorization service, sandbox, malware scanner, or secrets boundary.
- Current evidence does not establish an FP8 refusal rate, standardized cyber competence, production readiness, long-context quality, multimodal quality, or reliable MTP acceleration.
- Model behavior can shift substantially with prompts, sampling, runtime versions, quantization kernels, speculative settings, and agent scaffolding.
The deployer is responsible for authorization, least privilege, isolation, network policy, credential handling, human approval gates, monitoring, incident response, and compliance with applicable law.
License and disclaimer
Use and redistribution of this checkpoint are governed by the Qwen Community License 1.0. Review the license before use. This research release is provided without a warranty of correctness, fitness, security, or non-infringement.
Report reproducible model or packaging issues through this repository's Discussions page without including secrets, client data, live targets, or sensitive exploit details.
- Downloads last month
- 143

