PE-T2I 9B NVFP4 (W4A16)
A 4-bit weight-only NVFP4 quant of Qwen/Qwen-Image-2.1-PE-T2I (commit f3ed7985), the official prompt rewriter for Qwen-Image 2.1. It turns a short image request into a long detailed English prompt plus an aspect ratio, as JSON after a think block. Built with Qwen. Non-commercial use only, under the original's Qwen Research License. See NOTICE for what changed.
11.87 GB on disk vs 18.82 GB for the original bf16. Only the 96 language-model MLP projections are 4-bit. The vision tower, all attention and linear-attention layers, norms, embeddings and lm_head stay bf16, bit-identical to the original.
Results vs the bf16 original
Tested on one NVIDIA GB10 (DGX Spark), vLLM, same flags for both, 50 prompts at 1:1. The pass/fail bars were written down before the test ran. Two passed, two failed, all four are reported.
| Test | bf16 | NVFP4 | Bar | Result |
|---|---|---|---|---|
| Valid JSON + closed think block, greedy | 100% / 100% | 100% / 100% | within 2 pts | PASS |
| Valid JSON + closed think, sampled (official recipe, 3 seeds x 50) | 150/150 | 150/150 | (extra) | - |
| Blind image test: 12 prompts, both rewrites rendered with the same seed, judged side by side without knowing which was which | 4 wins | 5 wins (3 ties) | not worse | PASS |
| Teacher-forced token agreement with bf16 | 98.83% (bf16 vs itself) | 93.54% | 95% | FAIL |
| Decode speed, aggregate at 16 concurrent | 144.5 tok/s | 188.2 tok/s (1.30x) | 2x | FAIL |
| Decode speed, single stream | 11.8 tok/s | 17.8 tok/s (1.51x) | 1.3x | pass |
What the two fails mean: the rewrites drift from the bf16 wording early (most disagreements are near-ties, median log-prob gap 0.25 nats), but they come out the same length (median 430 vs 428 words) and rendered images were judged no worse. The speedup is modest because the vLLM build used here (0.27.2rc0 for sm121) reported no native FP4 support and ran these weights through its Marlin weight-only kernel, which dequantizes to bf16. GB10 does have FP4 tensor cores; reaching them needs an sm_121a-targeted stack and most likely an FP4-activation (W4A4) quant, not tested here. Both servers ran with --enforce-eager. Weights in vLLM memory: 10.33 GiB vs 16.8 GiB.
Serving with vLLM
Tested with a vLLM 0.27.2rc0 build for sm121 (GB10). vLLM detects the ModelOpt W4A16_NVFP4 checkpoint from config.json.
vllm serve kurtholes/PE-T2I-9B-NVFP4-W4A16 \
--language-model-only --max-model-len 20480 --kv-cache-dtype fp8
Prompt it like the original: system = system_prompt.txt (shipped here, unchanged), user = your request, the shipped chat template with thinking on. Sampling per the original card: temperature 1.0, top_p 0.95, top_k 20, up to 16256 new tokens. The answer after </think> is JSON:
{"rewritten_prompt": "<long detailed English prompt>", "wh_ratio": "16:9"}
Loading through plain transformers was not tested.
How it was made
NVIDIA ModelOpt 0.43.0, NVFP4 weight-only: block size 16, e4m3 block scales, one fp32 global scale per tensor, max algorithm, on exactly model.language_model.layers.{0..31}.mlp.{gate,up,down}_proj. The stock MLP config was not used, because its *mlp* pattern also catches the vision tower's MLPs. Calibration: 64 prompts in the real chat format (for weight-only max, the scales come from the weights). Checks after export: unquantized tensors bit-identical to the original (664 tensors), dequantized 4-bit weights at 9.4 to 9.5% relative error vs the originals, and the checkpoint routed by vLLM's own config code to exactly the 64 fused MLP linears.
Files
| File | sha256 (first 16) |
|---|---|
| model.safetensors | f683f3f7641f9c5d |
| config.json | 54c51a9237c5061e |
| hf_quant_config.json | 26b758d1a5d868f2 |
Full list in SHA256SUMS.
- Downloads last month
- 301
Model tree for kurtholes/PE-T2I-9B-NVFP4-W4A16
Base model
Qwen/Qwen-Image-2.1-PE-T2I