PE-T2I 9B NVFP4 (W4A16)

A 4-bit weight-only NVFP4 quant of Qwen/Qwen-Image-2.1-PE-T2I (commit f3ed7985), the official prompt rewriter for Qwen-Image 2.1. It turns a short image request into a long detailed English prompt plus an aspect ratio, as JSON after a think block. Built with Qwen. Non-commercial use only, under the original's Qwen Research License. See NOTICE for what changed.

11.87 GB on disk vs 18.82 GB for the original bf16. Only the 96 language-model MLP projections are 4-bit. The vision tower, all attention and linear-attention layers, norms, embeddings and lm_head stay bf16, bit-identical to the original.

Results vs the bf16 original

Tested on one NVIDIA GB10 (DGX Spark), vLLM, same flags for both, 50 prompts at 1:1. The pass/fail bars were written down before the test ran. Two passed, two failed, all four are reported.

Test bf16 NVFP4 Bar Result
Valid JSON + closed think block, greedy 100% / 100% 100% / 100% within 2 pts PASS
Valid JSON + closed think, sampled (official recipe, 3 seeds x 50) 150/150 150/150 (extra) -
Blind image test: 12 prompts, both rewrites rendered with the same seed, judged side by side without knowing which was which 4 wins 5 wins (3 ties) not worse PASS
Teacher-forced token agreement with bf16 98.83% (bf16 vs itself) 93.54% 95% FAIL
Decode speed, aggregate at 16 concurrent 144.5 tok/s 188.2 tok/s (1.30x) 2x FAIL
Decode speed, single stream 11.8 tok/s 17.8 tok/s (1.51x) 1.3x pass

What the two fails mean: the rewrites drift from the bf16 wording early (most disagreements are near-ties, median log-prob gap 0.25 nats), but they come out the same length (median 430 vs 428 words) and rendered images were judged no worse. The speedup is modest because the vLLM build used here (0.27.2rc0 for sm121) reported no native FP4 support and ran these weights through its Marlin weight-only kernel, which dequantizes to bf16. GB10 does have FP4 tensor cores; reaching them needs an sm_121a-targeted stack and most likely an FP4-activation (W4A4) quant, not tested here. Both servers ran with --enforce-eager. Weights in vLLM memory: 10.33 GiB vs 16.8 GiB.

Serving with vLLM

Tested with a vLLM 0.27.2rc0 build for sm121 (GB10). vLLM detects the ModelOpt W4A16_NVFP4 checkpoint from config.json.

vllm serve kurtholes/PE-T2I-9B-NVFP4-W4A16 \
  --language-model-only --max-model-len 20480 --kv-cache-dtype fp8

Prompt it like the original: system = system_prompt.txt (shipped here, unchanged), user = your request, the shipped chat template with thinking on. Sampling per the original card: temperature 1.0, top_p 0.95, top_k 20, up to 16256 new tokens. The answer after </think> is JSON:

{"rewritten_prompt": "<long detailed English prompt>", "wh_ratio": "16:9"}

Loading through plain transformers was not tested.

How it was made

NVIDIA ModelOpt 0.43.0, NVFP4 weight-only: block size 16, e4m3 block scales, one fp32 global scale per tensor, max algorithm, on exactly model.language_model.layers.{0..31}.mlp.{gate,up,down}_proj. The stock MLP config was not used, because its *mlp* pattern also catches the vision tower's MLPs. Calibration: 64 prompts in the real chat format (for weight-only max, the scales come from the weights). Checks after export: unquantized tensors bit-identical to the original (664 tensors), dequantized 4-bit weights at 9.4 to 9.5% relative error vs the originals, and the checkpoint routed by vLLM's own config code to exactly the 64 fused MLP linears.

Files

File sha256 (first 16)
model.safetensors f683f3f7641f9c5d
config.json 54c51a9237c5061e
hf_quant_config.json 26b758d1a5d868f2

Full list in SHA256SUMS.

Downloads last month
301
Safetensors
Model size
7B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kurtholes/PE-T2I-9B-NVFP4-W4A16

Quantized
(6)
this model