Confidential Qwen3.8-27B: a recipe for fast, verifiable inference on a GPU TEE

This is a tested recipe for serving Qwen/Qwen3.8-27B-FP8 (revision 017b9c7a) inside a hardware TEE: an Intel TDX confidential VM with an NVIDIA H200 in Confidential Computing (CC) mode. Clients encrypt prompts end to end to a key that the TEE generates and proves in hardware. Clients can check that proof, and every answer's receipt, offline.

No weights are re-uploaded here. You fetch the upstream weights and check every file against src/ops/confidential/phala/qwen38-27b-fp8.manifest.json (80 files, root 2fc31dd7…02218) before serving them.

Measured on 28 Sep 2026 on a Phala Cloud H200 (dstack-nvidia-0.5.9, driver 580.95.05, TCB UpToDate).

1. The vLLM settings that matter under GPU CC

Setting Why
VLLM_USE_V2_MODEL_RUNNER=0 vLLM 0.29's default V2 runner reads inputs through UVA views of pinned host memory. Under CC the GPU can't read guest memory directly, so those views go stale and the model returns garbage (vllm#57224). We reproduced this in the CVM: a host write never reached the GPU view.
no --enforce-eager The common workaround adds eager mode, but CUDA graphs are fine under CC. Eager holds every stream at 8–10 tok/s because kernel launches dominate.
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' + --kv-cache-dtype fp8_e4m3 Qwen3.8's MTP heads nearly double single-stream speed. Its greedy output was byte-identical to the reference on our correctness set.

compose/docker-compose.vllm-cc.yml is a minimal compose with these settings. lab/uvafix/ has an opt-in patch that makes the V2 runner correct under CC by replacing its UVA buffers with explicit H2D copies (the same idea as upstream PR #57415). It was as fast as V1 + graphs in our run.

2. Measurements

Correctness gate. Every engine config ran 8 fixed greedy prompts and was compared with the V1 + eager output. The prompts are arithmetic, a list, EN→FR translation, a synthetic clinical note, a contract summary, JSON extraction, code and counting. A config passes when no output is garbage and the mean similarity is at least 0.4. Only configs that pass are benchmarked (lab/lab.py).

Engine lab (inside the CVM). The prompt is ~1,575 tokens with 512 output tokens and ignore_eos. Figures are output tok/s.

Config Correctness c=1 per stream c=8 aggregate c=32 aggregate c=32 p50 latency
V1 + eager (common workaround) reference 9.8 71 248 65.3 s
V1 + CUDA graphs 8/8 sane, 7/8 identical 76.8 402 1,312 12.4 s
V1 + graphs + MTP-3 + fp8 KV 8/8 identical 140 628 1,722 8.9 s
V2 + graphs + UVA fix (lab/uvafix) 8/8 sane, 7/8 identical 75.5 486 1,308 12.5 s
V2 + UVA fix + MTP-3 + fp8 KV 8/8 sane, 7/8 identical 163 844 1,939 8.1 s

End to end through the sealed, attested path. This is natural generation with thinking on, 512-token answers and a ~1,575-token prompt. Each request is HPKE-sealed on the client, opened inside the TEE, answered, sealed and signed by the attested key, then opened and signature-checked on the client (scripts/confidential/soak.py, 10 min per level).

Concurrency Requests Errors p50 latency p95 latency Output tok/s Prompt tok/s
1 167 0 3.35 s 4.47 s 141 449
8 993 0 4.58 s 6.25 s 836 2,658
32 2,169 0 8.41 s 11.54 s 1,821 5,793

For comparison, the common workaround (V1 + eager) through the same path gave 53.4 s p50 at c=1 (9.2 tok/s).

3. What a client verifies

  1. GET /v1/confidential/key returns a key document and an evidence bundle. The bundle holds:
    • the TDX quote, whose REPORTDATA is the key binding;
    • Intel PCS collateral;
    • the dstack event log, whose RTMR3 replay gives the compose hash;
    • an NVIDIA NRAS GPU token, with nonce = SHA-256(binding).
  2. python -m decosa_api.confidential.verify --bundle bundle.json --policy policy.json --key-doc key_doc.json --at collected runs offline: the vendor roots are pinned, and the bundle carries the collateral.
  3. A fresh challenge (POST /v1/confidential/attest {nonce}) proves the TEE is live now.
  4. The client HPKE-seals the prompt to kem_pub. The answer is signed by the attested Ed25519 key. The receipt (decosa.confidential_receipt.v1) names the evidence hash, the policy hash, the weights and the ciphertext hashes. scripts/confidential/verify_calls.py re-checks every receipt of a run offline.

The scripts run from the src/ directory with Python 3.11+ and the cryptography package (for example cd src && python scripts/confidential/demo_client.py --help); they find decosa_api and pouw_inference next to them.

scripts/confidential/pin_policy.py pins a policy after first boot. Before pinning, it checks two things: the deployed compose must equal your build byte for byte, and its hash must equal the RTMR3 event.

Only trust a pinned policy. A first-boot ("pilot") policy leaves the compose hash and MRTD/RTMR0-2 open, so it would accept any deployment on genuine hardware. Pin the policy to the measured deployment (compose hash, MRTD, RTMR0-2) before trusting it. Our newer client code (next release of this recipe) refuses unpinned policies by default; until then, treat any policy with open compose hashes or measurements as untrusted.

4. Honest limits

  • A TEE proves which code and weights ran, not that the answer is right.
  • Physical interposer attacks (TEE.fail, DDRop) are out of scope for Intel, AMD and NVIDIA.
  • The cloud host still sees metadata (sizes, timing) and can stop the VM.
  • The provider's in-TEE OS (here dstack) is measured and reproducible, but it is their code.
  • Healthcare use needs a BAA with the host. A TEE doesn't remove that.

5. Contents and licences

  • src/: verifier (TDX, SEV-SNP, Azure HCL, dstack, NVIDIA NRAS), key binding, receipts, in-CVM sealbox and scripts. Apache-2.0 (src/LICENSE).
  • lab/: the correctness gate and benchmark, and the UVA fix. Apache-2.0.
  • compose/: a minimal vLLM-on-CC compose. Apache-2.0.
  • Model: Qwen/Qwen3.8-27B-FP8, Apache-2.0, fetched from the upstream repo. Engine: vLLM, Apache-2.0.
  • src/pouw_inference/: the sealed-request client (HPKE sealing, sealed requests, Ed25519 signing), a subset of the pouw-inference-token package. Apache-2.0.

Copyright 2026 Seth Pratt (Decosa). Licensed under the Apache License, Version 2.0: see LICENSE and NOTICE. "Decosa" is a trademark; the licence covers the code, not the name.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support