Confidential Qwen3.8-27B: a recipe for fast, verifiable inference on a GPU TEE
This is a tested recipe for serving Qwen/Qwen3.8-27B-FP8 (revision
017b9c7a) inside a hardware TEE: an Intel TDX confidential VM with an NVIDIA H200 in Confidential Computing (CC) mode.
Clients encrypt prompts end to end to a key that the TEE generates and proves in hardware. Clients can check that proof,
and every answer's receipt, offline.
No weights are re-uploaded here. You fetch the upstream weights and check every file against src/ops/confidential/phala/qwen38-27b-fp8.manifest.json
(80 files, root 2fc31dd7…02218) before serving them.
Measured on 28 Sep 2026 on a Phala Cloud H200 (dstack-nvidia-0.5.9, driver 580.95.05, TCB UpToDate).
1. The vLLM settings that matter under GPU CC
| Setting | Why |
|---|---|
VLLM_USE_V2_MODEL_RUNNER=0 |
vLLM 0.29's default V2 runner reads inputs through UVA views of pinned host memory. Under CC the GPU can't read guest memory directly, so those views go stale and the model returns garbage (vllm#57224). We reproduced this in the CVM: a host write never reached the GPU view. |
no --enforce-eager |
The common workaround adds eager mode, but CUDA graphs are fine under CC. Eager holds every stream at 8–10 tok/s because kernel launches dominate. |
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' + --kv-cache-dtype fp8_e4m3 |
Qwen3.8's MTP heads nearly double single-stream speed. Its greedy output was byte-identical to the reference on our correctness set. |
compose/docker-compose.vllm-cc.yml is a minimal compose with these settings. lab/uvafix/ has an opt-in patch that
makes the V2 runner correct under CC by replacing its UVA buffers with explicit H2D copies (the same idea as upstream
PR #57415). It was as fast as V1 + graphs in our run.
2. Measurements
Correctness gate. Every engine config ran 8 fixed greedy prompts and was compared with the V1 + eager output. The
prompts are arithmetic, a list, EN→FR translation, a synthetic clinical note, a contract summary, JSON extraction, code
and counting. A config passes when no output is garbage and the mean similarity is at least 0.4. Only configs that pass
are benchmarked (lab/lab.py).
Engine lab (inside the CVM). The prompt is ~1,575 tokens with 512 output tokens and ignore_eos. Figures are
output tok/s.
| Config | Correctness | c=1 per stream | c=8 aggregate | c=32 aggregate | c=32 p50 latency |
|---|---|---|---|---|---|
| V1 + eager (common workaround) | reference | 9.8 | 71 | 248 | 65.3 s |
| V1 + CUDA graphs | 8/8 sane, 7/8 identical | 76.8 | 402 | 1,312 | 12.4 s |
| V1 + graphs + MTP-3 + fp8 KV | 8/8 identical | 140 | 628 | 1,722 | 8.9 s |
V2 + graphs + UVA fix (lab/uvafix) |
8/8 sane, 7/8 identical | 75.5 | 486 | 1,308 | 12.5 s |
| V2 + UVA fix + MTP-3 + fp8 KV | 8/8 sane, 7/8 identical | 163 | 844 | 1,939 | 8.1 s |
End to end through the sealed, attested path. This is natural generation with thinking on, 512-token answers and
a ~1,575-token prompt. Each request is HPKE-sealed on the client, opened inside the TEE, answered, sealed and signed by
the attested key, then opened and signature-checked on the client (scripts/confidential/soak.py, 10 min per level).
| Concurrency | Requests | Errors | p50 latency | p95 latency | Output tok/s | Prompt tok/s |
|---|---|---|---|---|---|---|
| 1 | 167 | 0 | 3.35 s | 4.47 s | 141 | 449 |
| 8 | 993 | 0 | 4.58 s | 6.25 s | 836 | 2,658 |
| 32 | 2,169 | 0 | 8.41 s | 11.54 s | 1,821 | 5,793 |
For comparison, the common workaround (V1 + eager) through the same path gave 53.4 s p50 at c=1 (9.2 tok/s).
3. What a client verifies
GET /v1/confidential/keyreturns a key document and an evidence bundle. The bundle holds:- the TDX quote, whose REPORTDATA is the key binding;
- Intel PCS collateral;
- the dstack event log, whose RTMR3 replay gives the compose hash;
- an NVIDIA NRAS GPU token, with nonce = SHA-256(binding).
python -m decosa_api.confidential.verify --bundle bundle.json --policy policy.json --key-doc key_doc.json --at collectedruns offline: the vendor roots are pinned, and the bundle carries the collateral.- A fresh challenge (
POST /v1/confidential/attest {nonce}) proves the TEE is live now. - The client HPKE-seals the prompt to
kem_pub. The answer is signed by the attested Ed25519 key. The receipt (decosa.confidential_receipt.v1) names the evidence hash, the policy hash, the weights and the ciphertext hashes.scripts/confidential/verify_calls.pyre-checks every receipt of a run offline.
The scripts run from the src/ directory with Python 3.11+ and the cryptography package (for example
cd src && python scripts/confidential/demo_client.py --help); they find decosa_api and pouw_inference next to them.
scripts/confidential/pin_policy.py pins a policy after first boot. Before pinning, it checks two things: the deployed
compose must equal your build byte for byte, and its hash must equal the RTMR3 event.
Only trust a pinned policy. A first-boot ("pilot") policy leaves the compose hash and MRTD/RTMR0-2 open, so it would accept any deployment on genuine hardware. Pin the policy to the measured deployment (compose hash, MRTD, RTMR0-2) before trusting it. Our newer client code (next release of this recipe) refuses unpinned policies by default; until then, treat any policy with open compose hashes or measurements as untrusted.
4. Honest limits
- A TEE proves which code and weights ran, not that the answer is right.
- Physical interposer attacks (TEE.fail, DDRop) are out of scope for Intel, AMD and NVIDIA.
- The cloud host still sees metadata (sizes, timing) and can stop the VM.
- The provider's in-TEE OS (here dstack) is measured and reproducible, but it is their code.
- Healthcare use needs a BAA with the host. A TEE doesn't remove that.
5. Contents and licences
src/: verifier (TDX, SEV-SNP, Azure HCL, dstack, NVIDIA NRAS), key binding, receipts, in-CVM sealbox and scripts. Apache-2.0 (src/LICENSE).lab/: the correctness gate and benchmark, and the UVA fix. Apache-2.0.compose/: a minimal vLLM-on-CC compose. Apache-2.0.- Model: Qwen/Qwen3.8-27B-FP8, Apache-2.0, fetched from the upstream repo. Engine: vLLM, Apache-2.0.
src/pouw_inference/: the sealed-request client (HPKE sealing, sealed requests, Ed25519 signing), a subset of the pouw-inference-token package. Apache-2.0.
Copyright 2026 Seth Pratt (Decosa). Licensed under the Apache License, Version 2.0: see LICENSE and NOTICE.
"Decosa" is a trademark; the licence covers the code, not the name.