Eikos-27B-FP8

Eikos overview: JevBench hard accuracy, error when at least 90% confident and long context, against Jev and Laya

FP8 build of Eikos-27B: weights in FP8 (per channel) with dynamic FP8 activations (per token), made with llm-compressor (FP8_DYNAMIC, compressed-tensors format). No calibration data. The vision tower, MTP weights, embeddings and LM head are kept in higher precision. The prompt format, the letter readout and the calibration (calib.json, T = 1) are the same as the bf16 model.

Eikos answers typed decisions (yes/no, one of N options, ordinal scores) about a given state in one forward pass, with a calibrated probability for every option. See the Eikos-27B card for what the model does, how it was trained, the full evaluation and its limitations.

Use

Requires vLLM ≥ 0.30.0. Older builds return wrong answers when several long requests are batched together on this hybrid (Gated DeltaNet) architecture.

hf download caiovicentino1/Eikos-27B-FP8 --local-dir Eikos-27B-FP8
bash Eikos-27B-FP8/serve_vllm.sh $PWD/Eikos-27B-FP8 8001      # vLLM engine: letter readout + hybrid prefix cache
python Eikos-27B-FP8/serve.py --model $PWD/Eikos-27B-FP8 --vllm-url http://127.0.0.1:8001 --port 8000   # HTTP API on :8000

The HTTP API, agent sessions and the question types are the same as for Eikos-27B.

Images (serve v1.3): this build reads images too; its vision tower is kept in bf16. Send them in "images", as multipart files or inside the state: see Images for the format, the limits and the full results. Zero-shot, on the same public items for both builds:

300 items per set Eikos-27B (bf16) Eikos-27B-FP8
MME (yes/no) 89.0% 90.0%
MMStar 70.7% 70.7%
SEED-Bench-2-Plus 74.0% 74.3%
ScreenSpot-v2, no marks 70.7% 71.3%
ScreenSpot-v2, 4×4 grid drawn 61.0% 60.0%
ScreenSpot-Pro, no marks 66.0% 65.7%
Same answer as bf16: all / confident (≥0.9) — 97.6% / 100.0%

Images are harder than our text suites: the bf16 model is at least 90% confident on only 25% of these items. Of the 44 answers (out of 1,800) that differ from bf16, all are on items where bf16 itself was below 0.7 confidence.

Validation against bf16

Same 7,371 items for both builds (7 suites, never used in training), vLLM 0.30 with batching and prefix cache on. The release gate was fixed before looking at results: accuracy within 1 point of bf16, ECE within 0.01, and at least 97% of answers unchanged.

Eikos-27B (bf16) Eikos-27B-FP8
Size 55.6 GB 31.2 GB
JevBench public — original / hard 100.0 / 82.0 100.0 / 83.8
DecisionBench — medium / hard 89.1 / 78.2 88.7 / 78.8
General battery (9 tasks) 82.6 82.8
Finance (CUAD, sentiment, FinQA-judge) 85.4 85.3
Trade rules — seen / unseen 85.4 / 87.4 85.6 / 86.9
Compositional rules — same type / new domain / rulebooks 95.6 / 94.3 / 95.3 95.6 / 94.7 / 95.0
ECE (lower is better) 0.043 0.042
≥0.90 confidence: decides / error 44.1% / 2.5% 44.2% / 2.6%
Same answer as bf16 (all / confident ≥0.9) — 98.8% / 100.0%

It passes our release gate.

Third-party benchmarks

We ran other groups' public benchmarks on this build, with their own harnesses and scorers, in September 2026. Official placement depends on each maintainer.

Decision Index 0.2.1: 38 benchmarks in five areas, each chance-corrected (0 = random guessing, 100 = perfect). We ran the full suite with the kit's own runner and scorer.

Index Knowledge & Reasoning Language Retrieval & Classification Tools & Automation Arts & Human Taste
Jev 57.89 51.3 62.0 55.4 75.1 37.7
Eikos-27B-FP8 55.46 39.9 63.3 55.9 74.4 39.8
  • All 151,476 requests were answered and none was refused. Requests went one at a time on one RTX PRO 6000, with a median of 124 ms per request.
  • On the 27 September board this would be 5th among open models; the best open model there is Surogate Rune 26B-A4B v3 at 57.44.
  • The gap to Jev is in knowledge and reasoning (GPQA Diamond, MMLU-Pro, BBH). Eikos is ahead of Jev in Language, Retrieval & Classification and Arts & Human Taste, and 0.7 behind in Tools & Automation.
  • Results and scripts: dataset. Submission: #19.

AgentRewardBench: judging whether a web agent completed its task. Test split, 1,106 trajectories, official scorer.

Judge Precision Recall F1
Eikos-27B-FP8 78.9 73.6 76.1
Jev (same inputs, through its API) 75.7 65.4 70.2
GPT-4o, Axtree (best published LLM judge) 69.8 83.1 75.9
  • Eikos read the same text as the published Axtree judges, and the four official questions went out as typed questions in one request.
  • Judgments and method: dataset. Leaderboard submission: #12.

Jevals suite 0.1.0: Decision Score (100 = perfect, 0 = guessing the base rates). 300 items × 5 repeats per task, using the published states byte for byte.

PubMedQA (yes/no) Banking77 (choice, 77 options) HelpSteer2 (score)
Gemini 3.8 Flash 73.0 74.1 4.6
Eikos-27B-FP8 71.4 66.8 8.3
Jev 69.0 67.8 9.2
  • Eikos is within Jev's confidence intervals on all three tasks.
  • On these tasks Eikos is underconfident: its mean confidence is below its accuracy, and the Brier-based score penalizes that.
  • Run records: gist. Listing request: #2.

JevBench v1.4.2: a setup check with the stock typesafe adapter on the 231 public items gave easy 48/48, standard 72/72 and hard 93/111. The official run, which adds held-out and sealed items, speed and cost, is requested: #117.

License

MIT for our contributions (LICENSE). The base model, Qwen3.8-27B, is Apache-2.0 (LICENSE-Qwen); attributions are in NOTICE. Not legal, tax or investment advice.

Downloads last month
221
Safetensors
Model size
28B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for caiovicentino1/Eikos-27B-FP8

Base model

Qwen/Qwen3.8-27B
Quantized
(4)
this model

Spaces using caiovicentino1/Eikos-27B-FP8 3

Collection including caiovicentino1/Eikos-27B-FP8