Eikos-4B-MLX-4bit

Eikos overview: JevBench hard accuracy, error when at least 90% confident and long context, against Jev and Laya

MLX 4-bit build of Eikos-4B: MLX weights quantized to 4 bits (group size 64; mlx_lm convert -q --q-bits 4 --q-group-size 64) for Apple Silicon. The conversion keeps the text model only (no vision tower, no MTP weights), and the embeddings and LM head are quantized like the other layers. The prompt format, the letter readout and the calibration (calib.json, T = 1) are the same as the bf16 model.

Eikos answers typed decisions (yes/no, one of N options, ordinal scores) about a given state in one forward pass, with a calibrated probability for every option. See the Eikos-4B card for what the model does, how it was trained, the full evaluation and its limitations.

Use

pip install mlx mlx-lm
hf download caiovicentino1/Eikos-4B-MLX-4bit --local-dir Eikos-4B-MLX-4bit
python Eikos-4B-MLX-4bit/mlx_decide.py Eikos-4B-MLX-4bit          # runs a few demo decisions
import sys; sys.path.insert(0, "Eikos-4B-MLX-4bit")
from decision_core import options_of
from mlx_decide import MLXDecider

d = MLXDecider("Eikos-4B-MLX-4bit")
q = {"type": "noul", "instructions": "Under rule 4.2, can this order be executed as submitted?",
     "criteria": {"true": "complies with rule 4.2", "false": "breaches rule 4.2"}}
probs, n_tokens = d.dist("Order ticket #A-2231 ... Rule 4.2: ... Approvals on file: none.", q, options_of(q))

mlx_decide.py uses the same letter readout, prompt and calibration as the GPU builds. It also answers several questions over one state in a single cached batch (dist_many_cached).

Text only: this build has no vision weights, so it does not read images. Images need a GPU build with vLLM and serve v1.3: see Images.

Validation against bf16

Same 7,371 items for both builds (7 suites, never used in training), MLX run with its CUDA backend on Linux. The release gate was fixed before looking at results: accuracy within 1 point of bf16, ECE within 0.01, and at least 97% of answers unchanged.

Eikos-4B (bf16) Eikos-4B-MLX-4bit
Size 9.3 GB 2.4 GB
JevBench public — original / hard 91.7 / 73.9 95.8 / 73.0
DecisionBench — medium / hard 77.1 / 66.6 78.8 / 64.8
General battery (9 tasks) 75.7 75.3
Finance (CUAD, sentiment, FinQA-judge) 74.7 76.0
Trade rules — seen / unseen 74.5 / 76.0 72.4 / 76.0
Compositional rules — same type / new domain / rulebooks 96.0 / 91.5 / 91.7 95.0 / 90.5 / 90.7
ECE (lower is better) 0.033 0.021
≥0.90 confidence: decides / error 34.7% / 2.5% 32.5% / 2.7%
Same answer as bf16 (all / confident ≥0.9) — 92.5% / 99.8%

It passes our release gate on accuracy and calibration, but not on answer agreement (92.5% of answers are the same as bf16; the gate asks for 97%). Most changed answers (93%) are on items where the bf16 model itself was unsure (confidence below 0.7); on decisions the bf16 model takes with confidence ≥0.9, 99.8% of answers are the same.

License

MIT for our contributions (LICENSE). The base model, Qwen3.5-4B, is Apache-2.0 (LICENSE-Qwen); attributions are in NOTICE. Not legal, tax or investment advice.

Downloads last month
66
Safetensors
Model size
4B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for caiovicentino1/Eikos-4B-MLX-4bit

Finetuned
Qwen/Qwen3.5-4B
Quantized
(4)
this model

Collection including caiovicentino1/Eikos-4B-MLX-4bit