Instructions to use caiovicentino1/Eikos-4B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use caiovicentino1/Eikos-4B-MLX-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download caiovicentino1/Eikos-4B-MLX-4bit --local-dir Eikos-4B-MLX-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Eikos-4B-MLX-4bit
MLX 4-bit build of Eikos-4B: MLX weights quantized to 4 bits (group size 64; mlx_lm convert -q --q-bits 4 --q-group-size 64) for Apple Silicon.
The conversion keeps the text model only (no vision tower, no MTP weights), and the embeddings and LM head
are quantized like the other layers. The prompt format, the
letter readout and the calibration (calib.json, T = 1) are the same as the bf16 model.
Eikos answers typed decisions (yes/no, one of N options, ordinal scores) about a given state in one forward pass, with a calibrated probability for every option. See the Eikos-4B card for what the model does, how it was trained, the full evaluation and its limitations.
Use
pip install mlx mlx-lm
hf download caiovicentino1/Eikos-4B-MLX-4bit --local-dir Eikos-4B-MLX-4bit
python Eikos-4B-MLX-4bit/mlx_decide.py Eikos-4B-MLX-4bit # runs a few demo decisions
import sys; sys.path.insert(0, "Eikos-4B-MLX-4bit")
from decision_core import options_of
from mlx_decide import MLXDecider
d = MLXDecider("Eikos-4B-MLX-4bit")
q = {"type": "noul", "instructions": "Under rule 4.2, can this order be executed as submitted?",
"criteria": {"true": "complies with rule 4.2", "false": "breaches rule 4.2"}}
probs, n_tokens = d.dist("Order ticket #A-2231 ... Rule 4.2: ... Approvals on file: none.", q, options_of(q))
mlx_decide.py uses the same letter readout, prompt and calibration as the GPU builds. It also answers several
questions over one state in a single cached batch (dist_many_cached).
Text only: this build has no vision weights, so it does not read images. Images need a GPU build with vLLM and serve v1.3: see Images.
Validation against bf16
Same 7,371 items for both builds (7 suites, never used in training), MLX run with its CUDA backend on Linux. The release gate was fixed before looking at results: accuracy within 1 point of bf16, ECE within 0.01, and at least 97% of answers unchanged.
| Eikos-4B (bf16) | Eikos-4B-MLX-4bit | |
|---|---|---|
| Size | 9.3 GB | 2.4 GB |
| JevBench public — original / hard | 91.7 / 73.9 | 95.8 / 73.0 |
| DecisionBench — medium / hard | 77.1 / 66.6 | 78.8 / 64.8 |
| General battery (9 tasks) | 75.7 | 75.3 |
| Finance (CUAD, sentiment, FinQA-judge) | 74.7 | 76.0 |
| Trade rules — seen / unseen | 74.5 / 76.0 | 72.4 / 76.0 |
| Compositional rules — same type / new domain / rulebooks | 96.0 / 91.5 / 91.7 | 95.0 / 90.5 / 90.7 |
| ECE (lower is better) | 0.033 | 0.021 |
| ≥0.90 confidence: decides / error | 34.7% / 2.5% | 32.5% / 2.7% |
| Same answer as bf16 (all / confident ≥0.9) | — | 92.5% / 99.8% |
It passes our release gate on accuracy and calibration, but not on answer agreement (92.5% of answers are the same as bf16; the gate asks for 97%). Most changed answers (93%) are on items where the bf16 model itself was unsure (confidence below 0.7); on decisions the bf16 model takes with confidence ≥0.9, 99.8% of answers are the same.
License
MIT for our contributions (LICENSE). The base model, Qwen3.5-4B, is Apache-2.0 (LICENSE-Qwen);
attributions are in NOTICE. Not legal, tax or investment advice.
- Downloads last month
- 66
4-bit
