H2O-Lightning-4B

H2O-Lightning-4B is a 4B-parameter decision model from H2O.ai, built on Qwen/Qwen3.5-4B. It answers typed decision questions about a record (a document, a ticket, a policy, a conversation) and, from v1.2, about images that come with the record (photos, screenshots, scanned documents, charts). It returns a probability for every option:

  • choice: picks one of a set of named options;
  • yes/no (noul): gives the probability that a statement is true;
  • score: gives an ordinal level on a stated scale, with its probability distribution.

It runs on unmodified vLLM 0.30.0 with a small standard-library shim in front (h2o_lightning_shim.py). Each decision is one forward pass and one output token, so the cost is input tokens only.

Measured

Measured on one NVIDIA RTX PRO 4500 Blackwell (32 GB) with serve.sh as shipped (vLLM 0.30.0; text temperature 0.8, image temperature 0.65, images up to 1,048,576 pixels).

JevBench public items result
JevBench (text) the 231 public items, JevBench's own CLI 205 / 231 (88.7 %): easy 48/48 · standard 71/72 · hard 86/111
ImageJevBench the 8 published examples (the only public items released with images) 8 / 8
image validation sets 1,877 questions over 16 public licence-clean sets (table below) 87.4 %

Details follow.

Text: JevBench's public set

From JevBench's own CLI, with --adapter typesafe, on the 231 public items (datasets/public/easy.jsonl, original.jsonl and hard.jsonl; dataset hash dc3995d8…) at JevBench commit bb05a33, with the commands below.

RTX PRO 4500 Blackwell 32 GB
correct 205 / 231 (88.7 %)
by tier: easy / standard / hard 48/48 · 71/72 · 86/111
answered and valid 231 / 231
yes/no answers with P(yes) strictly between 0.2 and 0.8 0 of 74
mean input tokens per decision 654
latency p50 / p95, standard tier, serial 0.033 s / 0.035 s
latency p50 / p95, all 231 items, serial 0.034 s / 0.221 s

Latency depends on the GPU.

Identity check for an evaluator: a correct setup reproduces 205/231 with the tier split above. Five hard items sit on near-ties (top two within 0.03) that GPU numerics can flip. Four are wrong here: hard-opus-a-temporal_numeric-07, hard-opus-a-temporal_numeric-12, hard-opus-c-temporal_numeric-04 and hard-sol-b-temporal_numeric-03. One is right: hard-sol-c-judge_hard-13.

Images

(a) The published ImageJevBench examples: 8 / 8. ImageJevBench's own harness is not public, and only 8 of its items are published with their images (on the benchmark's page: two everyday photos, four screenshots with five labelled click markers, one geometry diagram, one financial table). We read all 8 through this API: 8 correct. These 8 items are not a benchmark score.

(b) Public validation sets with human or source ground truth. Multiple-choice questions we built from each dataset's own annotations (human labels, boxes, receipt fields, table cells, a chart's own data); none come from a model. None of these images or questions was used for training. Read through this API, one image per request, as a data URI in state.

set (licence) what is asked n correct ECE
Open Images V7 validation (CC BY 4.0 annotations) is there a ‹class› in the photo (human-verified labels) 150 88.0 % 0.034
Open Images V7 validation how many ‹class› in the photo (human boxes, 2-8) 150 63.3 % 0.121
VizWiz validation (CC BY 4.0) can the question be answered from this photo 120 84.2 % 0.057
VizWiz validation the answer (6+ of 10 annotators) 120 100.0 % 0.015
TextVQA validation (CC BY 4.0) text in the photo (6+ of 10 annotators) 120 99.2 % 0.024
CORD v2 test + validation (CC BY 4.0) receipt total 95 99.0 % 0.014
CORD v2 test + validation one item's price 29 96.5 % 0.043
CORD v2 test + validation number of item lines 113 85.0 % 0.050
TAT-QA dev (CC BY 4.0) financial-report table questions (the table shown as an image) 120 95.8 % 0.037
DocLayNet test (CDLA-Permissive-1.0) which section heading is on the page 120 100.0 % 0.011
DocLayNet test how many tables are on the page (0-3) 120 75.8 % 0.057
DocLayNet test does the page contain a picture 120 90.0 % 0.039
Our World in Data charts (CC BY 4.0) which country is highest or lowest (the chart's own data) 100 98.0 % 0.058
Our World in Data charts did a value rise between two years 100 95.0 % 0.069
Our World in Data charts roughly what is a value 100 98.0 % 0.066
Rico app screens (CC BY 4.0) which of 5 marked elements reaches a goal (human widget captions) 200 65.0 % 0.094
all 1,877 87.4 % 0.017

ECE: 10-bin expected calibration error of the top answer's probability.

Run it

Hardware measured: one RTX PRO 4500 Blackwell (32 GB); the weights are 9.1 GB. Install vLLM 0.30.0 in a fresh virtual environment (for example python3 -m venv .venv && . .venv/bin/activate):

pip install "vllm==0.30.0+cu129" "torchcodec==0.16.0+cu129" --extra-index-url https://wheels.vllm.ai/0.30.0/cu129 --extra-index-url https://download.pytorch.org/whl/cu129

The shim needs Python 3.8 or later and nothing else.

hf download h2oai/h2o-lightning-4b --revision v1.2.1 --local-dir h2o-lightning-4b

1. Serve the model. Keep config.json as shipped: its "head_dtype": "float32" makes vLLM compute the logits in fp32. Inputs over the context limit get HTTP 422. The last two options set the image limits (up to 4 images per request, each scaled to at most 1.6 megapixels inside vLLM).

vllm serve ./h2o-lightning-4b --served-model-name h2oai/h2o-lightning-4b --host 127.0.0.1 --port 8000 \
  --max-model-len 40960 --gpu-memory-utilization 0.90 \
  --limit-mm-per-prompt '{"image": 4, "video": 0}' --mm-processor-kwargs '{"max_pixels": 1605632}'

If vLLM stops at startup because FlashInfer cannot compile its sampling kernels (an older system CUDA toolkit), start it with VLLM_USE_FLASHINFER_SAMPLER=0 in the environment: the decisions do not use the sampler.

2. Start the shim on the same machine:

python3 h2o-lightning-4b/h2o_lightning_shim.py --config h2o-lightning-4b/serve_config.json --vllm http://127.0.0.1:8000 --port 8741

bash h2o-lightning-4b/serve.sh runs steps 1 and 2 together. Before answering, the shim checks that vLLM serves h2oai/h2o-lightning-4b and that every label is one token at the answer slot; if not, it answers 503, so a misconfigured server stops a run instead of scoring it. Before the first image request it also checks that vLLM's chat template renders the same prompt as the text path.

3. Warm up until this returns HTTP 200:

curl -s http://127.0.0.1:8741/v1/systemone -H 'Content-Type: application/json' -d '{"state": "warm-up",
  "questions": {"decision": {"type": "noul", "instructions": "Is this a warm-up?",
  "criteria": {"true": "yes", "false": "no"}}}}'

4. Run JevBench's harness:

cd <JEVBENCH> && python3 -m jevbench.cli run \
  --tasks datasets/public/easy.jsonl,datasets/public/original.jsonl,datasets/public/hard.jsonl \
  --adapter typesafe --endpoint http://127.0.0.1:8741 --key-env '' --model h2oai/h2o-lightning-4b \
  --cost-basis self_hosted_gpu --reserve-usd 0 --results <OUT>/results.jsonl --raw-dir <OUT>/raw \
  --ledger <OUT>/ledger.jsonl --manifest <OUT>/manifest.json --run-label h2o-lightning-4b --delay-s 0

The API. Send requests to POST /v1/systemone. One request can carry several questions about the same state:

{"state": "Ticket 4471: since 09:10 the checkout page returns HTTP 500 for every customer; no orders have gone through in 40 minutes. Severity guide: P1 = a revenue path fully down; P2 = degraded but working; P3 = cosmetic.",
 "questions": {
   "priority":  {"type": "choice", "instructions": "Which priority does the severity guide assign?",
                 "criteria": {"p1": "Revenue path fully down", "p2": "Degraded but working", "p3": "Cosmetic"}},
   "all_users": {"type": "noul", "instructions": "The problem affects every customer.",
                 "criteria": {"true": "Affects everyone", "false": "Affects only some"}},
   "urgency":   {"type": "score", "instructions": "How urgent is this ticket?", "criteria": ["Low", "Medium", "High"]}}}

The response (rounded here):

{"model": "h2oai/h2o-lightning-4b",
 "answers": {"priority": {"type": "choice", "choice": "p1",
                          "probabilities": {"p1": 0.944, "p2": 0.040, "p3": 0.017}, "confidence": 0.916},
             "all_users": {"type": "noul", "noul": 0.929},
             "urgency": {"type": "score", "score": 1.789, "legend": {"0": "Low", "1": "Medium", "2": "High"},
                         "probabilities": {"0": 0.033, "1": 0.146, "2": 0.821}, "confidence": 0.732}},
 "probability_source": "native", "effort_used": {"tier": "fast", "passes_per_decision": 1},
 "usage": {"input_tokens": 231, "output_tokens": 0, "latency_ms": 70.3}}

usage.input_tokens counts the prompt head that the questions share once.

Images. Put each image in the request as a data:image/...;base64,... URI. It can go anywhere inside state (as a string, an object value or a list item), or in a top-level images list next to state and questions. The images list also takes bare base64 strings and http(s) URLs (vLLM fetches a URL itself, so the server needs network access for those; a URL it cannot fetch gets HTTP 422). Chat-style parts ({"type": "image_url", "image_url": {"url": ...}}) and multipart/form-data (the JSON in a request field, each image as a file part) work too.

  • Each image taken out of state leaves [image N] in its place, so the record can refer to it. Images in the top-level list are numbered first.
  • Up to 4 images per request and 20 MB per image.
  • Each image is scaled to at most 1,048,576 pixels, and image answers use their own temperature, 0.65 (serve_config.json, "image"); text answers keep 0.8.
  • A request without an image takes exactly the text path above.

This example uses example_receipt.png from this repository:

import base64, json, urllib.request
img = base64.b64encode(open("h2o-lightning-4b/example_receipt.png", "rb").read()).decode()
req = {"state": {"receipt": "data:image/png;base64," + img,
                 "claim": {"employee": "J. Ortiz", "category": "office supplies", "amount": "444.26"}},
       "questions": {
         "matches":  {"type": "noul", "instructions": "The receipt total equals claim.amount.",
                      "criteria": {"true": "Same amount", "false": "Different amount"}},
         "items":    {"type": "choice", "instructions": "How many item lines are on the receipt?",
                      "criteria": {"three": "3", "four": "4", "five": "5", "six": "6"}},
         "legible":  {"type": "score", "instructions": "How legible is the receipt?",
                      "criteria": ["Unreadable", "Partly readable", "Fully readable"]}}}
r = urllib.request.urlopen(urllib.request.Request("http://127.0.0.1:8741/v1/systemone", json.dumps(req).encode(),
                                                  {"Content-Type": "application/json"}))
print(json.dumps(json.load(r), indent=1))

The response (rounded here):

{"model": "h2oai/h2o-lightning-4b",
 "answers": {"matches": {"type": "noul", "noul": 0.964},
             "items": {"type": "choice", "choice": "six",
                       "probabilities": {"three": 0.006, "four": 0.012, "five": 0.024, "six": 0.958}, "confidence": 0.943},
             "legible": {"type": "score", "score": 1.926,
                         "legend": {"0": "Unreadable", "1": "Partly readable", "2": "Fully readable"},
                         "probabilities": {"0": 0.011, "1": 0.052, "2": 0.937}, "confidence": 0.906}},
 "probability_source": "native", "effort_used": {"tier": "fast", "passes_per_decision": 1},
 "usage": {"input_tokens": 585, "output_tokens": 0, "latency_ms": 141.0, "images": 1}}

usage.input_tokens includes the image tokens; usage.images counts the images read.

Status codes. An input the shim cannot take gets HTTP 422, which the runner scores as one wrong answer. That covers:

  • over the context limit;
  • more than 255 options;
  • an unknown question type or a malformed question;
  • more than 4 images, or an image that is not valid base64 image data.

An empty request gets 400. vLLM's 401, 403 and 429 pass through. Any other vLLM failure is a 502, which counts toward the runner's three-consecutive-failures stop.

Tests (no GPU): python3 h2o-lightning-4b/test_shim.py.

Limitations

  • Evaluated mostly in English. Decides best with up to about 16 options; up to 255 are accepted.
  • The probabilities are calibrated for decision questions of this kind, not for open-ended text.
  • Not a chat model: it answers one decision per question and does not generate explanations.
  • Images: no video input; counting many small objects and reading values off line charts are the weakest image skills we measured.

Disclosures

  • No JevBench or ImageJevBench data, public or otherwise, was used for training.
  • No rules keyed to any benchmark item's wording, id or answer; no network calls at serving time (except fetching an image URL a request supplies).
  • Price basis for an evaluator: Qwen3.5-4B at bf16, one output token per decision; 654 mean input tokens on the public text set. Image tokens are counted in usage.input_tokens.

License

Apache-2.0. Built on Qwen/Qwen3.5-4B by the Qwen team (Alibaba Cloud), Apache-2.0; served with unmodified vLLM (Apache-2.0).

Downloads last month
98
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for h2oai/h2o-lightning-4b

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(863)
this model
Quantizations
1 model