Instructions to use h2oai/h2o-lightning-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use h2oai/h2o-lightning-4b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="h2oai/h2o-lightning-4b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("h2oai/h2o-lightning-4b") model = AutoModelForMultimodalLM.from_pretrained("h2oai/h2o-lightning-4b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use h2oai/h2o-lightning-4b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "h2oai/h2o-lightning-4b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "h2oai/h2o-lightning-4b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/h2oai/h2o-lightning-4b
- SGLang
How to use h2oai/h2o-lightning-4b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "h2oai/h2o-lightning-4b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "h2oai/h2o-lightning-4b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "h2oai/h2o-lightning-4b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "h2oai/h2o-lightning-4b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use h2oai/h2o-lightning-4b with Docker Model Runner:
docker model run hf.co/h2oai/h2o-lightning-4b
H2O-Lightning-4B
H2O-Lightning-4B is a 4B-parameter decision model from H2O.ai, built on Qwen/Qwen3.5-4B. It answers typed decision questions about a record (a document, a ticket, a policy, a conversation) and, from v1.2, about images that come with the record (photos, screenshots, scanned documents, charts). It returns a probability for every option:
- choice: picks one of a set of named options;
- yes/no (
noul): gives the probability that a statement is true; - score: gives an ordinal level on a stated scale, with its probability distribution.
It runs on unmodified vLLM 0.30.0 with a small standard-library shim in front (h2o_lightning_shim.py). Each
decision is one forward pass and one output token, so the cost is input tokens only.
Measured
Measured on one NVIDIA RTX PRO 4500 Blackwell (32 GB) with serve.sh as shipped (vLLM 0.30.0; text temperature
0.8, image temperature 0.65, images up to 1,048,576 pixels).
| JevBench | public items | result |
|---|---|---|
| JevBench (text) | the 231 public items, JevBench's own CLI | 205 / 231 (88.7 %): easy 48/48 · standard 71/72 · hard 86/111 |
| ImageJevBench | the 8 published examples (the only public items released with images) | 8 / 8 |
| image validation sets | 1,877 questions over 16 public licence-clean sets (table below) | 87.4 % |
Details follow.
Text: JevBench's public set
From JevBench's own CLI, with --adapter typesafe, on the 231 public items (datasets/public/easy.jsonl,
original.jsonl and hard.jsonl; dataset hash dc3995d8…) at JevBench commit bb05a33, with the commands below.
| RTX PRO 4500 Blackwell 32 GB | |
|---|---|
| correct | 205 / 231 (88.7 %) |
| by tier: easy / standard / hard | 48/48 · 71/72 · 86/111 |
| answered and valid | 231 / 231 |
| yes/no answers with P(yes) strictly between 0.2 and 0.8 | 0 of 74 |
| mean input tokens per decision | 654 |
| latency p50 / p95, standard tier, serial | 0.033 s / 0.035 s |
| latency p50 / p95, all 231 items, serial | 0.034 s / 0.221 s |
Latency depends on the GPU.
Identity check for an evaluator: a correct setup reproduces 205/231 with the tier split above. Five hard items
sit on near-ties (top two within 0.03) that GPU numerics can flip. Four are wrong here:
hard-opus-a-temporal_numeric-07, hard-opus-a-temporal_numeric-12, hard-opus-c-temporal_numeric-04 and
hard-sol-b-temporal_numeric-03. One is right: hard-sol-c-judge_hard-13.
Images
(a) The published ImageJevBench examples: 8 / 8. ImageJevBench's own harness is not public, and only 8 of its items are published with their images (on the benchmark's page: two everyday photos, four screenshots with five labelled click markers, one geometry diagram, one financial table). We read all 8 through this API: 8 correct. These 8 items are not a benchmark score.
(b) Public validation sets with human or source ground truth. Multiple-choice questions we built from each
dataset's own annotations (human labels, boxes, receipt fields, table cells, a chart's own data); none come from a
model. None of these images or questions was used for training. Read through this API, one image per request, as a
data URI in state.
| set (licence) | what is asked | n | correct | ECE |
|---|---|---|---|---|
| Open Images V7 validation (CC BY 4.0 annotations) | is there a ‹class› in the photo (human-verified labels) | 150 | 88.0 % | 0.034 |
| Open Images V7 validation | how many ‹class› in the photo (human boxes, 2-8) | 150 | 63.3 % | 0.121 |
| VizWiz validation (CC BY 4.0) | can the question be answered from this photo | 120 | 84.2 % | 0.057 |
| VizWiz validation | the answer (6+ of 10 annotators) | 120 | 100.0 % | 0.015 |
| TextVQA validation (CC BY 4.0) | text in the photo (6+ of 10 annotators) | 120 | 99.2 % | 0.024 |
| CORD v2 test + validation (CC BY 4.0) | receipt total | 95 | 99.0 % | 0.014 |
| CORD v2 test + validation | one item's price | 29 | 96.5 % | 0.043 |
| CORD v2 test + validation | number of item lines | 113 | 85.0 % | 0.050 |
| TAT-QA dev (CC BY 4.0) | financial-report table questions (the table shown as an image) | 120 | 95.8 % | 0.037 |
| DocLayNet test (CDLA-Permissive-1.0) | which section heading is on the page | 120 | 100.0 % | 0.011 |
| DocLayNet test | how many tables are on the page (0-3) | 120 | 75.8 % | 0.057 |
| DocLayNet test | does the page contain a picture | 120 | 90.0 % | 0.039 |
| Our World in Data charts (CC BY 4.0) | which country is highest or lowest (the chart's own data) | 100 | 98.0 % | 0.058 |
| Our World in Data charts | did a value rise between two years | 100 | 95.0 % | 0.069 |
| Our World in Data charts | roughly what is a value | 100 | 98.0 % | 0.066 |
| Rico app screens (CC BY 4.0) | which of 5 marked elements reaches a goal (human widget captions) | 200 | 65.0 % | 0.094 |
| all | 1,877 | 87.4 % | 0.017 |
ECE: 10-bin expected calibration error of the top answer's probability.
Run it
Hardware measured: one RTX PRO 4500 Blackwell (32 GB); the weights are 9.1 GB. Install vLLM 0.30.0 in a fresh virtual environment
(for example python3 -m venv .venv && . .venv/bin/activate):
pip install "vllm==0.30.0+cu129" "torchcodec==0.16.0+cu129" --extra-index-url https://wheels.vllm.ai/0.30.0/cu129 --extra-index-url https://download.pytorch.org/whl/cu129
The shim needs Python 3.8 or later and nothing else.
hf download h2oai/h2o-lightning-4b --revision v1.2.1 --local-dir h2o-lightning-4b
1. Serve the model. Keep config.json as shipped: its "head_dtype": "float32" makes vLLM compute the logits in
fp32. Inputs over the context limit get HTTP 422. The last two options set the image limits (up to 4 images
per request, each scaled to at most 1.6 megapixels inside vLLM).
vllm serve ./h2o-lightning-4b --served-model-name h2oai/h2o-lightning-4b --host 127.0.0.1 --port 8000 \
--max-model-len 40960 --gpu-memory-utilization 0.90 \
--limit-mm-per-prompt '{"image": 4, "video": 0}' --mm-processor-kwargs '{"max_pixels": 1605632}'
If vLLM stops at startup because FlashInfer cannot compile its sampling kernels (an older system CUDA toolkit), start it
with VLLM_USE_FLASHINFER_SAMPLER=0 in the environment: the decisions do not use the sampler.
2. Start the shim on the same machine:
python3 h2o-lightning-4b/h2o_lightning_shim.py --config h2o-lightning-4b/serve_config.json --vllm http://127.0.0.1:8000 --port 8741
bash h2o-lightning-4b/serve.sh runs steps 1 and 2 together. Before answering, the
shim checks that vLLM serves h2oai/h2o-lightning-4b and that every label is one token at the answer slot; if not, it
answers 503, so a misconfigured server stops a run instead of scoring it. Before the first image request it also checks
that vLLM's chat template renders the same prompt as the text path.
3. Warm up until this returns HTTP 200:
curl -s http://127.0.0.1:8741/v1/systemone -H 'Content-Type: application/json' -d '{"state": "warm-up",
"questions": {"decision": {"type": "noul", "instructions": "Is this a warm-up?",
"criteria": {"true": "yes", "false": "no"}}}}'
4. Run JevBench's harness:
cd <JEVBENCH> && python3 -m jevbench.cli run \
--tasks datasets/public/easy.jsonl,datasets/public/original.jsonl,datasets/public/hard.jsonl \
--adapter typesafe --endpoint http://127.0.0.1:8741 --key-env '' --model h2oai/h2o-lightning-4b \
--cost-basis self_hosted_gpu --reserve-usd 0 --results <OUT>/results.jsonl --raw-dir <OUT>/raw \
--ledger <OUT>/ledger.jsonl --manifest <OUT>/manifest.json --run-label h2o-lightning-4b --delay-s 0
The API. Send requests to POST /v1/systemone. One request can carry several questions about the same state:
{"state": "Ticket 4471: since 09:10 the checkout page returns HTTP 500 for every customer; no orders have gone through in 40 minutes. Severity guide: P1 = a revenue path fully down; P2 = degraded but working; P3 = cosmetic.",
"questions": {
"priority": {"type": "choice", "instructions": "Which priority does the severity guide assign?",
"criteria": {"p1": "Revenue path fully down", "p2": "Degraded but working", "p3": "Cosmetic"}},
"all_users": {"type": "noul", "instructions": "The problem affects every customer.",
"criteria": {"true": "Affects everyone", "false": "Affects only some"}},
"urgency": {"type": "score", "instructions": "How urgent is this ticket?", "criteria": ["Low", "Medium", "High"]}}}
The response (rounded here):
{"model": "h2oai/h2o-lightning-4b",
"answers": {"priority": {"type": "choice", "choice": "p1",
"probabilities": {"p1": 0.944, "p2": 0.040, "p3": 0.017}, "confidence": 0.916},
"all_users": {"type": "noul", "noul": 0.929},
"urgency": {"type": "score", "score": 1.789, "legend": {"0": "Low", "1": "Medium", "2": "High"},
"probabilities": {"0": 0.033, "1": 0.146, "2": 0.821}, "confidence": 0.732}},
"probability_source": "native", "effort_used": {"tier": "fast", "passes_per_decision": 1},
"usage": {"input_tokens": 231, "output_tokens": 0, "latency_ms": 70.3}}
usage.input_tokens counts the prompt head that the questions share once.
Images. Put each image in the request as a data:image/...;base64,... URI. It can go anywhere inside state (as a
string, an object value or a list item), or in a top-level images list next to state and questions. The images
list also takes bare base64 strings and http(s) URLs (vLLM fetches a URL itself, so the server needs network access
for those; a URL it cannot fetch gets HTTP 422). Chat-style parts ({"type": "image_url", "image_url": {"url": ...}}) and multipart/form-data (the JSON in a
request field, each image as a file part) work too.
- Each image taken out of
stateleaves[image N]in its place, so the record can refer to it. Images in the top-level list are numbered first. - Up to 4 images per request and 20 MB per image.
- Each image is scaled to at most 1,048,576 pixels, and image answers use their own temperature, 0.65 (
serve_config.json,"image"); text answers keep 0.8. - A request without an image takes exactly the text path above.
This example uses example_receipt.png from this repository:
import base64, json, urllib.request
img = base64.b64encode(open("h2o-lightning-4b/example_receipt.png", "rb").read()).decode()
req = {"state": {"receipt": "data:image/png;base64," + img,
"claim": {"employee": "J. Ortiz", "category": "office supplies", "amount": "444.26"}},
"questions": {
"matches": {"type": "noul", "instructions": "The receipt total equals claim.amount.",
"criteria": {"true": "Same amount", "false": "Different amount"}},
"items": {"type": "choice", "instructions": "How many item lines are on the receipt?",
"criteria": {"three": "3", "four": "4", "five": "5", "six": "6"}},
"legible": {"type": "score", "instructions": "How legible is the receipt?",
"criteria": ["Unreadable", "Partly readable", "Fully readable"]}}}
r = urllib.request.urlopen(urllib.request.Request("http://127.0.0.1:8741/v1/systemone", json.dumps(req).encode(),
{"Content-Type": "application/json"}))
print(json.dumps(json.load(r), indent=1))
The response (rounded here):
{"model": "h2oai/h2o-lightning-4b",
"answers": {"matches": {"type": "noul", "noul": 0.964},
"items": {"type": "choice", "choice": "six",
"probabilities": {"three": 0.006, "four": 0.012, "five": 0.024, "six": 0.958}, "confidence": 0.943},
"legible": {"type": "score", "score": 1.926,
"legend": {"0": "Unreadable", "1": "Partly readable", "2": "Fully readable"},
"probabilities": {"0": 0.011, "1": 0.052, "2": 0.937}, "confidence": 0.906}},
"probability_source": "native", "effort_used": {"tier": "fast", "passes_per_decision": 1},
"usage": {"input_tokens": 585, "output_tokens": 0, "latency_ms": 141.0, "images": 1}}
usage.input_tokens includes the image tokens; usage.images counts the images read.
Status codes. An input the shim cannot take gets HTTP 422, which the runner scores as one wrong answer. That covers:
- over the context limit;
- more than 255 options;
- an unknown question type or a malformed question;
- more than 4 images, or an image that is not valid base64 image data.
An empty request gets 400. vLLM's 401, 403 and 429 pass through. Any other vLLM failure is a 502, which counts toward the runner's three-consecutive-failures stop.
Tests (no GPU): python3 h2o-lightning-4b/test_shim.py.
Limitations
- Evaluated mostly in English. Decides best with up to about 16 options; up to 255 are accepted.
- The probabilities are calibrated for decision questions of this kind, not for open-ended text.
- Not a chat model: it answers one decision per question and does not generate explanations.
- Images: no video input; counting many small objects and reading values off line charts are the weakest image skills we measured.
Disclosures
- No JevBench or ImageJevBench data, public or otherwise, was used for training.
- No rules keyed to any benchmark item's wording, id or answer; no network calls at serving time (except fetching an image URL a request supplies).
- Price basis for an evaluator: Qwen3.5-4B at bf16, one output token per decision; 654 mean input tokens on the
public text set. Image tokens are counted in
usage.input_tokens.
License
Apache-2.0. Built on Qwen/Qwen3.5-4B by the Qwen team (Alibaba Cloud), Apache-2.0; served with unmodified vLLM (Apache-2.0).
- Downloads last month
- 98