Rene-1 31B FP8

One forward pass reads a document and returns a calibrated probability for every option of every question you ask. Nothing is generated.

Rene-1 is a decision model. You send a document and a set of typed questions: pick one option, answer yes or no, or place the document on a scale. It reads the document and every question in one pass, each question on its own, and answers each with a probability for every option you declared. There is no generated text to parse. The weights are stored in FP8, so the whole model runs on one GPU.

Skill 64.18 balanced skill on 37 of the 40 Decision Index 0.2 benchmarks; on the same benchmarks the board's rows score Jev 54.75 and AutoJev-27B 53.90
Calibration median calibration error 0.046 across those 37 benchmarks
Latency 92.1 ms at the median for 5 questions over 641 input tokens, on one NVIDIA B200
Size 33.3 GB of FP8 weights; 44.5 GB at peak on typical requests (one L40S, H100, H200 or B200), 67.9 GB at the token limit (one H100, H200 or B200)
Licence Apache 2.0

Quickstart

Install

Install PyTorch first, built for your CUDA version (torch>=2.14.0, from pytorch.org), then:

pip install "transformers>=5.17.0" "compressed-tensors>=0.19.0" "accelerate>=1.15.0" "safetensors>=0.8.0" "pillow>=12.3.0"

Rene-1 needs an NVIDIA GPU that runs FP8 and holds its weights, such as those under Hardware, below. On any other device it refuses to load, and says why.

Python

Every classifier gets asked the hot dog question sooner or later. One request asks three typed questions about a dish: yes or no, which cuisine, and how messy it is to eat.

from transformers import AutoModel

model = AutoModel.from_pretrained("salfatigroup/rene-1-31b-fp8", trust_remote_code=True, device_map="auto")

state = {"dish": "A grilled sausage in a split bun with mustard"}
questions = {
  "is_hotdog": {
    "type": "yesno",
    "instructions": "Is the food in `dish` a hot dog?",
    "criteria": {
      "true": "A sausage served in a sliced bun.",
      "false": "Anything else."
    }
  },
  "cuisine": {
    "type": "choice",
    "instructions": "Which cuisine does `dish` belong to?",
    "criteria": {
      "american": "US street food or diner food.",
      "german": "German cooking.",
      "other": "Anything else."
    }
  },
  "messiness": {
    "type": "score",
    "instructions": "How messy is `dish` to eat by hand?",
    "criteria": [
      "clean",
      "a little messy",
      "very messy"
    ]
  }
}

answers = model.decide(state, questions)
print(answers["is_hotdog"])
The response to the hotdog request

The response:

{
  "model": "rene-1-31b-fp8",
  "answers": {
    "is_hotdog": {
      "type": "yesno",
      "probability": 0.94
    },
    "cuisine": {
      "type": "choice",
      "choice": "american",
      "confidence": 0.63,
      "probabilities": {
        "american": 0.75,
        "german": 0.22,
        "other": 0.03
      }
    },
    "messiness": {
      "type": "score",
      "score": 1.14,
      "confidence": 0.63,
      "legend": {
        "0": "clean",
        "1": "a little messy",
        "2": "very messy"
      },
      "probabilities": {
        "0": 0.06,
        "1": 0.75,
        "2": 0.19
      }
    }
  },
  "usage": {
    "input_tokens": 354,
    "output_tokens": 0
  }
}

Each response carries:

  • type on every answer.
  • probability for a yes or no question: the probability of true.
  • probabilities for a choice, keyed by option, and for a score, keyed by level number as a string, "0" for the lowest. They sit on a two-decimal grid and sum to 1.
  • legend for a score: each level number with its text.
  • choice, the most likely option. Its confidence is (K x top probability - 1) / (K - 1) over K options: 0 for an even split, 1 for certainty.
  • score, the expected level, counting levels from 0. Its confidence is 1 minus the expected distance from the most likely level, scaled so an even spread scores 0.
  • usage.output_tokens, which is 0 because nothing is generated.

trust_remote_code=True runs the model code in this repository (modeling_rene.py and the files next to it). Read it first, and pin the commit you read with revision= in production.

  • model.decide(state, questions) returns the answers, keyed by your question names.
  • model.decide_batch(requests) answers up to 64 requests, 65,536 packed tokens in total, in one call. On a GPU a batched answer can differ from the single call's by at most 0.01 on the two-decimal grid.
  • model.respond(request) and model.respond_batch(requests) return the whole response object, model and usage included.
  • model.render(state, questions) shows the exact sequence the model reads.
  • The tokenizer: AutoTokenizer.from_pretrained("salfatigroup/rene-1-31b-fp8", subfolder="trunk"), or model.tokenizer.
  • At load it warns once if an installed package differs from the versions it was tested with, listed in release.json.

Images (experimental)

A field of state can hold an image: a PIL image or the file's bytes in Python, or the JSON form below for model.respond (PNG, JPEG or WebP). Refer to it by name in backticks, like any other field. Up to 4 images per request; each adds about 280 input tokens. URLs and file paths are refused, so the code never fetches anything. Image input is experimental: Rene-1 was trained on text only, image answers have not been evaluated, and the model says so once when it serves one.

from PIL import Image

questions = {
  "is_hotdog": {
    "type": "yesno",
    "instructions": "Is the food in `photo` a hot dog?"
  },
  "dish": {
    "type": "choice",
    "instructions": "Which dish is in `photo`?",
    "criteria": {
      "hot_dog": "A sausage in a sliced bun.",
      "corn_dog": "A battered sausage on a stick.",
      "sandwich": "Filling between slices of bread.",
      "other": "Anything else."
    }
  }
}

answers = model.decide({"photo": Image.open("lunch.png")}, questions)
print(answers["is_hotdog"])

The same request as JSON:

{
  "state": {
    "photo": {
      "type": "image",
      "base64": "<the file, base64>",
      "media_type": "image/png"
    }
  },
  "questions": {
    "is_hotdog": {
      "type": "yesno",
      "instructions": "Is the food in `photo` a hot dog?"
    },
    "dish": {
      "type": "choice",
      "instructions": "Which dish is in `photo`?",
      "criteria": {
        "hot_dog": "A sausage in a sliced bun.",
        "corn_dog": "A battered sausage on a stick.",
        "sandwich": "Filling between slices of bread.",
        "other": "Anything else."
      }
    }
  }
}
The response for `hot-dog-drawing.png`, one of the drawings in the video above
{
  "model": "rene-1-31b-fp8",
  "answers": {
    "is_hotdog": {
      "type": "yesno",
      "probability": 0.78
    },
    "dish": {
      "type": "choice",
      "choice": "hot_dog",
      "confidence": 0.94,
      "probabilities": {
        "hot_dog": 0.96,
        "corn_dog": 0.0,
        "sandwich": 0.01,
        "other": 0.03
      }
    }
  },
  "usage": {
    "input_tokens": 590,
    "output_tokens": 0
  }
}
The response for `corn-dog-drawing.png`, one of the drawings in the video above
{
  "model": "rene-1-31b-fp8",
  "answers": {
    "is_hotdog": {
      "type": "yesno",
      "probability": 0.34
    },
    "dish": {
      "type": "choice",
      "choice": "other",
      "confidence": 0.31,
      "probabilities": {
        "hot_dog": 0.1,
        "corn_dog": 0.4,
        "sandwich": 0.01,
        "other": 0.49
      }
    }
  },
  "usage": {
    "input_tokens": 580,
    "output_tokens": 0
  }
}

The same model as a transformers pipeline, which returns the whole response object:

from transformers import pipeline

rene = pipeline(model="salfatigroup/rene-1-31b-fp8", trust_remote_code=True, device_map="auto")
response = rene({"state": state, "questions": questions})

Real decisions

Route a support ticket

Pick the queue, flag the tickets that cannot wait, and read the customer's mood, in one pass.

The request and the response

The request:

{
  "state": {
    "subject": "Charged twice for the annual plan",
    "body": "I moved to the annual plan yesterday and my card shows the same charge twice. I only want one plan. Please refund the duplicate before it posts.",
    "plan": "business"
  },
  "questions": {
    "queue": {
      "type": "choice",
      "instructions": "Which team should handle the ticket in `subject` and `body`?",
      "criteria": {
        "billing": "Charges, refunds, invoices and plan changes.",
        "technical": "Errors, outages, bugs and integrations.",
        "account": "Sign-in, access, security and profile settings.",
        "sales": "Questions from people who have not bought yet."
      }
    },
    "urgent": {
      "type": "yesno",
      "instructions": "Does the ticket in `body` need a reply within the hour?",
      "criteria": {
        "true": "Money taken in error, a security problem, or a service that is down.",
        "false": "Anything that can wait for the normal queue."
      }
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated is the customer who wrote `body`?",
      "criteria": [
        "calm",
        "annoyed",
        "angry"
      ]
    }
  }
}

The response:

{
  "model": "rene-1-31b-fp8",
  "answers": {
    "queue": {
      "type": "choice",
      "choice": "billing",
      "confidence": 1.0,
      "probabilities": {
        "billing": 1.0,
        "technical": 0.0,
        "account": 0.0,
        "sales": 0.0
      }
    },
    "urgent": {
      "type": "yesno",
      "probability": 0.62
    },
    "frustration": {
      "type": "score",
      "score": 0.54,
      "confidence": 0.19,
      "legend": {
        "0": "calm",
        "1": "annoyed",
        "2": "angry"
      },
      "probabilities": {
        "0": 0.56,
        "1": 0.34,
        "2": 0.1
      }
    }
  },
  "usage": {
    "input_tokens": 435,
    "output_tokens": 0
  }
}

Review a contract clause

Classify the clause, check whether its right runs both ways, and place its risk on a four-level scale.

The request and the response

The request:

{
  "state": {
    "clause": "Either party may end this contract for convenience with thirty days' written notice. On termination the customer pays all fees accrued up to the termination date, and the provider refunds any prepaid fees for the period after it."
  },
  "questions": {
    "clause_type": {
      "type": "choice",
      "instructions": "What kind of clause is `clause`?",
      "criteria": {
        "termination": "When and how the contract can end.",
        "liability": "Caps on, or exclusions of, damages.",
        "payment": "Fees, invoicing and payment terms.",
        "confidentiality": "Duties to keep information secret.",
        "other": "Anything else."
      }
    },
    "mutual": {
      "type": "yesno",
      "instructions": "Can both parties use the right that `clause` grants?"
    },
    "customer_risk": {
      "type": "score",
      "instructions": "How much risk does `clause` put on the customer?",
      "criteria": [
        "low",
        "moderate",
        "high",
        "severe"
      ]
    }
  }
}

The response:

{
  "model": "rene-1-31b-fp8",
  "answers": {
    "clause_type": {
      "type": "choice",
      "choice": "termination",
      "confidence": 0.97,
      "probabilities": {
        "termination": 0.98,
        "liability": 0.0,
        "payment": 0.01,
        "confidentiality": 0.0,
        "other": 0.01
      }
    },
    "mutual": {
      "type": "yesno",
      "probability": 0.85
    },
    "customer_risk": {
      "type": "score",
      "score": 0.38,
      "confidence": 0.62,
      "legend": {
        "0": "low",
        "1": "moderate",
        "2": "high",
        "3": "severe"
      },
      "probabilities": {
        "0": 0.77,
        "1": 0.13,
        "2": 0.05,
        "3": 0.05
      }
    }
  },
  "usage": {
    "input_tokens": 401,
    "output_tokens": 0
  }
}

Check a post against your policy

The policy is part of the request, so changing a rule needs no retraining.

The request and the response

The request:

{
  "state": {
    "policy": "Harassment: posts must not insult, threaten or demean a person or a group. Criticism of ideas, products and the public actions of public figures is allowed.",
    "post": "This update broke my workflow again. Whoever shipped this release has clearly never used the product."
  },
  "questions": {
    "breaks_policy": {
      "type": "yesno",
      "instructions": "Does `post` break `policy`?"
    },
    "content": {
      "type": "choice",
      "instructions": "What is `post` mainly doing?",
      "criteria": {
        "product_criticism": "Criticising a product, a release or a decision.",
        "personal_attack": "Insulting or demeaning a person or a group.",
        "threat": "Threatening harm.",
        "spam": "Advertising or repeated content.",
        "other": "Anything else."
      }
    },
    "severity": {
      "type": "score",
      "instructions": "If `post` breaks `policy`, how serious is it?",
      "criteria": [
        "not a violation",
        "mild",
        "serious"
      ]
    }
  }
}

The response:

{
  "model": "rene-1-31b-fp8",
  "answers": {
    "breaks_policy": {
      "type": "yesno",
      "probability": 0.25
    },
    "content": {
      "type": "choice",
      "choice": "product_criticism",
      "confidence": 0.55,
      "probabilities": {
        "product_criticism": 0.64,
        "personal_attack": 0.31,
        "threat": 0.01,
        "spam": 0.01,
        "other": 0.03
      }
    },
    "severity": {
      "type": "score",
      "score": 0.46,
      "confidence": 0.31,
      "legend": {
        "0": "not a violation",
        "1": "mild",
        "2": "serious"
      },
      "probabilities": {
        "0": 0.61,
        "1": 0.33,
        "2": 0.06
      }
    }
  },
  "usage": {
    "input_tokens": 412,
    "output_tokens": 0
  }
}

Questions and answers

Type You declare You get back
yesno an instruction, descriptions of true and false, or both type and probability, the probability of true
choice 1 to 249 options, each with an optional description, and an optional instruction type, probabilities keyed by option, choice and confidence
score 2 to 10 ordered levels, lowest first, and an optional instruction type, probabilities keyed by level number, legend, score (the expected level) and confidence
  • state is a string, an object or a list. Instructions refer to its fields by name in backticks.
  • Questions are keyed by your own names, and answers come back under the same keys.
  • A choice question with a single option is answered with certainty; the model does not read it.
  • An unknown field in a request or a question is refused, not ignored, so a typo fails loudly.
  • Each question sees the document and itself, never the other questions. The document plus any one question can hold 32,512 tokens, and a whole request 65,280, counted on the text you send (objects as compact JSON).
  • A request the model refuses raises an error that says why.

Results

Decision Index 0.2

The Decision Index is a public panel of 40 decision benchmarks in 5 areas. Its headline, balanced skill, corrects each benchmark for chance and weighs the areas equally. We rebuilt 37 of the 40 benchmarks from their public sources, 91.81 of its 100 points: 28 with the index's own kit (apolinario/decision-index at 8e12d6d7); BANKING77 and Habermas with the kit's own code, run on source files its pinned copy does not hold; and the 7 benchmarks new in Decision Index 0.2 (MMLU-Pro, BBH, RAGTruth, PhishNChips, HoVer, When2Call and New Yorker) with our own code. We evaluated Rene-1 on those rows at temperature 1, with no calibration set, on NVIDIA B200. The other rows are the index's own board, the maintainers' runs, averaged over the same 37 benchmarks; Rene-1 has not been scored by them. The edition is pinned to Decision Index 0.2, the board of 2026-09-24. The index has since published Decision Index 0.2.1 (board of 2026-09-27), a different panel: 38 benchmarks, with SGD and RouterBench out of the index; new area weights; 13 benchmarks weighted 1.2 inside their area; and changed scoring, chance or row counts on MuSR, HLE, ACOS, RAGTruth, BRIGHT, ToolRet and Home appliances. This page uses Decision Index 0.2 only and never mixes the two.

Not evaluated (3):

  • ToolRet: its 0.2 row selection is not published, so an evaluation would not follow the 0.2 protocol.
  • BRIGHT: its 0.2 row selection is not published, so an evaluation would not follow the 0.2 protocol.
  • HLE: its rows need an access request on the Hub, which we did not make.
Model The 37 benchmarks evaluated All 40
Rene-1 31B FP8 64.18 58.61, the 3 not evaluated scored as zero
Jev, board 54.75 51.67, as published
AutoJev-27B, board 53.90 50.94, as published
Surogate Rune 26B-A4B, board 50.08 47.23, as published
Decider chat · Qwen3.6-27B, board 48.74 46.08, as published
Winnow-12B, board 47.72 45.05, as published
  • On the same 37 benchmarks, Rene-1 places 1 of 51 against all 50 rows of the Decision Index 0.2 board.

By area

Area Benchmarks evaluated Rene-1 31B FP8 Jev AutoJev-27B Surogate Rune 26B-A4B Decider chat · Qwen3.6-27B Winnow-12B
Knowledge & Reasoning 9 of 10 55.8 55.7 45.5 44.6 40.9 37.8
Language Understanding 10 of 10 67.5 59.7 61.9 55.6 54.9 55.4
Retrieval & Classification 6 of 7 61.8 48.1 46.4 50.7 44.2 46.4
Tools & Automation 5 of 6 80.4 72.4 75.8 66.0 68.4 68.8
Arts & Human Taste 7 of 7 55.4 37.9 39.9 33.5 35.3 30.2

Every column averages skills over the benchmarks evaluated here, so the board rows are on the same 37 as the first.

Calibration

At temperature 1 and with no calibration set, the median calibration error (ECE, 10 equal-mass bins) across the 37 benchmarks is 0.046. The last column below gives it per benchmark.

Every benchmark

Skill is each benchmark's own score corrected for chance, (score - chance) / (1 - chance), from 0 at chance to 100 for a perfect answer. A score below chance is floored at 0 and marked (floored). The best score on each row is in bold. ForecastBench, where a lower score is better, enters as clip((0.25 - Brier) / (0.25 - 0)) x coverage, so a higher skill is better there too.

Area Benchmark Skill computed from Rene-1 Jev AutoJev-27B Surogate Rune 26B-A4B Decider chat · Qwen3.6-27B Winnow-12B Calibration error, Rene-1
Knowledge & Reasoning CRUXEval accuracy 76.9 57.1 60.2 52.7 42.9 51.6 0.025
Knowledge & Reasoning CLadder accuracy 76.0 45.3 49.0 40.6 32.7 37.1 0.028
Knowledge & Reasoning GSM8K accuracy 73.7 75.6 61.1 70.1 49.3 44.7 0.046
Knowledge & Reasoning BBH accuracy 72.1 89.7 68.3 62.3 60.5 58.4 0.043
Knowledge & Reasoning MMLU-Pro accuracy 63.6 80.5 60.1 58.7 58.2 49.2 0.025
Knowledge & Reasoning MuSR accuracy 57.3 46.0 39.3 45.0 40.6 38.1 0.143
Knowledge & Reasoning GPQA Diamond accuracy 34.0 71.4 32.6 26.5 39.5 18.4 0.106
Knowledge & Reasoning SATA-Bench case exact accuracy 33.5 25.4 28.9 33.7 33.9 32.7 0.086
Knowledge & Reasoning ChessBench accuracy 14.8 9.8 9.8 11.3 10.3 10.0 0.097
Language Understanding FinEntity macro-F1 91.7 80.8 89.6 77.9 73.3 87.8 0.125
Language Understanding HellaSwag accuracy 90.4 92.7 92.0 89.6 93.9 80.2 0.046
Language Understanding WinoGrande accuracy 86.7 83.9 70.3 62.1 66.4 48.2 0.020
Language Understanding ContractNLI macro-F1 76.9 59.1 68.4 67.4 62.4 60.8 0.029
Language Understanding NLI4CT macro-F1 66.7 69.0 70.5 61.4 65.0 66.8 0.027
Language Understanding ANLI macro-F1 65.9 62.2 56.0 56.8 61.9 50.2 0.052
Language Understanding RAGTruth F1 on hallucinated class 64.8 60.1 66.6 61.7 57.6 61.1 0.013
Language Understanding VAST macro-F1 63.8 46.9 56.2 41.8 37.6 52.9 0.229
Language Understanding iSarcasmEval Sarcasm F1 · track A, English 53.9 36.3 49.0 36.4 30.3 45.6 0.081
Language Understanding ACOS case exact accuracy 13.7 6.2 0.5 0.5 0.8 0.8 0.017
Retrieval & Classification CLINC150 macro-F1 94.8 89.2 87.7 86.9 85.0 86.9 0.044
Retrieval & Classification BANKING77 macro-F1 92.0 79.5 78.8 75.9 74.8 74.8 0.015
Retrieval & Classification HoVer accuracy 77.6 45.7 48.4 43.0 36.0 44.1 0.023
Retrieval & Classification Amazon ESCI macro-F1 47.0 43.8 43.7 41.1 34.4 41.1 0.070
Retrieval & Classification PhishNChips accuracy 41.1 25.1 19.9 57.5 35.1 31.8 0.527
Retrieval & Classification SGD macro-F1 18.3 5.1 0.0 0.0 0.0 0.0 0.423
Tools & Automation BFCL case exact accuracy 95.1 94.3 96.8 92.8 96.7 92.9 0.002
Tools & Automation Home appliances case exact accuracy 90.0 52.5 70.6 38.8 41.9 54.4 0.005
Tools & Automation API-Bank accuracy 84.1 88.0 83.8 80.1 79.3 66.7 0.122
Tools & Automation When2Call accuracy 81.9 74.6 75.6 65.7 71.3 77.6 0.021
Tools & Automation RouterBench selected quality (quality objective) 51.0 52.7 52.4 52.7 52.6 52.2 0.166
Arts & Human Taste BPoMP accuracy 95.7 81.8 87.8 81.7 75.7 66.8 0.002
Arts & Human Taste New Yorker accuracy 83.2 62.6 62.8 57.6 65.9 56.0 0.031
Arts & Human Taste POP909 accuracy 53.2 15.9 37.3 12.9 17.5 7.3 0.372
Arts & Human Taste Habermas accuracy 45.5 21.5 15.5 18.0 14.2 15.6 0.087
Arts & Human Taste Humicroedit accuracy 41.2 23.7 24.8 26.2 21.7 22.5 0.068
Arts & Human Taste ForecastBench Brier (lower is better) 35.2 30.6 22.1 14.2 25.6 23.0 0.027
Arts & Human Taste cfcolor accuracy 33.8 28.8 28.8 24.2 26.8 20.0 0.093

JevBench

JevBench at commit 1bcc55e, run on its 231 public items (48 easy, 72 standard, 111 hard) and scored with its own v1.3.0 formulas, the last it defines on the public items alone. Its judge tier is not public, so the benchmark's own rule spreads that weight over the other tiers. The benchmark's own board also counts sealed items that only its maintainers run, so its board rows are not set beside this local run.

Model Intelligence Calibration Speed Cost Score
Rene-1 31B FP8 85.86 85.91 89.11 not scored not scored
  • On the 111 hard items Rene-1 is right on 77.5%. Its weakest kinds of hard item are date and number arithmetic (2 of 15 right) and long policy documents (12 of 19 right).
  • Speed follows the benchmark's rule for a self-hosted GPU: it times the standard and judge items one request at a time. The judge items are not public, so Rene-1 is timed here on the 72 standard items, on NVIDIA B200, and each latency is multiplied by 2 with 0.15 s added before scoring. Measured: p50 97.6 ms and p95 102.9 ms.
  • Cost needs a public per-token price, which this page does not set, so Cost and the combined Score are left out.
Tier Items Rene-1
Easy 48 100.0%
Standard 72 100.0%
Hard 111 77.5%

Latency and throughput

Measured on 1x NVIDIA B200 (NVIDIA B200, driver 595.91.07) with torch 2.14.0+cu130 (CUDA 13.0), transformers 5.17.0 and compressed-tensors 0.19.0. Each row times model.decide end to end, as a caller sees it: request checks, tokenization, the one forward pass and the answer arithmetic, synchronised with the GPU, after warm-up calls. p95 is the nearest rank. The weights take 33.3 GB on disk; peak GPU memory in these runs was 39.3 GB allocated and 44.5 GB reserved. Loading took 12.9 s.

Request Questions Input tokens p50 p95 Mean Calls timed
Yes or no, short document 1 478 91.1 ms 93.6 ms 91.4 ms 30
Mixed questions, short document 5 641 92.1 ms 93.1 ms 92.2 ms 30
Many mixed questions, longer document 20 2,006 296.1 ms 296.6 ms 296.1 ms 30
One choice with many options, short document 1 1,233 147.8 ms 147.8 ms 147.8 ms 30
Yes or no, long document 1 4,278 676.6 ms 677.3 ms 676.7 ms 30
Request Batch Requests per second Input tokens per second Batches timed
Mixed questions, short document 8 17.24 11,050 10
Mixed questions, short document 32 17.61 11,290 10

Throughput runs model.decide_batch on distinct requests and divides by the median batch time.

Hardware

FP8 matrix multiplies need compute capability 8.9 or newer: Ada Lovelace, Hopper or Blackwell. Rene-1 checks the GPU when it loads and refuses anything else, rather than run a slower copy that was never measured. At 44.5 GB peak, one GPU holds the model: L40S (48 GB), H100 (80 GB), H200 (141 GB) or B200 (180 GB). L4 (24 GB) has the compute capability but not the memory. A request at the token limit (62,976 input tokens) peaked at 67.9 GB, so requests that long need H100 (80 GB), H200 (141 GB) or B200 (180 GB).

Intended use

  • Classifying, routing and triaging text against labels you define in each request.
  • Many questions about one document in a single pass: checks on a contract, a ticket, a transcript or a record.
  • Grading on an ordinal scale, with the uncertainty attached.
  • Thresholds and review queues: act on the confident answers, send the uncertain ones to a person.

Out of scope

  • The sole basis for a decision about a person's access to credit, work, housing, healthcare, education, legal status or benefits. Keep a person in the loop, and test on your own data first.
  • Free text of any kind: summaries, explanations, chat. Rene-1 writes nothing.
  • Multi-step arithmetic, date arithmetic and long chains of reasoning.
  • Security and fraud triage without your own evaluation.
  • Languages other than English. It was not evaluated on them.
  • Images. The model code accepts an image inside state as an experimental input, but the decision layer was trained on text only and image answers have not been evaluated; the model warns once when it serves one.

Limitations

  • Its weakest benchmarks by skill are ACOS (13.7), ChessBench (14.8) and SGD (18.3).
  • The board's leading row scores higher than Rene-1 on 8 of the 37 benchmarks evaluated; the table under Results names them.
  • Its calibration is weakest on PhishNChips (calibration error 0.527), SGD (calibration error 0.423) and POP909 (calibration error 0.372). Check its confidence on data like these before you threshold it.
  • On the hard items of the third-party benchmark in Results it is right on 77.5%; it misses most on date and number arithmetic and long policy documents.
  • It gives no reasons. A probability is not an explanation.
  • It has not been tested for sensitivity to the order or the names of the options.
  • Speed depends on the GPU and on the shape of the request. Measure yours.

How it was made

Rene-1 is a 31B transformer fine-tuned in full from open weights: the language model was trained for one epoch on 396,500 rows from 56 task families under a log-score loss, to put its probability on the right option. The trained weights are stored in compressed-tensors FP8_DYNAMIC format: E4M3 weights with one scale per output channel, activations quantized per token at run time, so no calibration data. The 410 linear layers of the decoder are FP8; the embeddings, the norms and the output layer that scores the answers stay at higher precision.

Files

Path What it holds
config.json the model config (model_type rene) and the auto_map that points transformers at the code
configuration_rene.py, modeling_rene.py, pipeline_rene.py, decision_contract.py, decision_prompt.py, decision_image.py the code trust_remote_code=True runs: loading, the GPU check, the request contract, the prompt layout, the experimental image input and decide
requirements.txt the dependencies, with torch left to your own CUDA build
head.safetensors the output layer that scores the answer labels
release.json provenance: the quantization, the contract's limits, the tested versions and the sha256 of every file
trunk/ the transformer in compressed-tensors FP8: config, weight shards and index, tokenizer, chat template, generation and processor config
LICENSE the Apache License 2.0

Licence

Apache License 2.0: see LICENSE. The weights in this repository are modified from Apache 2.0 open weights; the change notices (NOTICE) and the training data's sources and terms (ATTRIBUTIONS.md) will be added in a later update.

Citation

@misc{salfati2026rene1,
  title        = {Rene-1: A 31B Decision Reader That Knows When It Is Sure},
  author       = {Salfati, Elon},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/salfatigroup/rene-1-31b-fp8}}
}
Downloads last month
97
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using salfatigroup/rene-1-31b-fp8 1

Evaluation results