Nace.AI
NewDrex is now generally available

Introducing Drex.
A small model that decides.

Facts in, a probability for every option out. One pass, and no text to parse.

Get started
51.73
Decision Index 0.2, rank 1
+0.06
over Jev 1.13.0
22
of 39 tests won
<6B
total parameters
Decision Index 0.2Drex and the top 5 of 50 entries
51.73Drex
51.67
50.94
47.23
46.08
45.73
Drex<6BJevn/aAutoJev27.8BRune25.8BDecider27.8BJevfire27.8B

The axis starts at 44, not zero. Total parameters under each name; Jev's size is not published.

0:00 / 0:00

First on the public Decision Index. Ahead of every other model.

Decision Index 0.2 scores 40 tests of deciding rather than writing: contracts, causes, tool calls, routing, judgment. Every score is chance-corrected, so guessing at random scores zero. Drex leads all 50 entries with under 6B parameters; every other model in the top ten with a published size has 12B to 36B.

Drex 1.0Jev 1.13.0AutoJev-27BSurogate RuneDecider chat
OverallDecision Index 0.2, 40 tests. Every cell is chance-corrected: 0% is random guessing.51.7351.6750.9447.2346.08
Knowledge & ReasoningGSM8K · accuracy85.1%75.6%61.1%70.1%49.3%
Knowledge & ReasoningGPQA Diamond · accuracy15.7%71.4%32.6%26.5%39.5%
Language UnderstandingContractNLI · macro-F180.5%59.1%68.4%67.4%62.4%
Language UnderstandingANLI · macro-F145.7%62.2%56.0%56.8%61.9%
Retrieval & ClassificationBANKING77 · macro-F186.0%79.5%78.8%75.9%74.8%
Retrieval & ClassificationBRIGHT · nDCG@1012.6%15.1%15.5%14.3%14.7%
Tools & AutomationToolRet · nDCG@1042.2%39.1%40.7%38.7%38.5%
Tools & AutomationBFCL · case exact accuracy89.2%94.3%96.8%92.8%96.7%
Arts & Human TasteHumicroedit · accuracy31.5%23.7%24.8%26.2%21.7%
Arts & Human TasteBPoMP · accuracy50.1%81.8%87.8%81.7%75.7%
Knowledge & ReasoningChessBench · accuracy15.7%9.8%9.8%11.3%10.3%
Knowledge & ReasoningMuSR · accuracy31.7%46.0%39.3%45.0%40.6%
Language UnderstandingFinEntity · macro-F184.6%80.8%89.6%77.9%73.3%
Language UnderstandingWinoGrande · accuracy61.3%83.9%70.3%62.1%66.4%
Retrieval & ClassificationCLINC150 · macro-F194.1%89.2%87.7%86.9%85.0%
Tools & AutomationAPI-Bank · accuracy73.5%88.0%83.8%80.1%79.3%
Tools & AutomationRouterBench · selected quality (quality objective)52.8%52.7%52.4%52.7%52.6%
Arts & Human TasteForecastBench · Brier skill vs always 0.512.9%30.6%22.1%14.2%25.6%
Arts & Human TastePOP909 · accuracy73.8%15.9%37.3%12.9%17.5%
Knowledge & ReasoningSATA-Bench · case exact accuracy21.3%25.4%28.9%33.7%33.9%
Knowledge & ReasoningCRUXEval · accuracy62.4%57.1%60.2%52.7%42.9%
Language UnderstandingHellaSwag · accuracy91.2%92.7%92.0%89.6%93.9%
Language UnderstandingiSarcasmEval · Sarcasm F1 · track A, English64.4%36.3%49.0%36.4%30.3%
Tools & AutomationHome appliances · case exact accuracy16.3%52.5%70.6%38.8%41.9%
Retrieval & ClassificationSGD · macro-F110.1%5.1%0.0%0.0%0.0%
Knowledge & ReasoningHLE · accuracy0.0%4.4%0.0%0.0%0.0%
Tools & AutomationWhen2Call · accuracy88.1%74.6%75.6%65.7%71.3%
Language UnderstandingACOS · case exact accuracy0.2%6.3%0.5%0.5%0.8%
Arts & Human Tastecfcolor · accuracy29.6%28.8%28.8%24.2%26.8%
Knowledge & ReasoningMMLU-Pro · accuracy37.8%80.5%60.1%58.7%58.2%
Knowledge & ReasoningCLadder · accuracy80.7%45.3%49.0%40.6%32.7%
Language UnderstandingNLI4CT · macro-F153.3%69.0%70.5%61.4%65.0%
Language UnderstandingVAST · macro-F164.2%46.9%56.2%41.8%37.6%
Knowledge & ReasoningBBH · accuracy47.4%89.7%68.3%62.3%60.5%
Retrieval & ClassificationAmazon ESCI · macro-F145.4%43.8%43.7%41.1%34.4%
Arts & Human TasteHabermas · accuracy41.3%21.5%15.5%18.0%14.2%
Language UnderstandingRAGTruth · F1 on hallucinated class72.3%60.1%66.6%61.7%57.6%
Retrieval & ClassificationHoVer · accuracy70.8%45.7%48.4%43.0%36.0%
Arts & Human TasteNew Yorker · accuracy65.9%62.6%62.8%57.6%65.9%
Retrieval & ClassificationPhishNChips · accuracy—25.1%19.9%57.5%35.1%

Drex was scored on 2026-09-25 with the official scorer. PhishNChips has not been scored for Drex yet and is left out of its Retrieval average; unanswered rows everywhere else count as wrong. Other entries are the public board's published scores, read 2026-09-24.

Ask it a question. It tells you how likely each answer is.

Send Drex the facts of a case and the options you would accept. One pass through the model returns a probability for each one. It never writes text, so there is nothing to parse and nothing it can make up.

You send

Invoice INV-88214 bills 1,200 units at $14.10. PO-5512 authorised 1,200 at $13.90. Goods receipt confirms 1,200 delivered, 2 days late.

What should happen to this invoice?

hold for buyer · approve · dispute

You get back

hold for buyer61%
approve30%
dispute9%

A probability for every option. Nothing written, nothing to parse. Example values; recorded answers are further down the page.

It is not a chatbot. It is the other kind of thinking.

Psychologists call the deliberate, step-by-step kind System 2 and the practised, intuitive kind System 1. A chat model is System 2: it writes its way to a decision. A decision model is System 1: it has seen the shape of the decision before and answers in one pass, with how sure it is.

How it answers

A chat model: Thinks out loud, one word at a time, then names an answer at the end.

Drex: Reads the case once and scores every option in the same pass.

What comes back

A chat model: A sentence, or JSON you parse and hope names one of your options.

Drex: A probability for each option you offered. Nothing else can come back.

A distribution, not a verdict.

Because the answer is a probability for every option, you can see how close the call was. Threshold it, route the uncertain cases to a person, and act on the confident ones without reading a word.

Each board is an example case, and each pile is the probability on one option. Recorded answers are further down the page.

Accounts payableWhat should happen to this invoice?
 hold for buyer
 approve
 dispute

So it is like Jev? Same kind of model. Ahead of it on the index.

Jev is the decision model the public index is named after. Drex answers the same requests in the same shape, scores higher on the index, and reads 5× fewer tokens to do it.

On the indexDecision Index 0.2, Jev 1.13.0 at 51.67
+0.00points ahead, winning 22 of 39 tests
Tokens read per decisionmedian of eight identical requests
Drex70 tokens
Jev 1.13.0367 tokens

It also beats Jev at board games, 117 to 92.

Eight games, 256 matches, every legal move offered as an option. Neither model looks ahead; each move is one pass over the board. Drex won more matches than Jev in five of the eight, and in the total.

Drex 11747 drawsJev 92
  • Checkers19–9–473.4%
  • Four in a Row22–0–1068.8%
  • Chess8–22–259.4%
  • Clobber18–0–1456.3%
  • Nine Men's Morris13–10–956.3%
  • Reversi15–0–1746.9%
  • Dots and Boxes14–0–1843.8%
  • Lines of Action8–6–1834.4%

A head-to-head we ran against Jev 1.13.0 with the Kaggle Game Arena harness, OpenSpiel rules. Every game has 16 openings, each played from both sides, so neither model keeps the first move. Score counts a win as one and a draw as a half.

See it answer. Real cases, recorded replies.

Pick a case. The same text and questions went to both columns, and neither wrote a sentence back.

Choice + yes/no on the same customer message the official Jev docs example.

The case

Customer: I was charged twice last Tuesday and I am furious. Order #4419. I want a refund today.

The questions

topic choice

What is the issue about?

billing · bug · shipping

urgent noul

Escalate to a human now?

recorded 2026-09-24

Drex

72 tokens read

topic choice

billing0.997
bug0.001
shipping0.001

urgent noul

noyes0.732

Jev 1.13.0

365 tokens read

topic choice

billing1.000
bug0.000
shipping0.000

urgent noul

noyes0.490

Run your own case.

Paste a case, name the options, and get the probabilities back. Free to try, no card needed.

Get started

On topic, Jev returns a flat 1.000 and 0.000 where Drex returns a distribution you can threshold.

Runs where your data lives.

Small enough to own, and to keep. Every way of running Drex answers the same request in the same shape.

Self-hosted

In your own cloud account, at the edge or on-premises. The model fits on one accelerator, so the decision layer never leaves your network.

Managed

On Nace-managed infrastructure, with capacity, scaling and upgrades handled for you.

Hybrid

Keep the decisions that touch regulated data in your environment and send the rest to managed capacity. Same request, same answer.

Tuned on your decisions

The results on this page come from training on decision data. We tune Drex on your labelled outcomes and ship you the weights.

Questions, answered.

What is a decision model, and how is it different from an LLM?

A decision model answers closed questions with numbers instead of writing text. You send the facts and the options; one forward pass returns a probability for every option. There is no token stream, no chain of thought and no JSON to repair, so the failure modes of chat models (truncation, malformed output, invented options) do not arise. Drex can only ever answer with options you gave it.

What kinds of questions can it answer?

Three. Pick one of several: named options in, the chosen one plus a probability for each out. Yes or no: the probability that the answer is yes. Rate on a scale: ordered levels in, the expected level plus the spread across levels out. Several questions about one case can share a single pass.

How does it compare to Jev?

Drex scores 51.73 on Decision Index 0.2, ahead of Jev 1.13.0 at 51.67 and of the strongest community entry, AutoJev-27B, at 50.94. It wins 22 of the 39 benchmarks scored so far, with its largest margins on chord recognition, causal inference, sarcasm, multi-hop fact checking and contract reasoning, and it trails on graduate-level science, broad subject knowledge and multi-step reasoning puzzles. Every result, including every loss, is in the table on this page.

How big is it?

Under 6B parameters, the smallest model in the top ten of Decision Index 0.2. Every other top-ten model with a published size has 12B to 36B parameters, and Jev's size is not published. The strongest other model under 6B on the board, Hopper, scores 36.71, fifteen points below Drex. Sizes are total parameters: mixture-of-experts models such as Decider 35B-A3B use only part of their weights per request, but still have to hold all of them in memory.

Is it trained on the benchmarks it is measured on?

It is trained on the official training splits of the index benchmarks alongside verifiable procedural data, and evaluated only on the held-out splits with the official scorer. That is how the leaderboard is designed to be run. It also explains the shape of the results: the model is strong where it has seen the kind of decision before, which is exactly the property that makes it worth tuning on your decisions.

Can it run inside our environment, on our decisions?

Yes. It fits on a single accelerator, so it runs in your cloud account, at the edge or on-premises, as well as on Nace-managed infrastructure or a mix of the two. Every deployment answers the same request in the same shape. We also tune Drex on your labelled outcomes and ship you the weights.

Put a decision model where the guesswork is.

Free to try today. Tuned on your decisions and served in your environment when you are ready.