Nativ (נתיב)

The best Hebrew decision model for commercial use (as of October 2026). It beats Laya multilingual, the leading commercially licensed alternative, on every test below, and runs on a laptop CPU.

Give it any Hebrew text, a question and the options, and it returns a probability for every option in a single forward pass. No generation, no prompt parsing, no GPU.

נתיב (nativ) is Hebrew for "path": the model chooses which way a request should go.

Size 364M parameters (dicta-il/neodictabert + an option scorer)
Speed ~70 ms per decision on a laptop CPU
Output a probability per option, with its calibration error reported for every task
Options any Hebrew text, 2 to 7 per question
License CC BY 4.0, commercial use allowed
Benchmark nativ-bench
Training data hebrew-intent-emotion (translated part)

Usage

decider.py is included in this repo (requires torch, transformers, safetensors).

from decider import HebrewDecider

nativ = HebrewDecider.from_pretrained("path/to/nativ-he-decision")
nativ.decide(
    state="שלום, ההזמנה שלי הייתה אמורה להגיע ביום שלישי ועדיין לא קיבלתי אותה. אפשר לבדוק מה קורה?",
    question="לאיזו מחלקה להעביר את הפנייה?",
    options=["משלוחים", "החזרות והחלפות", "חיובים ותשלומים", "תמיכה טכנית"],
)
# one probability per option, in the order given
  • Options are free Hebrew text. Short descriptions work better than label names ("לקבוע שעון מעורר" rather than "alarm_set").
  • decide_many takes a list of {"state", "question", "options"} dicts.
  • The input is limited to 1,024 tokens. Long states are truncated; the question and options are kept.

Results

The two nativ-bench tests (nativ-bench) use scenarios, rules and numbers that do not appear in Nativ's training data. The last four tests are public Hebrew datasets Nativ never trained on.

  • † - marks a model that trained on other examples from the same dataset.
  • 🟢 - marks the tests where Nativ scores highest.

Accuracy (higher is better)

The share of questions answered correctly, from 0 to 1: 0.95 means 95 of every 100 answers are right.

test (dataset) Nativ Laya multilingual Laya-Hebrew DeepSeek V4 Flash
what does the user want, 4 options (MASSIVE) 0.962 † 🟢 0.644 0.887 0.905
does the paragraph answer the question (HeQ) 0.852 † 🟢 0.582 0.694 † 0.847
sentiment of a comment (OnlpLab) 0.905 † 🟢 0.680 0.822 † 0.838
personal information in a message (nativ-bench) 0.951 🟢 0.566 0.848 0.890
request vs policy (nativ-bench) 0.919 🟢 0.298 0.406 0.624
topic of a sentence, 7 options (SIB-200) 0.755 0.583 0.789 0.858
does one sentence follow from another (HebNLI) 0.473 0.453 0.606 † 0.582
positive / negative / neutral (HebrewSentiment) 0.648 0.429 0.830 † 0.611
reading comprehension, 4 options (Belebele) 0.551 0.318 0.758 0.883

Calibration error (lower is better)

How far the model's confidence is from how often it is actually right, from 0 to 1. When a well-calibrated model says 0.9, it is right about 90% of the time; an error of 0.10 means its confidence is off by about 10 points on average. Low error means the probabilities can be trusted as thresholds, for example "send to a person below 0.8".

test (dataset) Nativ Laya multilingual Laya-Hebrew DeepSeek V4 Flash
what does the user want (MASSIVE) .010 🟢 .085 .028 .050
does the paragraph answer the question (HeQ) .022 🟢 .338 .061 .096
sentiment of a comment (OnlpLab) .011 🟢 .062 .064 .132
personal information in a message (nativ-bench) .069 .356 .028 .069
request vs policy (nativ-bench) .065 🟢 .373 .365 .311
topic of a sentence (SIB-200) .073 🟢 .112 .096 .109
does one sentence follow from another (HebNLI) .200 .213 .104 .291
positive / negative / neutral (HebrewSentiment) .104 .146 .062 .282
reading comprehension (Belebele) .124 .229 .026 .086

The models compared:

  • Laya multilingual: 322M, Apache 2.0. Its training data is not published.
  • Laya-Hebrew: 378M, CC BY-NC-SA (non-commercial). Trained on HebNLI, HeQ, Hebrew sentiment data, GoEmotions, CLINC150, Banking77, RACE, BoolQ and other datasets.
  • DeepSeek V4 Flash: a large LLM, zero-shot through an API, scored by the probability of each option's letter. A reference point, not a model that runs on a CPU.

Two MIT-licensed zero-shot classifiers were also tested, mDeBERTa-v3-xnli and bge-m3-zeroshot-v2.0-c, and scored below Nativ on every test they ran (for example, intent: 0.542 and 0.754).

Speed (lower is better), one decision at a time, fp32 on an Apple M2 Pro CPU: Nativ 70 ms for a short message, 130 ms for a paragraph · Laya-Hebrew 77 / 160 ms.

What to use it for

use question options
route a support message לאיזו מחלקה להעביר את הפנייה? משלוחים / החזרות / חיובים / תמיכה טכנית
catch personal details before logging האם ההודעה מכילה מידע מזהה אישי? כן / לא
check a request against a policy האם הבקשה עומדת במדיניות? עומדת / לא עומדת / חסר מידע
check a retrieved passage (RAG) האם הקטע עונה על השאלה? כן / לא
tag a topic or an intent מה הנושא של הטקסט? your own list
sentiment of reviews and comments מה הסנטימנט של התגובה? חיובית / שלילית / לא קשורה

A low top probability is a useful signal to hand the case to a person.

Training

The encoder reads [CLS] state [SEP] question [SEP] option1 [SEP] option2 [SEP] ... in one pass; each option is scored from the [CLS] vector and the mean of its tokens, and a softmax gives the probabilities.

Nativ was distilled (Hinton et al., 2015) from a fine-tuned DictaLM-3.0-1.7B-Instruct (Apache-2.0), from DeepSeek V4 Flash (MIT) for intent, emotion and 3-way sentiment, and from DictaLM-3.0-24B-Thinking for reading. Part of the data was generated, translated or checked by DictaLM-3.0-24B-Thinking and DeepSeek V4 Flash. During training every item's question and options are reworded at random, so the model learns to read the options.

task data license
routing MASSIVE he-IL train (11.5k), Hebrew descriptions of the 60 intents, some labels corrected by hand CC BY 4.0
sentiment OnlpLab Hebrew-Sentiment-Data, deduplicated (5.9k) MIT
grounded HeQ v1.1 train, answerable and unanswerable questions, balanced (15k) CC BY 4.0
personal information ~5k template messages rewritten by the 24B, ~2k messages written by the 24B generated
policy ~8k rule / request pairs, 16 rule types, labels computed in code, rewritten by the 24B generated
reading CosmosQA translated to Hebrew (24.8k), questions on HeQ paragraphs (5k) CC BY 4.0 / generated
intent CLINC150 (11.9k) and Banking77 (10k), translated to Hebrew; 2 to 7 options per item from 228 intents CC BY 3.0 / CC BY 4.0
emotion GoEmotions single-label comments, translated to Hebrew (5.8k); 2 to 7 options from 28 emotions Apache 2.0
sentiment with neutral 12k texts from the sources above, labelled positive / negative / neutral by DeepSeek V4 Flash as the sources

Generated and translated data were filtered for answer shortcuts, changed facts and broken translations. The Hebrew translations of CLINC150, Banking77 and GoEmotions are published as hebrew-intent-emotion.

Limitations

  • Reading comprehension of long passages is the weakest task.
  • Whether something counts as personal information depends on the application (order numbers, for example). Adjust the threshold or the options.
  • Hebrew only. Not meant as the only safeguard for decisions about people.

License

CC BY 4.0, as dicta-il/neodictabert. Training data: MASSIVE (Amazon, CC BY 4.0), HeQ (NNLP-IL, CC BY 4.0), OnlpLab Hebrew-Sentiment-Data (MIT), CosmosQA (AllenAI, CC BY 4.0), CLINC150 (CC BY 3.0), Banking77 (PolyAI, CC BY 4.0), GoEmotions (Google, Apache 2.0), and data generated with DictaLM-3.0 models (Dicta, Apache-2.0) and DeepSeek V4 Flash (MIT). Belebele (Meta, CC BY-SA 4.0), SIB-200, HebrewSentiment and HebNLI were used for evaluation only.

Downloads last month
104
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yoavipo/nativ-he-decision

Finetuned
(5)
this model

Datasets used to train yoavipo/nativ-he-decision