Nativ (נתיב)
The best Hebrew decision model for commercial use (as of October 2026). It beats Laya multilingual, the leading commercially licensed alternative, on every test below, and runs on a laptop CPU.
Give it any Hebrew text, a question and the options, and it returns a probability for every option in a single forward pass. No generation, no prompt parsing, no GPU.
נתיב (nativ) is Hebrew for "path": the model chooses which way a request should go.
| Size | 364M parameters (dicta-il/neodictabert + an option scorer) |
| Speed | ~70 ms per decision on a laptop CPU |
| Output | a probability per option, with its calibration error reported for every task |
| Options | any Hebrew text, 2 to 7 per question |
| License | CC BY 4.0, commercial use allowed |
| Benchmark | nativ-bench |
| Training data | hebrew-intent-emotion (translated part) |
Usage
decider.py is included in this repo (requires torch, transformers, safetensors).
from decider import HebrewDecider
nativ = HebrewDecider.from_pretrained("path/to/nativ-he-decision")
nativ.decide(
state="שלום, ההזמנה שלי הייתה אמורה להגיע ביום שלישי ועדיין לא קיבלתי אותה. אפשר לבדוק מה קורה?",
question="לאיזו מחלקה להעביר את הפנייה?",
options=["משלוחים", "החזרות והחלפות", "חיובים ותשלומים", "תמיכה טכנית"],
)
# one probability per option, in the order given
- Options are free Hebrew text. Short descriptions work better than label names ("לקבוע שעון מעורר" rather than "alarm_set").
decide_manytakes a list of{"state", "question", "options"}dicts.- The input is limited to 1,024 tokens. Long states are truncated; the question and options are kept.
Results
The two nativ-bench tests (nativ-bench) use scenarios, rules
and numbers that do not appear in Nativ's training data. The last four tests are public Hebrew datasets Nativ never
trained on.
- † - marks a model that trained on other examples from the same dataset.
- 🟢 - marks the tests where Nativ scores highest.
Accuracy (higher is better)
The share of questions answered correctly, from 0 to 1: 0.95 means 95 of every 100 answers are right.
| test (dataset) | Nativ | Laya multilingual | Laya-Hebrew | DeepSeek V4 Flash |
|---|---|---|---|---|
| what does the user want, 4 options (MASSIVE) | 0.962 † 🟢 | 0.644 | 0.887 | 0.905 |
| does the paragraph answer the question (HeQ) | 0.852 † 🟢 | 0.582 | 0.694 † | 0.847 |
| sentiment of a comment (OnlpLab) | 0.905 † 🟢 | 0.680 | 0.822 † | 0.838 |
| personal information in a message (nativ-bench) | 0.951 🟢 | 0.566 | 0.848 | 0.890 |
| request vs policy (nativ-bench) | 0.919 🟢 | 0.298 | 0.406 | 0.624 |
| topic of a sentence, 7 options (SIB-200) | 0.755 | 0.583 | 0.789 | 0.858 |
| does one sentence follow from another (HebNLI) | 0.473 | 0.453 | 0.606 † | 0.582 |
| positive / negative / neutral (HebrewSentiment) | 0.648 | 0.429 | 0.830 † | 0.611 |
| reading comprehension, 4 options (Belebele) | 0.551 | 0.318 | 0.758 | 0.883 |
Calibration error (lower is better)
How far the model's confidence is from how often it is actually right, from 0 to 1. When a well-calibrated model says 0.9, it is right about 90% of the time; an error of 0.10 means its confidence is off by about 10 points on average. Low error means the probabilities can be trusted as thresholds, for example "send to a person below 0.8".
| test (dataset) | Nativ | Laya multilingual | Laya-Hebrew | DeepSeek V4 Flash |
|---|---|---|---|---|
| what does the user want (MASSIVE) | .010 🟢 | .085 | .028 | .050 |
| does the paragraph answer the question (HeQ) | .022 🟢 | .338 | .061 | .096 |
| sentiment of a comment (OnlpLab) | .011 🟢 | .062 | .064 | .132 |
| personal information in a message (nativ-bench) | .069 | .356 | .028 | .069 |
| request vs policy (nativ-bench) | .065 🟢 | .373 | .365 | .311 |
| topic of a sentence (SIB-200) | .073 🟢 | .112 | .096 | .109 |
| does one sentence follow from another (HebNLI) | .200 | .213 | .104 | .291 |
| positive / negative / neutral (HebrewSentiment) | .104 | .146 | .062 | .282 |
| reading comprehension (Belebele) | .124 | .229 | .026 | .086 |
The models compared:
- Laya multilingual: 322M, Apache 2.0. Its training data is not published.
- Laya-Hebrew: 378M, CC BY-NC-SA (non-commercial). Trained on HebNLI, HeQ, Hebrew sentiment data, GoEmotions, CLINC150, Banking77, RACE, BoolQ and other datasets.
- DeepSeek V4 Flash: a large LLM, zero-shot through an API, scored by the probability of each option's letter. A reference point, not a model that runs on a CPU.
Two MIT-licensed zero-shot classifiers were also tested, mDeBERTa-v3-xnli and bge-m3-zeroshot-v2.0-c, and scored below Nativ on every test they ran (for example, intent: 0.542 and 0.754).
Speed (lower is better), one decision at a time, fp32 on an Apple M2 Pro CPU: Nativ 70 ms for a short message, 130 ms for a paragraph · Laya-Hebrew 77 / 160 ms.
What to use it for
| use | question | options |
|---|---|---|
| route a support message | לאיזו מחלקה להעביר את הפנייה? | משלוחים / החזרות / חיובים / תמיכה טכנית |
| catch personal details before logging | האם ההודעה מכילה מידע מזהה אישי? | כן / לא |
| check a request against a policy | האם הבקשה עומדת במדיניות? | עומדת / לא עומדת / חסר מידע |
| check a retrieved passage (RAG) | האם הקטע עונה על השאלה? | כן / לא |
| tag a topic or an intent | מה הנושא של הטקסט? | your own list |
| sentiment of reviews and comments | מה הסנטימנט של התגובה? | חיובית / שלילית / לא קשורה |
A low top probability is a useful signal to hand the case to a person.
Training
The encoder reads [CLS] state [SEP] question [SEP] option1 [SEP] option2 [SEP] ... in one pass; each option is
scored from the [CLS] vector and the mean of its tokens, and a softmax gives the probabilities.
Nativ was distilled (Hinton et al., 2015) from a fine-tuned DictaLM-3.0-1.7B-Instruct (Apache-2.0), from DeepSeek V4 Flash (MIT) for intent, emotion and 3-way sentiment, and from DictaLM-3.0-24B-Thinking for reading. Part of the data was generated, translated or checked by DictaLM-3.0-24B-Thinking and DeepSeek V4 Flash. During training every item's question and options are reworded at random, so the model learns to read the options.
| task | data | license |
|---|---|---|
| routing | MASSIVE he-IL train (11.5k), Hebrew descriptions of the 60 intents, some labels corrected by hand | CC BY 4.0 |
| sentiment | OnlpLab Hebrew-Sentiment-Data, deduplicated (5.9k) | MIT |
| grounded | HeQ v1.1 train, answerable and unanswerable questions, balanced (15k) | CC BY 4.0 |
| personal information | ~5k template messages rewritten by the 24B, ~2k messages written by the 24B | generated |
| policy | ~8k rule / request pairs, 16 rule types, labels computed in code, rewritten by the 24B | generated |
| reading | CosmosQA translated to Hebrew (24.8k), questions on HeQ paragraphs (5k) | CC BY 4.0 / generated |
| intent | CLINC150 (11.9k) and Banking77 (10k), translated to Hebrew; 2 to 7 options per item from 228 intents | CC BY 3.0 / CC BY 4.0 |
| emotion | GoEmotions single-label comments, translated to Hebrew (5.8k); 2 to 7 options from 28 emotions | Apache 2.0 |
| sentiment with neutral | 12k texts from the sources above, labelled positive / negative / neutral by DeepSeek V4 Flash | as the sources |
Generated and translated data were filtered for answer shortcuts, changed facts and broken translations. The Hebrew
translations of CLINC150, Banking77 and GoEmotions are published as
hebrew-intent-emotion.
Limitations
- Reading comprehension of long passages is the weakest task.
- Whether something counts as personal information depends on the application (order numbers, for example). Adjust the threshold or the options.
- Hebrew only. Not meant as the only safeguard for decisions about people.
License
CC BY 4.0, as dicta-il/neodictabert. Training data: MASSIVE
(Amazon, CC BY 4.0), HeQ (NNLP-IL, CC BY 4.0), OnlpLab Hebrew-Sentiment-Data (MIT), CosmosQA (AllenAI, CC BY 4.0),
CLINC150 (CC BY 3.0), Banking77 (PolyAI, CC BY 4.0), GoEmotions (Google, Apache 2.0), and data generated with
DictaLM-3.0 models (Dicta, Apache-2.0) and DeepSeek V4 Flash (MIT). Belebele (Meta, CC BY-SA 4.0), SIB-200, HebrewSentiment
and HebNLI were used for evaluation only.
- Downloads last month
- 104
Model tree for yoavipo/nativ-he-decision
Base model
dicta-il/neodictabert