Kev-0.8B for LiteRT

jaredpalmer/kev-0.8b is a decision model by Jared Palmer. It reads a text, called the state, and typed questions about it: yes/no (noul), multiple choice (choice) and rating (score). For each question it returns an answer with probabilities. It never generates text. A pointer head of two linear layers reads the hidden states of Qwen/Qwen3.5-0.8B-Base, adapted with a rank-16 LoRA, and turns them into those probabilities. Requests and responses follow the /v1/systemone shape of the author's server (github.com/jaredpalmer/kev).

This repository holds the backbone as three LiteRT graphs, for rows of up to 512, 1,024 and 2,048 tokens, with the LoRA folded into the weights. It also holds the pointer head, the tokenizer files, a Python host, test fixtures and the conversion scripts. The graphs return hidden states. The host builds the input rows and applies the head.

The reference for every check is the author's fp32 PyTorch code. The shipped 512-token file ran on the Apple M4 Max CPU (8 threads) and Metal GPU (float32 precision) with ai-edge-litert 2.2.0, and on the Galaxy S26 GPU (FP32 precision) with LiteRT 2.2.0, on 392 test questions. A near tie is a question whose two most likely options in the reference are 0.02 or less apart; 15 of the 392 are. On the other 377, the file gave the reference's most likely option every time, on all three. Of the 15 near ties, 13 kept the reference's answer. No option's probability, near ties included, differed from the reference by more than 0.0104.

An invented support ticket with three questions, the state abbreviated:

{
  "state": "Ticket #48213, opened by Mara Quellen.\n\nHi, I ordered the Thistlebeam desk lamp (order TB-20931) on September 14 and was charged twice on my card … Could you refund the duplicate charge? I need it sorted before my card statement closes on Friday. …",
  "questions": {
    "team": {"type": "choice", "instructions": "Which team should handle this ticket?",
             "criteria": {"billing": "Charges, refunds and invoices", "shipping": "Deliveries, tracking and lost parcels",
                          "returns": "Exchanges and sending a product back", "technical": "Product faults and setup help"}},
    "deadline": {"type": "noul", "instructions": "Does the customer ask for action by a specific deadline?",
                 "criteria": {"true": "The ticket names a day or date by which something must happen",
                              "false": "No deadline is stated"}},
    "mood": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["Calm", "Annoyed", "Angry"]}
  }
}

On desktop CPU, the L512 file returned this response. It is examples/run_example.expected.json, which leaves out latency_ms.

{
  "model": "kev-latest",
  "answers": {
    "team": {"type": "choice", "choice": "billing", "confidence": 0.7618,
             "probabilities": {"billing": 0.8213, "shipping": 0.014, "returns": 0.0883, "technical": 0.0764}},
    "deadline": {"type": "noul", "noul": 0.4461},
    "mood": {"type": "score", "score": 0.8542, "legend": {"0": "Calm", "1": "Annoyed", "2": "Angry"},
             "probabilities": {"0": 0.3672, "1": 0.4114, "2": 0.2214}, "confidence": 0.1172}
  },
  "usage": {"input_tokens": 216, "output_tokens": 189}
}

Other formats:

We have not checked their contents.

Files

File Bytes Role
kev-0.8b_rowprefill_L512_fp16fc_i8emb.tflite 1,264,068,368 Graph for rows of up to 512 tokens. The example uses it.
kev-0.8b_rowprefill_L1024_fp16fc_i8emb.tflite 1,269,023,216 Graph for rows of up to 1,024 tokens.
kev-0.8b_rowprefill_L2048_fp16fc_i8emb.tflite 1,285,227,888 Graph for rows of up to 2,048 tokens.
head/kev_0.8b_pointer_head.safetensors 2,099,632 Pointer head: q.weight [256, 1024], q.bias, k.weight, k.bias, float32.
head/kev_0.8b_pointer_head.json 2,105 Temperature, delimiter and pad token ids, the head formula and the graph contract.
tokenizer/tokenizer.json 19,989,325 The source repository's tokenizer.json, unchanged.
tokenizer/tokenizer_config.json 1,128 The source repository's tokenizer_config.json, unchanged.
host/kev_litert.py 32,091 Python host: request to rows, graph, pointer head and response. It also runs from the command line.
host/requirements-host.txt 393 Pinned packages for the Python host.
examples/ run_example.py, the example above, and run_example.expected.json, its expected response.
android/CardSnippet.kt 2,440 The Kotlin block below.
fixtures/ Test requests (156 with their text, 221 by reference) and the script that rebuilds the full set; the reference's results for all 402 questions; 12 tokenizer probes; the SemIf license.
conversion/ Conversion and check scripts, their environment files, and a README with the commands in order.
REPRODUCE.md 9,913 How to reproduce the files and the checks.
LICENSE 11,358 Apache License 2.0 text.
NOTICE 2,013 Attribution.
SHA256SUMS SHA-256 checksums of the files.

The three graphs hold the same weights and differ only in row length. A row is the state plus one question. Use the smallest graph that holds the row. The Python host picks it for each question and refuses a row longer than 2,048 tokens instead of cutting it.

Each graph holds the token embedding as an int8 table, the 24 layers of the backbone and the final RMSNorm. It has no LM head and returns the hidden state at every position. The 186 FULLY_CONNECTED weights are stored in float16, each behind a DEQUANTIZE operator (ai-edge-quantizer float casting). The embedding table, [248,320 × 1,024], is int8 with one scale per row. Activations, the 18 causal convolutions of the Gated DeltaNet layers and the delta rule stay in float32.

Minimal usage

Python: desktop CPU or GPU

The host needs tokenizers, numpy, safetensors and ai-edge-litert (2.2.0), pinned with their dependencies in host/requirements-host.txt. It does not need PyTorch or the kev package.

hf download litert-community/Kev-0.8B-LiteRT --local-dir Kev-0.8B-LiteRT
cd Kev-0.8B-LiteRT
pip install -r host/requirements-host.txt
python examples/run_example.py --check

examples/run_example.py sends the ticket request above on the CPU and prints the response. --check also compares it with examples/run_example.expected.json. The example needs only the L512 graph: add --exclude "*L1024*" "*L2048*" to the download to skip the other two.

In your own code:

import sys
sys.path.insert(0, "host")  # run from the repository root
from kev_litert import KevLiteRT

request = {"state": "Order #1182 arrived with a cracked screen.",
           "questions": {"refund": {"type": "noul", "instructions": "Should we offer a refund?"}}}

with KevLiteRT.from_dir(".") as kev:
    print(kev.decide(request))

from_dir finds the graphs that are present, the head and the tokenizer in this repository's layout. It compiles a graph only when a row needs it, and with the L512 file alone it serves rows of up to 512 tokens. decide() returns the response: model, answers, usage and latency_ms. The default is the CPU with 4 threads; threads= changes the count. accelerator="gpu" runs the graphs on the GPU with float32 precision (GpuOptions(enforce_f32=True)). A row that fits no graph raises RowTooLong, and a non-finite hidden state raises NonFiniteOutput.

From the command line:

python host/kev_litert.py --graph kev-0.8b_rowprefill_L512_fp16fc_i8emb.tflite \
    --head head/kev_0.8b_pointer_head.safetensors --tokenizer tokenizer/tokenizer.json --request request.json

--graph takes one or more of the three files. With only the L512 file, a row longer than 512 tokens is refused. --accel gpu selects the GPU at float32 precision, and --request - reads the request from stdin.

Kotlin: Android GPU with explicit FP32

// SPDX-License-Identifier: Apache-2.0
package com.kev.snippet

import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import com.google.ai.edge.litert.Environment
import java.io.File

/**
 * One Kev row-prefill graph (`*_rowprefill_L{512,1024,2048}_fp16fc_i8emb.tflite`) on the GPU with explicit FP32:
 * ids int32 [1, L] + valid float32 [1, L] -> hidden float32 [1, L, d] (d = 1024 for Kev-0.8B, 2560 for Kev-4B).
 * The default GPU precision (float16 activations) gives non-finite hidden states on some rows.
 */
class KevRowGraph(file: File, private val length: Int, env: Environment) : AutoCloseable {
  private val options = CompiledModel.Options(Accelerator.GPU).apply {
    gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
  }
  private val model = CompiledModel.create(file.absolutePath, options, env)
  private val inputs = listOf("ids", "valid").associateWith { model.createInputBuffer(it, "serving_default") }
  private val outputs = mapOf("hidden" to model.createOutputBuffer("hidden", "serving_default"))

  /**
   * row = [248060] + state + [248061] + instructions + for each option ([248049] + option + [248050]) + [248062], with
   * the Kev repository's tokenizer.json and no special tokens added; optionEnds = the index of each option's 248050.
   * Returns the hidden rows the pointer head reads: [decide (the last token), each option's 248050], d floats each.
   */
  fun readout(row: IntArray, optionEnds: IntArray): List<FloatArray> {
    require(row.size <= length && row.last() == DECIDE) { "a row ends with 248062 and fits $length tokens" }
    require(optionEnds.all { row[it] == OPTION_END }) { "optionEnds must point at 248050 tokens" }
    inputs.getValue("ids").writeInt(IntArray(length) { if (it < row.size) row[it] else PAD })
    inputs.getValue("valid").writeFloat(FloatArray(length) { if (it < row.size) 1f else 0f })
    model.run(inputs, outputs, "serving_default")
    val hidden = outputs.getValue("hidden").readFloat()
    val d = hidden.size / length
    return (listOf(row.size - 1) + optionEnds.toList()).map { hidden.copyOfRange(it * d, (it + 1) * d) }
  }

  override fun close() {
    (inputs.values + outputs.values).forEach { it.close() }
    model.close()
  }

  companion object {
    const val PAD = 248044 // <|endoftext|>, right padding
    const val OPTION_END = 248050
    const val DECIDE = 248062
  }
}

Before calling readout, the app builds row and optionEnds: it tokenizes the state, the instructions and each option with this repository's tokenizer.json (no special tokens added, <|name|> rewritten to <¦name¦>) and joins them with the delimiter ids of the host contract below. Afterwards, the app applies the pointer head in float32 to the returned vectors, with the weights in head/kev_0.8b_pointer_head.safetensors and the temperature in the JSON beside it, and takes the softmax.

Host contract

host/kev_litert.py implements this contract. It ports the request handling of the kev package at tag kev-1.0.

  1. Render the request. A JSON state becomes key: value lines. The options become text: no[: description] and yes[: description] for noul, name[: description] for choice, and the level text for score.
  2. Tokenize each text with tokenizer/tokenizer.json and add_special_tokens=False. Rewrite <|name|> in caller text to <¦name¦> before tokenizing, so that caller text cannot produce a delimiter.
  3. Build one row per question: [248060] + state + [248061] + instructions, then [248049] + option + [248050] for each option, then [248062]. The five delimiters are state (248060), question (248061), option start (248049), option end (248050) and decide (248062).
  4. Take a graph whose length L holds the row. Pad ids on the right with 248044. valid is 1.0 on real tokens and 0.0 on padding.
  5. Run the serving_default signature. The inputs are ids int32 [1, L] and valid float32 [1, L]. The output is hidden float32 [1, L, 1024], the hidden states after the final RMSNorm at every position. Positions are the constants 0 to L−1. The graph has no input or output for model state, so each row starts from zero.
  6. Read out in float32. The decide vector is the hidden state of the row's last real token (248062). Option k's vector is the hidden state of its closing 248050. Then z_k = ((h_opt_k · Wkᵀ + bk) · (h_decide · Wqᵀ + bq)) / 16 / T, with T = 2.3510958125672174, and p = softmax(z).
  7. Answer. Noul returns noul = p(yes). Choice returns choice (the most likely option), probabilities and confidence = (p_max − 1/K) / (1 − 1/K) for K options. Score returns score = Σ i·p_i over the levels counted from 0, legend, probabilities and confidence = max(0, 1 − E|level − mode| / D), where D is the mean absolute deviation of a uniform distribution over the levels. Numbers are rounded to 4 decimals, as the author's round_prob does.

Use this repository's tokenizer.json, which is the source repository's file, and read it with tokenizers.Tokenizer.from_file. It stores the tokenizer pipeline that transformers builds for the author's code (Qwen2Tokenizer). The base repository's own tokenizer.json, read the same way, gives different ids for text with combining marks, such as Devanagari, and for a few special-looking strings such as <think>. It differs on 4 of the 12 probes in fixtures/tokenizer_probes.json. The host refuses it.

On all 402 test questions, the host's token ids equal those of the author's encode. The tokenizer gives the expected ids for all 12 probes.

On the GPU, set float32 precision explicitly: GpuOptions(enforce_f32=True) in Python, CompiledModel.GpuOptions(precision = FP32) in Kotlin. The default precision gives NaN on some rows (see fp16 activations below).

Measured agreement and speed

Agreement with the author's fp32 code

The reference is the author's code at tag kev-1.0 on the CPU in float32, with the author's lock file (torch 2.8.0, transformers 5.17.0, peft 0.21.0): Checkpoint.load(cpu, dtype=float32), then forward in the row form. Probabilities are compared after the temperature.

The test set has 377 requests with 402 questions. The author's evals/v4/transfer-v4/development.jsonl gives 220 of them: records 1 to 60 (MMLU, 4 options), 20 from each of its 7 other sources and 20 score questions. SemIf authored144 gives 144 items, mapped to 3-option choice questions. The 12 invented records, written for this conversion, hold 37 questions of all three types. One more request is a control, described below. Without the control there are 401 questions: 280 choice, 93 noul and 28 score. The rows of 392 questions have at most 366 tokens. The other 9 questions sit on 3 long invented states, with rows of 1,486 to 1,805 tokens. Only the L2048 file holds them.

"Same most likely option" counts the questions that are not near ties. The 15 near-tie questions all have rows of at most 512 tokens. Max |Δp| is the largest absolute difference of any option's probability, and mean |Δp| is the mean over all options. The tolerance set before the runs was: the same most likely option on every question that is not a near tie, max |Δp| of at most 0.02 and mean |Δp| of at most 0.002.

File Runtime Questions Same most likely option Near ties Max |Δp| Mean |Δp|
fp32 graph, L2048 (not shipped; a conversion check) Mac CPU 401 386/386 15/15 4.4e-6 5.0e-7
L512 (shipped) Mac CPU, 8 threads 392 377/377 13/15 0.0104 9.1e-4
L512 (shipped) Mac GPU (Metal), float32 precision 392 377/377 13/15 0.0104 9.1e-4
L1024 (shipped) Mac GPU (Metal), float32 precision 392 377/377 13/15 0.0104 9.1e-4
L2048 (shipped) Mac GPU (Metal), float32 precision 401 386/386 (long rows 9/9) 13/15 0.0104 8.9e-4
L512 (shipped) Galaxy S26 GPU (OpenCL), FP32 precision 392 377/377 13/15 0.0104 9.1e-4
L1024 (shipped) Galaxy S26 GPU (OpenCL), FP32 precision 243 234/234 8/9 0.0104 8.9e-4
L2048 (shipped) Galaxy S26 GPU (OpenCL), FP32 precision 29 27/27 (long rows 9/9) 2/2 0.0104 8.3e-4
L512 (shipped) Galaxy S26 CPU, 4 threads 127 113/113 12/14 0.0104 9.5e-4

The S26 L1024 run stopped at a 10-minute limit after 243 of the 392 questions; the phone had warmed up and slowed down. The S26 L2048 run covered the 9 long questions and the 20 opening transfer-v4 questions. The S26 CPU run covered 127 questions, 128 rows with the control.

The two near-tie questions that changed have reference top-two gaps of 0.0020 and 0.00008. They are the same two on every runtime.

The int8 embedding table accounts for the difference from the reference. A file that keeps the FULLY_CONNECTED weights in float32 and makes only the table int8 also reaches a max |Δp| of 0.0103. On the same file, the Mac GPU and CPU agree within a max |Δp| of 4.8e-6. The S26 GPU agrees with the Mac CPU within 4.3e-6.

A control request with one word of its instructions changed ("correctly" to "incorrectly") differs from the reference by a max |Δp| of 0.048, outside the 0.02 tolerance, and every runtime detected the change.

A variant with dynamic-int8 FULLY_CONNECTED weights (0.78 GB) moved probabilities by up to 0.070 on the Mac GPU and 0.146 on the CPU, so it is not shipped.

Desktop (Apple M4 Max)

These times come from an Apple M4 Max (macOS 27.0) with ai-edge-litert 2.2.0 through the Python CompiledModel API, measured while no other GPU job ran. Other processes were running on the CPU (it was 80 to 89% idle at the start of each session). A call runs one question's row. A time is the wall clock of writing the inputs, running and reading the output back: the median of 20 calls after 5 warm-up calls.

Row File GPU (Metal), float32 precision CPU, 8 threads
About 130 tokens (one question of a 5-question request; rows of 128 to 142 tokens) L512 139.9 ms 585.2 ms
The 5-question request (sum of 5 calls) L512 699.9 ms 2,922.9 ms
300 tokens L512 140.0 ms 644.1 ms
1,000 tokens L1024 240.1 ms 1,149.4 ms
1,000 tokens L2048 473.9 ms 2,194.9 ms
1,805 tokens (the longest test row) L2048 474.2 ms 2,185.3 ms

Each graph computes all of its positions whatever the row length, so 130-token and 300-token rows take nearly the same time on L512, and 1,000-token and 1,805-token rows on L2048. The L2048 rows were measured in a second session in the same way. The initial GPU compile took 10.5 s for L512, 12.9 s for L1024 and 19.8 s for L2048. With the GPU at float32 precision, the process memory footprint was 3.8 to 3.9 GB. On Metal, is_fully_accelerated is true for all three files.

Galaxy S26

These times come from one Galaxy S26 (SM-S942Q, Android 16, 12 GB) with LiteRT 2.2.0 through the Kotlin CompiledModel API, on 2026-10-03. Calls and medians are defined as on the desktop. The 5-question request was timed as 20 requests (100 calls). Each timing run compiled the graph in a new process. It started at thermal status 0 (battery 28.5 to 37.8 °C) with no thermal cap on the GPU or CPU clocks, 181 to 241 s after the previous run. At launch the prime CPU cluster showed the limit that follows waking the screen (4.19 of 4.74 GHz); the timed calls began 9 to 14 s later. The phone was on USB with the screen on and the app in the foreground. By the end of each run, the thermal status had risen to 2 (3 after the L2048 run).

Row File GPU (OpenCL), FP32 precision CPU, 4 threads
About 130 tokens (one question of the 5-question request) L512 630 ms (min 623, max 668) 1,838 ms (min 1,204)
The 5-question request (sum of 5 calls) L512 3,150 ms (3,126 to 3,235) 9,207 ms (6,854 to 9,364)
300 tokens L512 663 ms (641 to 669) 1,843 ms
1,000 tokens L1024 1,345 ms (1,294 to 1,417) not measured
One call at any row length (n = 25) L2048 3,331 ms not measured

On the CPU, the 5 warm-up calls took 978 to 1,106 ms on the cool phone, and the shortest timed call took 1,204 ms. Within about 10 s of load the CPU clocks fell to 1.5 to 1.8 GHz, and the calls settled near 1.8 s. The L2048 graph computes all 2,048 positions whatever the row length. Its time is the median of 25 calls on different rows (rows 6 to 30 of the run). The 9 long rows among them (1,486 to 1,805 tokens) alone give 3,441 ms.

Cold compile with the GPU at FP32 took 13.9 s for L512, 17.2 s for L1024 and 23.4 s for L2048. On the CPU with 4 threads, L512 compiled in 3.3 s.

Every graph ran whole on the GPU delegate in one partition. Logcat shows Replacing 21059 out of 21059 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for L512, 25667 of 25667 nodes for L1024 and 34883 of 34883 for L2048.

On the GPU at FP32, the L2048 file gave the reference's most likely option on 29 of 29 questions: the 9 long ones and 20 transfer-v4 questions. Its max |Δp| was 0.0104, and the control moved by 0.048.

Continuous runs warm the phone and slow it. During the agreement runs on the same phone and files, at thermal status 2 to 3 with GPU and CPU clocks capped, one call took 1.06 s on L512 and up to 2.6 s on L1024.

On the GPU, run() returns at once, and the work finishes when the output is read. Time the read as well; the times above include it. If the screen locks, the app moves to the background CPU set and the computation stops. Keep the app in the foreground with the screen on.

fp16 activations

Run the GPU at float32 precision. At the default GPU precision (float16 activations), 18 of the 393 test rows gave non-finite (NaN) hidden states on the Mac (Metal, L1024 file). The Galaxy S26 GPU (OpenCL, L512 file) gave NaN on the same 18 rows. The Python host raises NonFiniteOutput instead of answering from such a row. No form of these graphs has been verified to run correctly with fp16 activations.

Limits

  • A row holds at most 2,048 tokens: the state plus one question. Longer states are not handled. The author's server accepts states of up to 65,536 tokens.
  • The state is computed again for every question. No graph here computes it once for reuse, so a request with Q questions takes Q calls.
  • The GPU needs float32 precision. The default precision gives NaN on some rows, and no form that runs with fp16 activations has been verified.
  • Probabilities differ from the reference by up to 0.0104, because of the int8 embedding table. A question whose two most likely options are 0.002 or less apart can change its answer: 2 of 401 did.
  • Kev-4B is in a separate repository, litert-community/Kev-4B-LiteRT; its 7.8 GB file did not fit the 12 GB Galaxy S26, where the low-memory killer stopped the process during compile (one try).
  • The agreement numbers measure how closely the conversion follows the author's fp32 code, not task accuracy. This conversion did not measure accuracy or calibration again; see the source model card. The temperature is the author's fitted value, unchanged.
  • Input languages and intended uses follow the source model card.
  • Measured on one Mac (Apple M4 Max) and one Galaxy S26.

Provenance, conversion and license

  • Source: jaredpalmer/kev-0.8b at tag v1.0 (commit bf75a6a8848ea6960ff2ed108d9ed44c2941174f) by Jared Palmer, Apache-2.0. Base: Qwen/Qwen3.5-0.8B-Base at revision dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68, Apache-2.0.
  • Model: the base with a rank-16 LoRA on 12 projection types and a pointer head of two linear layers with 256 dimensions. The base has 24 layers (18 Gated DeltaNet layers, a form of linear attention, and 6 full-attention layers), hidden size 1024 and a vocabulary of 248,320.
  • Weights: the author's scripts/merge_lora_checkpoint.py at tag kev-1.0 folded the LoRA into 186 weights in float32. Read with the author's loader, the folded checkpoint gives bit-identical probabilities to the adapter checkpoint on all 402 questions.
  • Graph: the Qwen3.5 text model of transformers 5.14.1, re-authored for export. The Gated DeltaNet chunk kernel is rewritten at rank 4 or less: the tail padding becomes a concat, and the diagonal and triangular masks become constants. Attention copies the GQA keys and values by concat. The hidden states before and after the rewrite differ by 3.2e-5. Exported with litert-torch 0.9.4 and quantized with ai-edge-quantizer 0.9.0.
  • Operators: 21,059 in L512, 25,667 in L1024 (the fp32 export's 25,481 plus 186 DEQUANTIZE) and 34,883 in L2048. None is CUSTOM, and no tensor is int64.
  • Scripts: conversion/ holds every conversion and check script, with its environment files and a README of the commands. REPRODUCE.md lists the sources, the environments and the steps.
  • Test data: fixtures/ holds the SemIf records (MIT, from github.com/TheoLeeCJ/SemIf at commit ca3ba65f) and the 12 invented records with their text. The 221 transfer-v4 records, the control included, are listed by reference only: file, tag, line, line SHA-256 and _meta.id. Their source datasets carry different licenses: on the Hub cards, tweet_eval is unknown and SciQ is CC BY-NC 3.0. fixtures/rebuild_requests.py restores them from the author's GitHub repository at tag kev-1.0.
  • Training data and evaluation: see the source model card.

License: Kev-0.8B and Qwen3.5-0.8B-Base are licensed under Apache 2.0. This repository is released under the same license (LICENSE), except the SemIf fixture records, which are MIT (fixtures/LICENSE-SemIf-MIT.txt). Attribution, including the code adapted from the kev package and from litert-torch, is in NOTICE.

Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for litert-community/Kev-0.8B-LiteRT

Finetuned
(2)
this model