cagliostro-v3.5

A 146M parameter decoder-only language model. It is cagliostro-v3's final checkpoint with one more short training stage: 0.39B tokens of v3's own cooldown data mixed with how-to and textbook text from Cosmopedia. Same architecture, same tokenizer, 75.39B tokens in total.

It scores 27.49 on the Open SLM Leaderboard metric, above every model listed there at the time of release.

Results

Zero-shot, measured with lm-evaluation-harness and the leaderboard's own ArithMark-3 script, on the float32 weights in this repository.

Benchmark Metric Score
HellaSwag acc_norm 43.41
ARC-Easy acc_norm 53.75
ARC-Challenge acc_norm 29.35
PIQA acc_norm 68.06
ArithMark-3 acc_norm 45.30
Open SLM Index 27.49

The Index is the leaderboard's own formula, (N(HellaSwag,25) + N(CombinedARC,25) + N(PIQA,50) + 0.65*N(ArithMark,25)) / 3.65 where N(v,c) = 100(v-c)/(100-c) and CombinedARC is the mean of ARC-Easy and ARC-Challenge.

For context, using the leaderboard's published figures for every other model:

Model Params Tokens Index
cagliostro-v3.5 146M 75.4B 27.49
SmolLM2-135M 135M 2T 27.13
cagliostro-v3 146M 75B 26.55
SmolLM-135M 135M 600B 25.74
GPT-X2.5-135M 135M 75B 25.17
Haidass1.5-143M 143M 400B 25.07

Index against the sub-150M field

Per-benchmark comparison

The lead over SmolLM2-135M is 0.36 Index on the leaderboard's figures. Re-evaluated in the same harness, SmolLM2-135M scores 27.01 and cagliostro-v3.5 is ahead by 0.48, but a paired bootstrap over benchmark items puts the 95% interval for that gap at [-0.82, +1.75]. On points it is first. On these five benchmarks the two models cannot be told apart with confidence, and we would not claim more than that.

cagliostro-v3 appears in the same-harness charts as it measures today, Index 26.31. Its leaderboard row lists 26.55.

What changed from v3

Field Value
Starting point cagliostro-v3 step 762,939, weights and AdamW state
Steps 4,000
Tokens per step 98,304
Tokens 0.39B
Learning rate warmup over 200 steps to 3e-4, then cosine to zero
Peak learning rate a tenth of v3's 3e-3
Optimizer AdamW, weight decay 0.01
Hardware one RTX PRO 6000 Blackwell
Wall clock about 41 minutes at 162,000 tokens per second
Source Share
Cosmopedia v1, WikiHow-style tutorials 30.0%
FineWeb-Edu (deduplicated) 20.35%
Cosmopedia v1, textbooks (OpenStax, Khan Academy, Stanford) 15.0%
Cosmopedia v2 13.75%
FineMath 3+ 8.25%
OpenMathInstruct-2 7.15%
SmolTalk 2.75%
DCLM-Baseline 2.75%

The bottom six rows are v3's cooldown mixture scaled to 55%. The WikiHow set is 174M tokens, so it was seen about 0.7 times. The textbook slice is the first 150M tokens of those three subsets.

What changed, benchmark by benchmark

The gain is mostly ArithMark-3 (+2.6 points) and HellaSwag (+0.9), with smaller gains on ARC-Challenge and PIQA. ARC-Easy moved the other way by about a point. That drop also shows up when v3 is trained the same way on its own cooldown mixture, so it comes from the extra low learning rate stage rather than from the new data.

The post-training run

How it was chosen

Six runs started from the same v3 checkpoint, identical except for the data. Before any of them finished we wrote down what a winner had to do: score above SmolLM2-135M's 27.13, beat SmolLM2-135M re-measured in the same harness, and beat v3 with the paired 95% interval above zero. Then a second run of the same recipe with a different data seed had to clear 27.13 as well.

All six post-training runs

Run Index
100% SmolTalk chat data 26.00
v3's own cooldown mixture 26.54
+ 30% WikiHow, seed 1 27.14
+ 30% WikiHow, seed 2 26.98
+ 30% WikiHow + 15% textbooks, seed 1 27.31
+ 30% WikiHow + 15% textbooks, seed 2 27.49

WikiHow alone passed once and then failed its second seed, so it was dropped. WikiHow with textbooks passed on both seeds. This repository holds the seed 2 checkpoint, the stronger of the two. The recipe averages 27.40 over both.

Chat data lost points, as it does for most base and instruct pairs on the leaderboard. The benchmarks score raw log-likelihood with no chat template, so text that looks like a conversation does not help.

Contamination checks

No benchmark training, validation or test split was used at any point.

Every new document went through the same 13-gram decontamination v3 used, against HellaSwag validation, ARC test, PIQA validation and ArithMark-3, dropping any document with a match. In a sample of 8,000 documents from the two Cosmopedia sources it dropped none.

The HellaSwag gain deserves a closer look, because about two thirds of HellaSwag is built from real WikiHow articles and Cosmopedia's WikiHow set was generated from real WikiHow titles.

HellaSwag gain by item source

19.3% of HellaSwag's WikiHow validation items share an article title with the Cosmopedia set. For the training split the figure is 18.5%, so the validation items are not singled out. The gain is about the same on items whose title appears in the training data (+1.45) as on items whose title does not (+1.24). ActivityNet items, which have nothing to do with WikiHow, barely moved (+0.09). That pattern is a model getting better at everyday how-to text, not one that has seen the test articles.

Sample outputs

Greedy decoding, float32, first sentence of each completion, unedited. These are all twelve prompts that were tried; the video shows four of them. The full completions are in samples.json.

Prompt Completion
To solve 47 + 68, first add the numbers together: \[ 47 + 68 = 115 \] Now, divide 115 by 2 to find the quotient: ...
A prime number is a positive integer greater than 1 that has no positive divisors other than 1 and itself.
Photosynthesis is the process by which plants convert sunlight into chemical energy.
How to boil an egg: First, you need to crack the egg into a bowl.
To tie a shoelace, start by wrapping the yarn around the hook of the shoelace.
The water cycle begins when rain falls onto the ground and turns into water droplets.
If a train travels 60 miles in 2 hours, its speed is 60 miles per hour.
Sam had 24 apples and gave away 9, so Sam now has 24 - 9 = 15 apples.
The capital of France is Paris.
Gravity is the force that pulls objects towards each other.
The area of a rectangle is found by multiplying its length by its width.
Plants need sunlight because they need to grow.

The misses are typical for this size. It adds 47 and 68 correctly and then divides by 2 for no reason, gets the train's speed wrong, starts boiling an egg by cracking it, and describes a shoelace as yarn.

Model details

Field Value
Parameters 146,352,000
Non-embedding parameters 85.7%
Layers 30
Hidden size 640
Intermediate size 1,536
Attention heads 10
Key/value heads 5
Attention Grouped query attention with cross-head subspace attenuation
Activation SwiGLU
Normalization RMSNorm, eps 1e-6
Positional encoding RoPE, theta 100,000
Context length 2,048
Vocabulary 32,768 BPE
Embeddings Tied input and output
Logit cap 15.0
Tokens seen 75.39B (75B pretraining, 0.39B post-training)
Weights float32 safetensors

The architecture is defined in this repository. trust_remote_code=True is required because CagliostroForCausalLM is not part of transformers.

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "bench-labs/cagliostro-v3.5"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True, dtype=torch.float32)

ids = tok("The capital of France is", return_tensors="pt")
out = model.generate(**ids, max_new_tokens=32, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))

This is a base model with no instruction tuning and no chat template. It completes text.

Reproducing the evaluation

pip install lm-eval
python -m lm_eval --model hf \
  --model_args pretrained=bench-labs/cagliostro-v3.5,dtype=float32,trust_remote_code=True \
  --tasks hellaswag,arc_easy,arc_challenge,piqa \
  --num_fewshot 0 --batch_size 64 --device cuda:0

ArithMark-3 uses the script linked from the leaderboard, pointed at the same model id with --dtype float32. Evaluate in float32. A bfloat16 round trip moves logits enough at this logit cap to change borderline multiple-choice answers.

These exact commands, run against this repository, return the numbers in the results table.

Limitations

English only. 2,048 token context. No instruction tuning, no safety tuning, no RLHF. At 146M parameters it confabulates freely and should not be relied on for factual questions. ARC-Easy is about a point below cagliostro-v3. The mathematics ability measured by ArithMark is arithmetic and short symbolic work, not general mathematical reasoning.

License

Apache-2.0. The training data is drawn from FineWeb-Edu and FineMath (ODC-By), DCLM-Baseline and OpenMathInstruct-2 (CC-BY-4.0), Cosmopedia v1, Cosmopedia v2 and SmolTalk (Apache-2.0).

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bench-labs/cagliostro-v3.5

Finetuned
(1)
this model

Datasets used to train bench-labs/cagliostro-v3.5