Instructions to use bench-labs/cagliostro-v3.5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bench-labs/cagliostro-v3.5 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="bench-labs/cagliostro-v3.5", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("bench-labs/cagliostro-v3.5", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bench-labs/cagliostro-v3.5 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bench-labs/cagliostro-v3.5" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bench-labs/cagliostro-v3.5", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/bench-labs/cagliostro-v3.5
- SGLang
How to use bench-labs/cagliostro-v3.5 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bench-labs/cagliostro-v3.5" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bench-labs/cagliostro-v3.5", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bench-labs/cagliostro-v3.5" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bench-labs/cagliostro-v3.5", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use bench-labs/cagliostro-v3.5 with Docker Model Runner:
docker model run hf.co/bench-labs/cagliostro-v3.5
cagliostro-v3.5
A 146M parameter decoder-only language model. It is cagliostro-v3's final checkpoint with one more short training stage: 0.39B tokens of v3's own cooldown data mixed with how-to and textbook text from Cosmopedia. Same architecture, same tokenizer, 75.39B tokens in total.
It scores 27.49 on the Open SLM Leaderboard metric, above every model listed there at the time of release.
Results
Zero-shot, measured with lm-evaluation-harness and the leaderboard's own ArithMark-3 script, on the float32 weights in this repository.
| Benchmark | Metric | Score |
|---|---|---|
| HellaSwag | acc_norm | 43.41 |
| ARC-Easy | acc_norm | 53.75 |
| ARC-Challenge | acc_norm | 29.35 |
| PIQA | acc_norm | 68.06 |
| ArithMark-3 | acc_norm | 45.30 |
| Open SLM Index | 27.49 |
The Index is the leaderboard's own formula, (N(HellaSwag,25) + N(CombinedARC,25) + N(PIQA,50) + 0.65*N(ArithMark,25)) / 3.65 where N(v,c) = 100(v-c)/(100-c) and CombinedARC is the mean of ARC-Easy and ARC-Challenge.
For context, using the leaderboard's published figures for every other model:
| Model | Params | Tokens | Index |
|---|---|---|---|
| cagliostro-v3.5 | 146M | 75.4B | 27.49 |
| SmolLM2-135M | 135M | 2T | 27.13 |
| cagliostro-v3 | 146M | 75B | 26.55 |
| SmolLM-135M | 135M | 600B | 25.74 |
| GPT-X2.5-135M | 135M | 75B | 25.17 |
| Haidass1.5-143M | 143M | 400B | 25.07 |
The lead over SmolLM2-135M is 0.36 Index on the leaderboard's figures. Re-evaluated in the same harness, SmolLM2-135M scores 27.01 and cagliostro-v3.5 is ahead by 0.48, but a paired bootstrap over benchmark items puts the 95% interval for that gap at [-0.82, +1.75]. On points it is first. On these five benchmarks the two models cannot be told apart with confidence, and we would not claim more than that.
cagliostro-v3 appears in the same-harness charts as it measures today, Index 26.31. Its leaderboard row lists 26.55.
What changed from v3
| Field | Value |
|---|---|
| Starting point | cagliostro-v3 step 762,939, weights and AdamW state |
| Steps | 4,000 |
| Tokens per step | 98,304 |
| Tokens | 0.39B |
| Learning rate | warmup over 200 steps to 3e-4, then cosine to zero |
| Peak learning rate | a tenth of v3's 3e-3 |
| Optimizer | AdamW, weight decay 0.01 |
| Hardware | one RTX PRO 6000 Blackwell |
| Wall clock | about 41 minutes at 162,000 tokens per second |
| Source | Share |
|---|---|
| Cosmopedia v1, WikiHow-style tutorials | 30.0% |
| FineWeb-Edu (deduplicated) | 20.35% |
| Cosmopedia v1, textbooks (OpenStax, Khan Academy, Stanford) | 15.0% |
| Cosmopedia v2 | 13.75% |
| FineMath 3+ | 8.25% |
| OpenMathInstruct-2 | 7.15% |
| SmolTalk | 2.75% |
| DCLM-Baseline | 2.75% |
The bottom six rows are v3's cooldown mixture scaled to 55%. The WikiHow set is 174M tokens, so it was seen about 0.7 times. The textbook slice is the first 150M tokens of those three subsets.
The gain is mostly ArithMark-3 (+2.6 points) and HellaSwag (+0.9), with smaller gains on ARC-Challenge and PIQA. ARC-Easy moved the other way by about a point. That drop also shows up when v3 is trained the same way on its own cooldown mixture, so it comes from the extra low learning rate stage rather than from the new data.
How it was chosen
Six runs started from the same v3 checkpoint, identical except for the data. Before any of them finished we wrote down what a winner had to do: score above SmolLM2-135M's 27.13, beat SmolLM2-135M re-measured in the same harness, and beat v3 with the paired 95% interval above zero. Then a second run of the same recipe with a different data seed had to clear 27.13 as well.
| Run | Index |
|---|---|
| 100% SmolTalk chat data | 26.00 |
| v3's own cooldown mixture | 26.54 |
| + 30% WikiHow, seed 1 | 27.14 |
| + 30% WikiHow, seed 2 | 26.98 |
| + 30% WikiHow + 15% textbooks, seed 1 | 27.31 |
| + 30% WikiHow + 15% textbooks, seed 2 | 27.49 |
WikiHow alone passed once and then failed its second seed, so it was dropped. WikiHow with textbooks passed on both seeds. This repository holds the seed 2 checkpoint, the stronger of the two. The recipe averages 27.40 over both.
Chat data lost points, as it does for most base and instruct pairs on the leaderboard. The benchmarks score raw log-likelihood with no chat template, so text that looks like a conversation does not help.
Contamination checks
No benchmark training, validation or test split was used at any point.
Every new document went through the same 13-gram decontamination v3 used, against HellaSwag validation, ARC test, PIQA validation and ArithMark-3, dropping any document with a match. In a sample of 8,000 documents from the two Cosmopedia sources it dropped none.
The HellaSwag gain deserves a closer look, because about two thirds of HellaSwag is built from real WikiHow articles and Cosmopedia's WikiHow set was generated from real WikiHow titles.
19.3% of HellaSwag's WikiHow validation items share an article title with the Cosmopedia set. For the training split the figure is 18.5%, so the validation items are not singled out. The gain is about the same on items whose title appears in the training data (+1.45) as on items whose title does not (+1.24). ActivityNet items, which have nothing to do with WikiHow, barely moved (+0.09). That pattern is a model getting better at everyday how-to text, not one that has seen the test articles.
Sample outputs
Greedy decoding, float32, first sentence of each completion, unedited. These are all twelve prompts that were tried; the video shows four of them. The full completions are in samples.json.
| Prompt | Completion |
|---|---|
| To solve 47 + 68, first | add the numbers together: \[ 47 + 68 = 115 \] Now, divide 115 by 2 to find the quotient: ... |
| A prime number is | a positive integer greater than 1 that has no positive divisors other than 1 and itself. |
| Photosynthesis is the process by which | plants convert sunlight into chemical energy. |
| How to boil an egg: First, | you need to crack the egg into a bowl. |
| To tie a shoelace, start by | wrapping the yarn around the hook of the shoelace. |
| The water cycle begins when | rain falls onto the ground and turns into water droplets. |
| If a train travels 60 miles in 2 hours, its speed is | 60 miles per hour. |
| Sam had 24 apples and gave away 9, so Sam now has | 24 - 9 = 15 apples. |
| The capital of France is | Paris. |
| Gravity is the force that | pulls objects towards each other. |
| The area of a rectangle is found by | multiplying its length by its width. |
| Plants need sunlight because | they need to grow. |
The misses are typical for this size. It adds 47 and 68 correctly and then divides by 2 for no reason, gets the train's speed wrong, starts boiling an egg by cracking it, and describes a shoelace as yarn.
Model details
| Field | Value |
|---|---|
| Parameters | 146,352,000 |
| Non-embedding parameters | 85.7% |
| Layers | 30 |
| Hidden size | 640 |
| Intermediate size | 1,536 |
| Attention heads | 10 |
| Key/value heads | 5 |
| Attention | Grouped query attention with cross-head subspace attenuation |
| Activation | SwiGLU |
| Normalization | RMSNorm, eps 1e-6 |
| Positional encoding | RoPE, theta 100,000 |
| Context length | 2,048 |
| Vocabulary | 32,768 BPE |
| Embeddings | Tied input and output |
| Logit cap | 15.0 |
| Tokens seen | 75.39B (75B pretraining, 0.39B post-training) |
| Weights | float32 safetensors |
The architecture is defined in this repository. trust_remote_code=True is required because CagliostroForCausalLM is not part of transformers.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "bench-labs/cagliostro-v3.5"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True, dtype=torch.float32)
ids = tok("The capital of France is", return_tensors="pt")
out = model.generate(**ids, max_new_tokens=32, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))
This is a base model with no instruction tuning and no chat template. It completes text.
Reproducing the evaluation
pip install lm-eval
python -m lm_eval --model hf \
--model_args pretrained=bench-labs/cagliostro-v3.5,dtype=float32,trust_remote_code=True \
--tasks hellaswag,arc_easy,arc_challenge,piqa \
--num_fewshot 0 --batch_size 64 --device cuda:0
ArithMark-3 uses the script linked from the leaderboard, pointed at the same model id with --dtype float32. Evaluate in float32. A bfloat16 round trip moves logits enough at this logit cap to change borderline multiple-choice answers.
These exact commands, run against this repository, return the numbers in the results table.
Limitations
English only. 2,048 token context. No instruction tuning, no safety tuning, no RLHF. At 146M parameters it confabulates freely and should not be relied on for factual questions. ARC-Easy is about a point below cagliostro-v3. The mathematics ability measured by ArithMark is arithmetic and short symbolic work, not general mathematical reasoning.
License
Apache-2.0. The training data is drawn from FineWeb-Edu and FineMath (ODC-By), DCLM-Baseline and OpenMathInstruct-2 (CC-BY-4.0), Cosmopedia v1, Cosmopedia v2 and SmolTalk (Apache-2.0).
- Downloads last month
- -
Model tree for bench-labs/cagliostro-v3.5
Base model
bench-labs/cagliostro-v3




