code-daemon-summary-v1

A compact bilingual (English / Russian) code-documentation writer — a distilled Qwen3.5-4B in GGUF. It writes:

  • one-sentence descriptions of a file's entities (functions, classes, fields) as a markdown bullet list;
  • module overviews — short prose plus ASCII architecture and data-flow diagrams;
  • hierarchical summaries — subsystem, product and whole-project digests built from smaller ones.

It is the long-output worker of the Code-Daemon daemon — a purpose-built component, not a general assistant. The output language follows the request; both languages were distilled first-class.

What makes it different: a 27B teacher's documentation style distilled into a 4B model that runs on a 6–8 GB laptop GPU at the stock base's speed — and a Q3 file that keeps the whole gain.

The weights under this id changed on 2026-09-16. Until then it served a Qwen3-4B distilled from Qwen2.5-7B-Instruct; its card is in this repository's git history. Do not mix figures across the two.

Numbers

200 held-out documentation prompts, scored against the teacher's answer, greedy decoding, every model run as its GGUF file through llama.cpp:

model (GGUF) size ROUGE-L token F1
this model, Q4_K_M 2.58 GB 0.548 0.628
this model, Q3_K_M + importance matrix 2.16 GB 0.543 0.623
stock Qwen3.5-4B, Q4_K_M 2.58 GB 0.464 0.555
the Qwen3-4B this id served before 2.38 GB 0.414 0.479
  • +0.084 ROUGE-L over its own base (95 % CI +0.066 … +0.101, 156 wins / 42 losses).
  • Q3 costs nothing measurable against Q4 (−0.004, CI crosses zero) and saves 0.4 GB — the domain importance matrix is what makes it hold.
  • The gain shrinks with prompt length and stays positive: +0.12 below 512 prompt tokens, +0.04 above 2 048.
  • In a real documentation pass over one Go repository it wrote tidier output than its base — diagram lines repeating an earlier one 26 % vs 42 %, bullets that only restate the name 1.7 % vs 2.3 % — but dropped 2 sections the base kept. One repository: a direction, not a verdict.

Speed and memory

Laptop RTX 5060 8 GB, llama.cpp CUDA:

Q4_K_M Q3_K_M
decode, one slot 88.6 tok/s 77.0 tok/s
VRAM, 28 parallel sequences 4 710 MiB 4 284 MiB
VRAM, 12 parallel sequences — 3 408 MiB

Inside the daemon, sharing the GPU with its other workers, the Q4 runs 86 tok/s on one slot and 154 tok/s across four — the same speed as the stock base.

The trade between the two files: Q3 is 13 % slower to decode and ~0.4 GB smaller on disk and in VRAM, at the same quality. Take Q4 when it fits, Q3 on a 6 GB card.

How to use it

The bundled chat template runs the model in non-thinking mode — use it as a normal ChatML model. If you build raw prompts, end them with the assistant tag and an empty think block, as below; that is the shape it was trained on. Greedy decoding, stop on <|im_end|>.

llama-cli -m code-daemon-summary-v1-Q4_K_M.gguf -c 8192 --temp 0 \
  -p '<|im_start|>system
You write one-sentence descriptions for code entities of a single file. Output ONLY a markdown bullet list, ONE bullet per entity: - **<EntityName>**: <one-sentence description>.<|im_end|>
<|im_start|>user
Entities: parseArray, encodeValue. File excerpt: <...><|im_end|>
<|im_start|>assistant
<think>

</think>

'

Output shapes: entity docs as - **Name**: one sentence.; module overviews as ## Overview prose plus ## Architecture / ## Flow diagrams; summaries as paragraph-length digests.

Sizing memory. The model is hybrid: 24 of its 32 blocks are gated DeltaNet and carry a recurrent state of 50 MiB per parallel sequence, whatever the context length. To fit a small card, cut sequences, not n_ctx. If you embed llama.cpp as a library, set n_outputs_max to your sequence count — the default reserves logits for a whole batch over a 248k vocabulary (~1.9 GB never used); llama-server already does this.

How it was made

  • Base: Qwen/Qwen3.5-4B — 32 blocks, every fourth full attention, the rest gated DeltaNet; ChatML.
  • Teacher: Qwen3.8-27B.
  • Method: sequence-level knowledge distillation on the bilingual documentation tasks above; merged into the base in full precision and converted to GGUF without the base's multi-token-prediction block (so there is no built-in speculative head). The Q3 is quantized with a domain importance matrix.

Files

file size use
code-daemon-summary-v1-Q4_K_M.gguf 2.58 GB recommended
code-daemon-summary-v1-Q3_K_M.gguf 2.16 GB 6 GB-VRAM cards

Both need llama.cpp with the qwen35 architecture (b10809 / v0.4.0 or newer).

License & attribution

Apache-2.0, matching the Qwen/Qwen3.5-4B base. Not legal advice — check the base and teacher model cards before redistributing. Base and teacher © the Qwen team; please also honour their cards.

Downloads last month
165
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for faxenoff/code-daemon-summary-v1

Finetuned
Qwen/Qwen3.5-4B
Quantized
(512)
this model