Instructions to use faxenoff/code-daemon-summary-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use faxenoff/code-daemon-summary-v1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf faxenoff/code-daemon-summary-v1:Q3_K_M # Run inference directly in the terminal: llama cli -hf faxenoff/code-daemon-summary-v1:Q3_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf faxenoff/code-daemon-summary-v1:Q3_K_M # Run inference directly in the terminal: llama cli -hf faxenoff/code-daemon-summary-v1:Q3_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf faxenoff/code-daemon-summary-v1:Q3_K_M # Run inference directly in the terminal: ./llama-cli -hf faxenoff/code-daemon-summary-v1:Q3_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf faxenoff/code-daemon-summary-v1:Q3_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf faxenoff/code-daemon-summary-v1:Q3_K_M
Use Docker
docker model run hf.co/faxenoff/code-daemon-summary-v1:Q3_K_M
- LM Studio
- Jan
- vLLM
How to use faxenoff/code-daemon-summary-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "faxenoff/code-daemon-summary-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "faxenoff/code-daemon-summary-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/faxenoff/code-daemon-summary-v1:Q3_K_M
- Ollama
How to use faxenoff/code-daemon-summary-v1 with Ollama:
ollama run hf.co/faxenoff/code-daemon-summary-v1:Q3_K_M
- Unsloth Desktop
- Pi
How to use faxenoff/code-daemon-summary-v1 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf faxenoff/code-daemon-summary-v1:Q3_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "faxenoff/code-daemon-summary-v1:Q3_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use faxenoff/code-daemon-summary-v1 with Docker Model Runner:
docker model run hf.co/faxenoff/code-daemon-summary-v1:Q3_K_M
- Lemonade
How to use faxenoff/code-daemon-summary-v1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull faxenoff/code-daemon-summary-v1:Q3_K_M
Run and chat with the model
lemonade run user.code-daemon-summary-v1-Q3_K_M
List all available models
lemonade list
- Hermes Agent
How to use faxenoff/code-daemon-summary-v1 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf faxenoff/code-daemon-summary-v1:Q3_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default faxenoff/code-daemon-summary-v1:Q3_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use faxenoff/code-daemon-summary-v1 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf faxenoff/code-daemon-summary-v1:Q3_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "faxenoff/code-daemon-summary-v1:Q3_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
code-daemon-summary-v1
A compact bilingual (English / Russian) code-documentation writer — a distilled Qwen3.5-4B in GGUF. It writes:
- one-sentence descriptions of a file's entities (functions, classes, fields) as a markdown bullet list;
- module overviews — short prose plus ASCII architecture and data-flow diagrams;
- hierarchical summaries — subsystem, product and whole-project digests built from smaller ones.
It is the long-output worker of the Code-Daemon daemon — a purpose-built component, not a general assistant. The output language follows the request; both languages were distilled first-class.
What makes it different: a 27B teacher's documentation style distilled into a 4B model that runs on a 6–8 GB laptop GPU at the stock base's speed — and a Q3 file that keeps the whole gain.
The weights under this id changed on 2026-09-16. Until then it served a Qwen3-4B distilled from Qwen2.5-7B-Instruct; its card is in this repository's git history. Do not mix figures across the two.
Numbers
200 held-out documentation prompts, scored against the teacher's answer, greedy decoding, every model run as its GGUF file through llama.cpp:
| model (GGUF) | size | ROUGE-L | token F1 |
|---|---|---|---|
| this model, Q4_K_M | 2.58 GB | 0.548 | 0.628 |
| this model, Q3_K_M + importance matrix | 2.16 GB | 0.543 | 0.623 |
| stock Qwen3.5-4B, Q4_K_M | 2.58 GB | 0.464 | 0.555 |
| the Qwen3-4B this id served before | 2.38 GB | 0.414 | 0.479 |
- +0.084 ROUGE-L over its own base (95 % CI +0.066 … +0.101, 156 wins / 42 losses).
- Q3 costs nothing measurable against Q4 (−0.004, CI crosses zero) and saves 0.4 GB — the domain importance matrix is what makes it hold.
- The gain shrinks with prompt length and stays positive: +0.12 below 512 prompt tokens, +0.04 above 2 048.
- In a real documentation pass over one Go repository it wrote tidier output than its base — diagram lines repeating an earlier one 26 % vs 42 %, bullets that only restate the name 1.7 % vs 2.3 % — but dropped 2 sections the base kept. One repository: a direction, not a verdict.
Speed and memory
Laptop RTX 5060 8 GB, llama.cpp CUDA:
| Q4_K_M | Q3_K_M | |
|---|---|---|
| decode, one slot | 88.6 tok/s | 77.0 tok/s |
| VRAM, 28 parallel sequences | 4 710 MiB | 4 284 MiB |
| VRAM, 12 parallel sequences | — | 3 408 MiB |
Inside the daemon, sharing the GPU with its other workers, the Q4 runs 86 tok/s on one slot and 154 tok/s across four — the same speed as the stock base.
The trade between the two files: Q3 is 13 % slower to decode and ~0.4 GB smaller on disk and in VRAM, at the same quality. Take Q4 when it fits, Q3 on a 6 GB card.
How to use it
The bundled chat template runs the model in non-thinking mode — use it as a normal ChatML
model. If you build raw prompts, end them with the assistant tag and an empty think block, as
below; that is the shape it was trained on. Greedy decoding, stop on <|im_end|>.
llama-cli -m code-daemon-summary-v1-Q4_K_M.gguf -c 8192 --temp 0 \
-p '<|im_start|>system
You write one-sentence descriptions for code entities of a single file. Output ONLY a markdown bullet list, ONE bullet per entity: - **<EntityName>**: <one-sentence description>.<|im_end|>
<|im_start|>user
Entities: parseArray, encodeValue. File excerpt: <...><|im_end|>
<|im_start|>assistant
<think>
</think>
'
Output shapes: entity docs as - **Name**: one sentence.; module overviews as ## Overview prose
plus ## Architecture / ## Flow diagrams; summaries as paragraph-length digests.
Sizing memory. The model is hybrid: 24 of its 32 blocks are gated DeltaNet and carry a
recurrent state of 50 MiB per parallel sequence, whatever the context length. To fit a small
card, cut sequences, not n_ctx. If you embed llama.cpp as a library, set n_outputs_max to your
sequence count — the default reserves logits for a whole batch over a 248k vocabulary (~1.9 GB
never used); llama-server already does this.
How it was made
- Base:
Qwen/Qwen3.5-4B— 32 blocks, every fourth full attention, the rest gated DeltaNet; ChatML. - Teacher:
Qwen3.8-27B. - Method: sequence-level knowledge distillation on the bilingual documentation tasks above; merged into the base in full precision and converted to GGUF without the base's multi-token-prediction block (so there is no built-in speculative head). The Q3 is quantized with a domain importance matrix.
Files
| file | size | use |
|---|---|---|
code-daemon-summary-v1-Q4_K_M.gguf |
2.58 GB | recommended |
code-daemon-summary-v1-Q3_K_M.gguf |
2.16 GB | 6 GB-VRAM cards |
Both need llama.cpp with the qwen35 architecture (b10809 / v0.4.0 or newer).
License & attribution
Apache-2.0, matching the Qwen/Qwen3.5-4B base.
Not legal advice — check the base and teacher model cards before redistributing. Base and teacher
© the Qwen team; please also honour their cards.
- Downloads last month
- 165
3-bit
4-bit