tiny-random-DeepseekV4-Flash

A tiny random model for testing, shrunk from deepseek-ai/DeepSeek-V4-Flash: the same architecture, quantization config and checkpoint layout at test sizes. Its key patterns, dtypes and tensor ranks match the real checkpoint's (scripts/extract_layout.py).

DeepSeek's own key names: routed experts packed FP4 (I8 (out, in/2) + E8M0 .scale per 32, via the triton_kernels reference) over block-FP8 attention / shared experts / MTP projections (e4m3 + E8M0 .scale per 128x128), a hash-routed first layer, hyper-connections and an mtp.0 block.

reference/ holds the same weights dequantized to bf16, under the unquantized model's keys: the reference to compare logits against, so a test measures what the load path and kernels add, not the quantization itself.

The weights are random; the outputs mean nothing. scripts/ rebuilds it from the real checkpoint's config.json.

Downloads last month
263
Safetensors
Model size
19.3M params
Tensor type
I64
路
F32
路
BF16
路
F8_E4M3
路
I8
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support