Quite a few models have been released in a month: GPT5.6, Grok 4.5, Gemini 3.6, Muse Spark 1.1, Kimi K3, Qwen 3.8 Max, Long cat 2.0, Laguna S 2.1, Ling 3.0, etc
But I will stick to testing small param LLM for my edge device.
Nanbeige4.2-3B that claims to beat the likes of Gemma4-E4B and Qwen3.5 9B.
Gonna test this model in the weekend. Tested many small LLM <4B in Jetson Orin Nano, the followings are the one I would consider: Ministral3-3B - Best overall with tool-calling, vision and tons of hallucinations (currently using) MiniCPM-V-4.6 1.3B - Good VLM, lacking in chatting/tool-calling Granite4.1-3B - slightly better at agentic tool-calling, faster inference speed Bonsai 8B Q1 - better at agentic tool-calling, similar speed as Ministral3-3B Bonsai 27B Q1 - better than 8B but way too slow Needle 26M - only use for query generation and tool execution, needs JAX Nanbeige4.2-3B - TBD