logan7000/cogrpo-n3-strict-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupC-qwen3-best Reinforcement Learning • 2B • Updated Aug 12 • 8 • 1