mathsci_300m_greenteam — iter 288 (~302M SFT tokens), THINK
(Formerly mathsci_500m_32k_sft_on_mtfinepdfs_iter288; renamed to reflect the ~300M-token SFT budget actually consumed at this checkpoint.)
PA green-team 32k think SFT, 3/5-of-training checkpoint (the campaign's strongest).
Lineage: NVIDIA-Nemotron-3-Super-120B-A12B Base (chat-init) → 1.0B-token finePDFs
midtraining at seq 32768 (nemo-midtrain-finepdfs-8b)
→ SFT on pa-minimal-green-team-SFT-500m
config math_sci_agentic_500m (500M tokens: 37.7% math_reasoning / 36.3% science_mcq /
16.7% science_research / 9.4% agentic_interactive, short-reasoning-first selection,
reasoning traces intact). Packed at 32,768 (zero truncation); 288 of 478 one-epoch iters,
GBS 32, lr 5e-6, TP1·CP4·EP4·PP22.
Evals (limit 200; vLLM tp4; prefill <think>): ifeval .755 · capabilities mean .600
(mmlu_pro .800, gsm8k .955, aime2025 .233, gpqa .495, cute .605, popqa .510) ·
eval truncation 26.5% @4096 cap · repetition 0.1%.
Tokenizer: geodesic-research/nemotron-think-tokenizer (think tags = special ids 12/13;
chat format byte-identical to upstream Nemotron).
- Downloads last month
- 1