arxiv:2605.15726
Chanuk Lee
tally0818
AI & ML interests
LLM post-training
Recent Activity
upvoted a paper 1 day ago
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning upvoted a paper 1 day ago
DAPD: Dual-Anchored Policy Distillation upvoted a paper 1 day ago
On-Policy Self-Distillation without Any SupervisionOrganizations
None yet