arXiv AI· Chen Henry Wu, Thomas Zhang, Aditi Raghunathan·· 22 小时前评分55
Off-Policy Merging Beats On-Policy Self-Distillation for Continual Learning
摘要
原文摘录(自动中文摘要暂不可用):Off-Policy Merging Beats On-Policy Self-Distillation for Continual Learning A long-standing goal of AI is a model that can continually learn and improve itself. On post-trained models, supervised finetuning (SFT) on new data often causes poor generalization and catastrophic forgetting. As such, the conventional wisdom is that on-policy training is a prerequi
推荐理由
自动中文摘要尚未通过校验,请核对原文与完整来源。
本站只提供摘要与原文入口。完整内容请阅读原文。
来源:arXiv AI · arxiv.org