跳到正文
Apple Machine Learning Research·· 4 天前精选评分85

RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback

摘要

这条英文动态主要涉及智能体工作流、模型能力与工程、教育应用。原文要点:The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RL...

本站只提供摘要与原文入口。完整内容请阅读原文。

来源:Apple Machine Learning Research · machinelearning.apple.com