跳到正文
arXiv AI Education· Jakub Macina, Manu Kapur, Mrinmaya Sachan·· 14 小时前评分60

The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning

摘要

原文摘录(自动中文摘要暂不可用):The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem wi

推荐理由

自动中文摘要尚未通过校验,请核对原文与完整来源。

本站只提供摘要与原文入口。完整内容请阅读原文。

来源:arXiv AI Education · arxiv.org