arXiv AI Education· Jakub Macina, Manu Kapur, Mrinmaya Sachan·· 14 小时前评分60
The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning
摘要
原文摘录(自动中文摘要暂不可用):The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning Large language models (LLMs) trained to answer questions are natively poor at teaching. Reinforcement Learning (RL) against a simulated student is a promising approach to improve their pedagogy, but existing RL-trained tutors reward the student's success on the tutored problem wi
推荐理由
自动中文摘要尚未通过校验,请核对原文与完整来源。
本站只提供摘要与原文入口。完整内容请阅读原文。
来源:arXiv AI Education · arxiv.org