跳到正文
  1. Apple Machine Learning Research85

    RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback

    这条英文动态主要涉及智能体工作流、模型能力与工程、教育应用。原文要点:The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RL...

    推荐理由:来自Apple Machine Learning Research的《RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback》。重点看智能体工作流、模型能力与工程、教育应用。摘要提到:这条英文动态主要涉及智能体工作流、模型能力与工程、教育应用。原文要点:The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill f...

  1. Apple Machine Learning Research89

    A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization

    这条英文动态主要涉及模型能力与工程、教育应用、文化创意。原文要点:Semi-supervised federated learning (SSFL) trains models on clients’ unlabeled data using a teacher to generate pseudo-labels, with a small labeled seed dataset on the server. Automatic Speech Recognition (ASR) is particularly fragile here: pseudo-label errors compound across the output sequence and across training rounds into divergence, leaving a large gap to fully-supervised FL. We show that closing this gap turns ...

    推荐理由:来自Apple Machine Learning Research的《A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization》。重点看模型能力与工程、教育应用、文化创意。摘要提到:这条英文动态主要涉及模型能力与工程、教育应用、文化创意。原文要点:Semi-supervised federated learning (SSFL) trains models on clients’ unlabeled data using a teacher to generate pseudo-labels, with a small labeled seed dataset on the server. Automatic Speech Recognition (ASR) is particularly fragile here: pseudo-label errors compound across the output sequence and across training rounds into divergence, leaving a large gap to fully-supervised FL. We ...

已经到底了