跳到正文
arXiv AI Education· Anjali Kantharuban, Jonas Mueller·· 4 天前评分40

CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

摘要

这条英文动态主要涉及智能体工作流、评测与基准。原文要点:Recent benchmarks rely on user simulators to evaluate AI agents in multi-turn interaction. While existing simulation techniques demonstrate surface fidelity to human style and behavior, ecologically valid interactive benchmarking also requires alignment in when and how agents fail across simulated and real user populations. We find that existing simulators lack outcome calibration: agreement with observed success rat...

本站只提供摘要与原文入口。完整内容请阅读原文。

来源:arXiv AI Education · arxiv.org