要点
What matters
Sierra 9 月 8 日开源 hyper-τ-bench:开发代理要从模拟企业的记录中恢复需求、构建客服代理,并在未见过的模拟流量上验收,而不只看能否生成代码。
Sierra open-sourced hyper-τ-bench on September 8. A developer agent must recover requirements from simulated business records, build a support agent, and be assessed on unseen simulated traffic—not merely whether it generates code.
范围与事实边界
Scope and limits
Sierra 报告的最强自动配置通过 23.9% 的评测模拟,专家与前沿模型的参考配置为 82.2%。这是特定基准和报告结果,不是任一供应商在真实客服中的通用成功率。
Sierra reports 23.9% of evaluation simulations for its strongest automated configuration and 82.2% for an expert-plus-frontier-model reference. Those are reported, benchmark-specific results—not a general production success rate for any vendor.
编辑工作台
Editorial workbench
把自家客服代理的验收拆成需求回收、工具调用、越权处理、人工交接和盲测五项;上线前用未参与编写答案的人维护一组隐藏案例。
Split a support-agent acceptance plan into requirement recovery, tool use, authorization handling, human handoff and blind tests; maintain hidden cases with people who did not author the answers.