Sierra · 2026-09-08 · REVIEWED 2026.09.10

Hyper-τ-bench 把“让代理造客服代理”变成可验收问题

Hyper-τ-bench evaluates agents that build customer-service agents

开源评测 · 非产品能力承诺

Open benchmark / not a product-capability claim

Jingwen He

要点

What matters

Sierra 9 月 8 日开源 hyper-τ-bench:开发代理要从模拟企业的记录中恢复需求、构建客服代理,并在未见过的模拟流量上验收,而不只看能否生成代码。

Sierra open-sourced hyper-τ-bench on September 8. A developer agent must recover requirements from simulated business records, build a support agent, and be assessed on unseen simulated traffic—not merely whether it generates code.

范围与事实边界

Scope and limits

Sierra 报告的最强自动配置通过 23.9% 的评测模拟,专家与前沿模型的参考配置为 82.2%。这是特定基准和报告结果,不是任一供应商在真实客服中的通用成功率。

Sierra reports 23.9% of evaluation simulations for its strongest automated configuration and 82.2% for an expert-plus-frontier-model reference. Those are reported, benchmark-specific results—not a general production success rate for any vendor.

编辑工作台

Editorial workbench

把自家客服代理的验收拆成需求回收、工具调用、越权处理、人工交接和盲测五项;上线前用未参与编写答案的人维护一组隐藏案例。

Split a support-agent acceptance plan into requirement recovery, tool use, authorization handling, human handoff and blind tests; maintain hidden cases with people who did not author the answers.

核验来源

Sources checked

← 返回今日← Back to today