AI Agent 评测闭环

原名:agent-platform-eval-flywheel

建立从线上样本、评测数据集、自动评估到改进验证的智能体迭代流程。

中文 Skills 技能说明

适合智能体已经有真实任务,需要用可重复数据判断版本是否变好。样本必须脱敏并得到授权,评分标准要能解释;发布新版本、调用外部模型或把线上数据纳入评测前,需确认账户、成本、数据去向和人工验收。

上游能力依据

上游原始适用说明:Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results before and after a fix, or when guidance is needed on Agent Platform eval methodology — including dataset schema, LLM-as-judge scoring, and common failure causes. For fine-tuning, use agent-platform-tuning. For general production deployment, use agent-platform-deploy.

上游 SKILL.md 主要章节(保留原文标题):

使用边界

先读取项目约定、依赖版本和现有测试,再提出最小改动。写文件、运行脚本、提交代码或发布前应检查差异并保留回退路径。

作者、翻译与许可证

原作者
google
中文翻译
CEOFans翻译
许可证
Apache-2.0
上游来源
https://github.com/google/skills/tree/8db47666a1307cb838f056ac03193f85c42e2c69/skills/cloud/agent-platform-eval-flywheel

适用范围

平台:linux、macos、windows;标签:Google Cloud、Agent、Platform、Eval

安全提示

基础静态扫描不等于绝对安全。技能可能调用命令、浏览器、云服务或本地文件,请在最小权限环境中使用,高风险操作必须人工确认。

如发现侵权、许可证或安全问题,可在本页前台提交投诉,管理员复核后可立即下架。