Arize 模型实验

原名:arize-experiment

创建、运行并比较模型或提示词实验,查看不同版本的真实表现。

中文 Skills 技能说明

适合模型升级、提示词调整或 RAG 改动后的对比验证。实验需固定数据集、指标、模型版本和参数,不能只选有利结果;运行外部实验会产生调用和数据写入,执行前确认账户与成本。

上游能力依据

上游原始适用说明:Creates, runs, and analyzes Arize experiments for evaluating and comparing model performance. Covers experiment CRUD, exporting runs, comparing results, and evaluation workflows using the ax CLI. Use when the user mentions create experiment, run experiment, compare models, model performance, evaluate AI, experiment results, benchmark, A/B test models, or measure accuracy.

上游 SKILL.md 主要章节(保留原文标题):

使用边界

先读取项目约定、依赖版本和现有测试,再提出最小改动。写文件、运行脚本、提交代码或发布前应检查差异并保留回退路径。

作者、翻译与许可证

原作者
GitHub, Inc. 与 awesome-copilot contributors
中文翻译
CEOFans翻译
许可证
MIT
上游来源
https://github.com/github/awesome-copilot/tree/83561bd7d8a46fcda0581aedabdf8eac7cb196b6/skills/arize-experiment

适用范围

平台:linux、macos、windows;标签:AI 工程、Arize 模型实验

安全提示

基础静态扫描不等于绝对安全。技能可能调用命令、浏览器、云服务或本地文件,请在最小权限环境中使用,高风险操作必须人工确认。

如发现侵权、许可证或安全问题,可在本页前台提交投诉,管理员复核后可立即下架。