Arize 模型评估器
原名:arize-evaluator
建立并运行正确性、相关性、忠实度和幻觉等模型评估流程。
中文 Skills 技能说明
适合对智能体回答、检索结果和生产追踪做持续质量评估。模型评分不能代替人工验收,评价标准与提示词要可解释;使用外部评审模型前确认费用、数据区域和敏感信息边界。
上游能力依据
上游原始适用说明:Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and continuous monitoring. Use when the user mentions create evaluator, LLM judge, hallucination, faithfulness, correctness, relevance, run eval, score spans, score experiment, trigger-run, column mapping, continuous monitoring, or improve evaluator prompt.
上游 SKILL.md 主要章节(保留原文标题):
- Prerequisites
- Concepts
- What is an Evaluator?
- What is a Task?
- Data Granularity
- How trace and session aggregation works
- The {conversation} template variable
- Multi-evaluator tasks
- Basic CRUD
- AI Integrations
使用边界
先读取项目约定、依赖版本和现有测试,再提出最小改动。写文件、运行脚本、提交代码或发布前应检查差异并保留回退路径。
作者、翻译与许可证
- 原作者
- GitHub, Inc. 与 awesome-copilot contributors
- 中文翻译
- CEOFans翻译
- 许可证
- MIT
- 上游来源
- https://github.com/github/awesome-copilot/tree/83561bd7d8a46fcda0581aedabdf8eac7cb196b6/skills/arize-evaluator
适用范围
平台:linux、macos、windows;标签:AI 工程、Arize 模型评估器
安全提示
基础静态扫描不等于绝对安全。技能可能调用命令、浏览器、云服务或本地文件,请在最小权限环境中使用,高风险操作必须人工确认。
如发现侵权、许可证或安全问题,可在本页前台提交投诉,管理员复核后可立即下架。