.claude/skills/ai-llm-evaluationRuntime, accounts, dependencies, permissions, network behavior and task quality remain untested.
LLM 应用质量评测与回归测试实操手册——从"感觉不错"到"可度量可门禁":评测全景与指标体系(正确性/相关性/忠实度/幻觉率/鲁棒性/效率)、评测集构建(黄金数据集/对抗样本/领域评测集/规模估算)、RAG 系统评测(RAGAS 四指标:忠实度/答案相关性/上下文精度/上下文召回)、幻觉检测与度量(事实性幻觉/提示幻觉/上下文矛盾分类与检测方法)、Prompt 回归测试(用例管理/回归门禁/漂移检测/版本对比)、模型对比选型(评测矩阵/成本质量权衡/多模型 A-B/上线决策)、评测流水线与报告(自动化评测/评分聚合/报告模板/上线门禁)。附零依赖本地工具一键查指标、建评测集清单、看 RAG 指标、出对比矩阵、生成评测报告模板。面向 AI 工程师、测试、产品与质量负责人——与 AI 安全红队测试(测安全)互补,本技能测质量。
These states come from the source or distribution context. None of the entries below are SkillVetAI compatibility test results.
These checks parse the fixed package against dated platform rules. They do not execute the Skill or verify task behavior.
.claude/skills/ai-llm-evaluationRuntime, accounts, dependencies, permissions, network behavior and task quality remain untested.
.agents/skills/ai-llm-evaluationRuntime, accounts, dependencies, permissions, network behavior and task quality remain untested.
skills/ai-llm-evaluationRuntime, accounts, dependencies, permissions, network behavior and task quality remain untested.
This command is recorded from the source ecosystem and resolves the registry's latest release. The fixed release shown on this page should be inspected before adoption.
clawhub install @zhaoxinghua09-cell/ai-llm-evaluationclawhub inspect @zhaoxinghua09-cell/ai-llm-evaluation --version 1.0.0This automated, non-executing scan is bound to this release hash. It is not a safety certification and may contain false positives or false negatives.
This is registry-supplied evidence for the recorded release, not an independent SkillVetAI scan. Check the canonical source for the full report, scanner versions, scope, and current moderation state.
The catalog stores hashes and an inventory summary for change detection. It does not republish the package contents.
sha256:e64d1eb193a6a5150fc4236ec1fbd594cee0f7ef74c8e637a533a2a470c5ef81ATTESTATION.mdLICENSE.mdmanifest.jsonreferences/01-评测全景与指标.mdreferences/02-评测集构建.mdreferences/03-RAG系统评测.mdreferences/04-幻觉检测与度量.mdreferences/05-Prompt回归测试.mdreferences/06-模型对比选型.mdreferences/07-评测流程与报告.mdreferences/08-FAQ.mdSECURITY_AUDIT.mdskill-card.mdSKILL.mdtools/llm_eval_toolkit.pyverify/gen_security_radar.pyverify/llm_eval_quality_test.pyverify/security_results.jsonverify/security-radar.svgv1.0.0: hands-on playbook for LLM application quality evaluation and regression testing - evaluation landscape and metrics (correctness, relevance, faithfulness, hallucination rate, robustness, efficiency), test-set construction (golden sets, adversarial samples, size estimation), RAG evaluation (RAGAS four metrics), hallucination detection and measurement, prompt regression testing with gates and drift detection, model comparison and selection with cost-quality trade-off and A-B testing, evaluation pipeline and reporting templates with launch gates; 5-command zero-dependency toolkit