evaluate-plugin
Установка
npx skills add https://github.com/openai/plugins/tree/5fd93af4cd0c623e020d0cc7e9ce178b4ac1f70f/plugins/plugin-eval/skills/evaluate-pluginСтавит скилл в текущий проект - CLI спросит, для каких агентов. С флагом -g - в домашнюю папку, для всех проектов.
Установи скилл «evaluate-plugin» из https://github.com/openai/plugins/tree/5fd93af4cd0c623e020d0cc7e9ce178b4ac1f70f/plugins/plugin-eval/skills/evaluate-plugin: скопируй эту папку целиком в .claude/skills/evaluate-plugin (для Codex - в .agents/skills/evaluate-plugin). Потом прочитай SKILL.md и коротко скажи, в каких задачах будешь его применять.
Вставьте в Claude Code или Codex, открытый в папке проекта.
В Библиотеке ВайбКода этот скилл открывает исходник: автоматической установки для его формата пока нет. Поставьте командой или промптом.
Скачать ВайбКод · Windows и macOS
Описание
Evaluate a local Codex plugin in engineer-friendly language. Use when the user says
SKILL.md
Исходник на GitHubEvaluate Plugin
Use this skill when the target is a plugin root with .codex-plugin/plugin.json.
Workflow
- Treat "Evaluate this plugin." as the default entrypoint.
- If the request comes in as natural chat language, use
plugin-eval start <plugin-root> --request "<user request>" --format markdownfirst so the user sees the routed local path. - Run
plugin-eval analyze <plugin-root> --format markdown. - Read
Fix Firstbefore drilling into manifest findings, nested skill findings, and code or coverage details. - If the plugin contains multiple skills, summarize the strongest and weakest ones explicitly.
- If the user wants measured usage, switch to "Help me benchmark this plugin." and use the starter benchmark flow.
- If the user wants trend data, compare two JSON outputs with
plugin-eval compare.
Chat Requests To Recognize
Evaluate this plugin.Audit this plugin.Why did this score that way?What should I fix first?Help me benchmark this plugin.What should I run next?
Commands
plugin-eval start <plugin-root> --request "Evaluate this plugin." --format markdown
plugin-eval analyze <plugin-root> --format markdown
plugin-eval start <plugin-root> --request "What should I run next?" --format markdown
plugin-eval compare before.json after.json
plugin-eval report result.json --format html --output ./plugin-eval-report.html
plugin-eval init-benchmark <plugin-root>
plugin-eval benchmark <plugin-root> --dry-run
Reference
../../references/chat-first-workflows.md