A repository for analyzing and generating awesome skills with evaluation and validation tools
This server has severe definition quality issues. Of 8 tools, 7 are skill evaluation/analysis tools with reasonable descriptions but incomplete schemas, and 1 is a Gerrit permission manager that appears to be a shell script wrapper with minimal MCP integration. Most tools lack explicit input/output schema documentation visible in the source. Parameter descriptions are present for some tools but many lack type information. No evidence of error handling guidance, recovery patterns, or output schema documentation. The skill_evaluator tool suite (tools 1-7) has descriptions ranging from 65-180 chars (within acceptable bounds) but schemas are incomplete, input types are visible (Path, string, dict) but no JSON Schema objects with property definitions. The apply_gerrit_permissions tool (tool 8) is a bash script with minimal MCP wrapper structure visible; it has enum constraint for template parameter but overall lacks production-grade MCP integration patterns.
Analyzes skill scripts for malicious code patterns and data exfiltration risks. Returns safety score, findings, and whether scripts exist.
Applies permission templates to Gerrit repositories via SSH or REST API. Supports standard, restricted, protected-branches, multi-team, and ci-cd templates. Can perform dry-run operations.
Evaluates whether a skill has proper documentation, usage instructions, and examples. Returns functionality score and findings based on content completeness.
Evaluates the trustworthiness of a skill's source based on origin (official sources like OpenAI/Anthropic score highest) and documentation quality. Returns trust score and findings.
Detects prompt injection risks in skill content by scanning for common injection patterns, hidden Unicode characters, and bidirectional text markers. Returns injection score (0-100) and list of findings.
Extracts YAML frontmatter from skill markdown files and returns the parsed metadata along with the body content
No visible JSON Schema definitions for tool inputs. Tools list parameter types (Path, string, dict) but lack formal JSON Schema objects with required fields, property definitions, and constraints. Rubric requires explicit schema for all tools.
Output schemas not documented. No visible specification of what each tool returns, fields returned, or data types. LLM cannot plan chained calls without knowing return structure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-15 | F | 45 | - | v1 |
Infers the category or domain of a skill based on its name and content
Evaluates skill files for security risks including prompt injection, code safety, data privacy, source trust, and functionality. Returns risk assessment with SAFE, USE_WITH_CAUTION, NOT_RECOMMENDED, or DANGEROUS classifications.
No error handling guidance. Tools do not document error conditions, recovery steps, or how to distinguish retryable vs fatal errors. Descriptions lack 'If X fails, try Y' recovery patterns.
apply_gerrit_permissions parameter 'template' has enum constraint (good), but tool accepts passwords/credentials as implicit environment variables (GERRIT_PASSWORD inferred in bash script). Credentials must never be tool parameters; use server-side secret injection.
infer_category tool name is vague. 'infer_category' does not state WHAT is inferred (category of a skill) or WHERE it comes from (name + content). Stronger name: 'categorize_skill' or 'classify_skill_domain'.
No pagination support visible on tools that may return large result sets (skill_evaluator analysis results, category inference). Descriptions do not mention limit/page parameters or result count caps. Risk of context window overflow.
Tool composition risk: apply_gerrit_permissions is described as a WRITE operation with dry_run support, but no companion read tool to verify current state before applying. Agent cannot safely inspect permissions before modifying them.