NPC-governed MCP server that exposes jinxes (NPC-specific tools) as MCP tools, with support for NPC switching via prompts
The npcpy MCP server exhibits pervasive quality issues across naming, descriptions, schemas, and error handling. All 11 tools have minimal descriptions (6-13 words, well below the 50-200 char LLM-optimized baseline), lack proper error handling guidance, and provide no output schema documentation. While input schemas are present with typed parameters, they are sparsely described. Tool naming follows some verb conventions but descriptions are generic and interchangeable, offering no guidance on when to select one tool over another. No evidence of idempotency markers, rate limiting, permission gates, or audit trails. The server is a thin passthrough to stub implementations with no production-grade tooling.
Analyzes the emotional tone of a product review.
Analyzes sentiment.
Checks if review appears authentic or potentially fake.
Extracts mentioned product features from review.
Recognizes images.
Organizes notes.
Parses documents.
All tool descriptions are critically short (6-20 chars). Examples: 'Translates text.' (14 chars), 'Analyzes sentiment.' (18 chars), 'Recognizes images.' (18 chars). Rubric baseline is 50-200 chars for LLM optimization. These minimal descriptions do not explain WHEN to use the tool, WHAT it returns, or HOW it differs from similar tools.
Parameter descriptions are absent or minimal. Example: translate_text has 'text' and 'target_language' with generic descriptions ('Text to translate', 'Target language for translation'). These do not specify format constraints, valid values, length limits, or disambiguation. Replace e.g. with enum constraints' and 'Every parameter needs a description explaining what it controls.' Input parameters lack depth needed for LLM self-correction and validation.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Reviews code.
Summarizes data.
Converts text to speech.
Translates text.
No documented output schemas. Tools are stub implementations (e.g., analyze_sentiment, translate_text have no visible return type hints or response documentation in examples/tool_use_example.py). LLMs need to know what fields to expect so they can plan downstream tool calls' and '100% of A+ tools have documented return types'. Agents cannot chain calls or extract specific fields when output structure is unknown.
No error handling or recovery guidance. Tools offer no actionable error messages. Example: image_recognition with invalid path produces no guidance on alternative formats or fallback tools.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Risk column indicates text_to_speech is WRITE, others are READ_ONLY, but these are not declared in schema or via tool annotations. Agents need to know which calls are safe to retry.' Current spec (2026-07-28) favors tool annotations for clarity.
No pagination, rate limits, or result caps. Tools like parse_documents and organize_notes accept arbitrary input without documented limits.
No permission gates, audit trails, or security model. No evidence of scope declarations or secret injection patterns.
Tool naming lacks specificity and disambiguation. Examples: 'review_code' is ambiguous (does it suggest changes, check syntax, assess quality?). No naming guidelines for agents to distinguish from similar tools.