MCP server for task guidance with hierarchical RAG and hybrid search
Task Guide MCP exhibits moderate definition quality with a critical language inconsistency flaw, partial schema completeness, and minimal error handling guidance. Of the 8 tools, 6 have complete input schemas with descriptions, but 2 parameter descriptions are in Korean, breaking the expected English interface contract. Descriptions are present but terse (average ~30 chars), below the baseline of 194 chars. Tool names follow verb_noun convention appropriately (create_*, update_*, list_*, get_*, delete_*, index_*, search_, build_*). No output schemas documented, agents cannot predict response structure for chaining. Error handling is not evident in the implementation review provided. Parameter descriptions lack actionable constraints (e.g., 'limit' defaults to 10 with no bounds documented, 'threshold' defaults to 0.5 with no explanation of scale). The server targets a specialized RAG + guidance domain and demonstrates attempt at structured design, but falls short of production-grade definition standards due to incomplete metadata and missing guidance.
Builds hierarchical structure of codebase
Creates a new task guidance
Deletes a task guidance
Retrieves a specific task guidance
Indexes codebase and documents for a guidance
Lists all task guidances
Performs hybrid search
Korean language in parameter descriptions breaks contract. 'get_guidance' and 'delete_guidance' have parameter 'id' described as '가이드 ID' instead of English 'Guide ID'. This breaks automated schema parsing and violates localization expectations for MCP tool interfaces.
No output schemas documented. Agents cannot predict response structure, fields available, or downstream tool parameter matches. Source code does not show response structure definitions.
Descriptions are too terse (average 30 chars vs baseline 194 chars). Most tool descriptions are 2-4 words ('Lists all task guidances', 'Deletes a task guidance') without context for WHEN to use them or WHAT they return. This forces LLMs to guess selection criteria.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Updates an existing task guidance
Parameter constraints missing for numeric values. 'search' tool has 'limit' (optional, default: 10) and 'threshold' (optional, default: 0.5) with no documented min/max bounds. Unbounded numbers let LLMs pass absurd values (limit=999999, threshold=0.001) that may break indexing logic.
No error handling guidance in tool descriptions. Tools like 'delete_guidance' and 'create_guidance' are destructive/state-modifying but provide no guidance on error recovery, retryability, or confirmation patterns. Should I ask the user? Or is this unrecoverable?
Parameter 'externalDocs' in 'index_guidance' is under-specified. It's an array of strings described as 'External document paths (optional)' but lacks format guidance (file paths? URLs? relative paths?). This invites malformed input and potential path traversal issues.
'build_hierarchy' tool name is vague. Does it return a tree structure? A flat list? A graph? 'build' is a weak verb, consider 'get_codebase_hierarchy' or 'analyze_directory_structure'. LLMs cannot infer from 'build_hierarchy' when to use it vs 'index_guidance'.
Enum values for 'type' parameter in 'search' are not explained. The description lists ['code', 'document', 'guidance', 'all'] but does not explain what each searches (indexed code vs RAG docs vs guidance text?) or when to use each. This forces LLMs to guess or call all variants.