MCP server for generating and embedding documentation from Git repositories, with semantic search capabilities
The server defines 3 tools with explicit schemas and descriptions. Tool names follow verb_noun conventions (embed_repo, query_embeddings, check_embed_status) and are action-oriented. All tools have non-trivial descriptions (90-120 chars). Input parameters are properly typed with JSON Schema and include descriptions. However, several critical gaps prevent a higher score: (1) output schemas are not documented, no indication of what embed_repo, query_embeddings, or check_embed_status return; (2) parameter descriptions lack format constraints and ranges (e.g., 'limit' is u64 with no bounds specified, should state 1 - 1000 or similar); (3) error handling guidance is missing, no indication of what errors are retryable vs. fatal, or what the LLM should do if embedding fails; (4) no mention of idempotency, dry-run, or confirmation for the destructive embed_repo operation; (5) 'query_embeddings' lacks pagination (no offset/limit pattern shown in response), which risks returning unbounded result sets. The repository input parsing is well-designed (accepts both full URLs and owner/repo shorthand via custom deserializer), which is a strength.
Check the status of an ongoing repository embedding operation
Generate and embed documentation from a Git repository
Perform semantic search on repository documentation embeddings
Output schemas not documented for any tool. Callers have no visibility into what embed_repo returns (success object? operation tracking data?), what query_embeddings yields (array of matches? pagination?), or what check_embed_status provides (status enum string? detailed progress object?). LLMs cannot plan downstream processing without knowing output structure.
Parameter 'limit' in query_embeddings lacks explicit bounds. Schema shows u64 type but no min/max constraints. This violates mxe:enforce-result-limits, unbounded integers let LLMs pass absurd values. Should explicitly state limit range (e.g., 1 - 100, default 10).
No error handling guidance. embed_repo may fail (invalid repo, network error, quota exceeded, rate limit). Tool descriptions do not state whether errors are retryable, what the LLM should do on failure, or provide recovery hints (e.g., 'If network fails, retry in 30s. If quota exceeded, contact admin'). Violates pattern:recovery-guide.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 74 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 46 | - | v1 |
embed_repo is a WRITE operation but lacks confirmation or dry-run semantics. Code shows it initiates background embedding (line 'tokio::spawn(async move {'), which is irreversible. No tool to cancel an operation in progress or preview what will be embedded. Violates pattern:confirmation-request for destructive operations.
query_embeddings returns results with default limit=10, but description does not clarify pagination strategy. Is there a next_cursor? Are results sorted by relevance score? If >10 matches exist, how does caller fetch the rest? Violates pattern:paginated-result, tools returning lists must explain pagination mechanism.
operation_id format in check_embed_status is inferred from code ('embed_{repo_name}_{uuid}') but not documented in parameter description. LLMs cannot know what ID to pass without calling embed_repo first. Should explicitly state format or accept repo_url as an alternative lookup key.